A Method for Analyzing the Market Segmentation of Tourism Consumers Based on Online Reviews

The method transforms online reviews into structured vectors using LDA and clustering algorithms, addressing the precision and comprehensiveness issues in tourism market segmentation, enabling targeted product strategies.

CN115018526BActive Publication Date: 2025-07-15HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210466609.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-07-15
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

The existing tourism market segmentation analysis methods are not accurate, the coverage information is not comprehensive, and the consumer demand information cannot be accurately obtained.

Method used

By obtaining online comment data, word segmentation processing and LDA theme model analysis, product attribute feature lexicon and emotional lexicon are constructed, SOM algorithm and slime mold algorithm are used for clustering, multi-dimensional score vectors are generated and segmented analysis is performed.

Benefits of technology

A more accurate market segmentation of tourism consumers can be achieved, which can better position consumers' consumption preferences and help online travel agencies formulate products and strategies suitable for different market segments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115018526B_ABST
    Figure CN115018526B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method for analyzing the market segmentation of tourism consumers based on online reviews, belonging to the field of tourism analysis. The method includes: obtaining online data reviews; performing word segmentation on the online data reviews to divide each sentence review in the online data reviews into multiple words; obtaining noun words in the words; obtaining new words in the online data reviews. This method can obtain online data reviews on the Internet and further analyze and summarize the online data reviews to divide tourism consumers into different segmented markets, and then can conduct more refined analysis of consumers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of tourism analysis, and particularly to a method for analyzing and segmenting the tourism consumer market based on online reviews. Background Art

[0002] The emergence of information technology and digital platforms has had a significant impact on various industries. The tourism industry is one of the most affected industries. By the end of 2019, the scale of online tourism users in China reached 403 million, an increase of 38 million compared with 2018, accounting for 40.7% of all Internet users. The development of the Internet has brought about great changes in the information retrieval and purchasing methods of travelers. Many small and medium-sized travel agencies have lost their markets in this wave. With the continuous development of Internet tourism, online travel agencies have emerged during this period. Online travel agencies attract more and more users with the advantages of low cost, customized services, wide coverage, convenience and speed. Some customers choose online travel agencies because of their convenience on mobile terminals, while other consumers choose them because online travel agencies can maximize customer value and provide better perceived value, thus enhancing loyalty. The booming development of online tourism has led to an increasing demand for consumer personalization, which poses new requirements for travel agency managers and marketers. Market segmentation has become an important link among them.

[0003] Nowadays, the review data of tourism products by consumers on online travel agency websites has become an important data source for travelers before they set out. Obtaining effective data from these reviews and improving products has also become an important method for online travel agencies to create more profits. After traveling, customers feedback their experiences in the form of reviews on social media such as online travel websites. These data not only include users' personal information and travel experiences, but also carry real demand information, thus becoming an important data source for segmenting the market.

[0004] The tourism industry's strong dependence on information data makes market segmentation driven by data an important research method. The rise of data analysis technology has also brought important research tools to this field. However, the existing methods for analyzing and segmenting the tourism market generally have low accuracy and incomplete coverage of information. They can only analyze by comparing data information according to some keywords, making it impossible for enterprises to obtain more accurate information when conducting demand analysis. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a method for analyzing and segmenting the tourism consumer market based on online reviews. This method can obtain online data reviews on the Internet, further analyze and summarize the online data reviews to obtain specific classifications of the online data reviews, and then divide the online tourism user group into several segmented markets, laying a foundation for enterprises to conduct demand analysis and product configuration improvement.

[0006] To achieve the above object, an embodiment of the present invention provides a method for analyzing the market segmentation of tourism consumers based on online reviews, and the method includes:

[0007] Obtain online data reviews;

[0008] Perform word segmentation on the online data reviews to divide each statement review in the online data reviews into multiple words;

[0009] Obtain the noun words in the words;

[0010] Obtain the new words in the online data reviews;

[0011] Combine the noun words and the new words into a noun phrase set;

[0012] Analyze the online data reviews through the LDA model to obtain different topics;

[0013] Incorporate the noun words and the new words in the noun phrase set as attribute feature words into the corresponding topics to obtain a product attribute feature word library;

[0014] Select some sentiment words in the sentiment dictionary as seed sentiment words;

[0015] Obtain the adjective words in the online data reviews;

[0016] Screen the adjective words that meet the standard in terms of the matching degree with the seed sentiment words;

[0017] Combine all the sentiment words in the sentiment dictionary and the adjective words that meet the standard into a sentiment word library;

[0018] Obtain the attribute feature words in the online data reviews that meet the product attribute feature word library;

[0019] Obtain the sentiment words and the adjective words in the online data reviews that meet the sentiment word library near the attribute feature words;

[0020] Obtain the intensity modifiers near the sentiment words in the online data reviews;

[0021] Calculate the evaluation values of the attribute feature words that meet the product attribute feature word library, the sentiment words and the adjective words that meet the sentiment word library, and the intensity modifiers respectively to convert each statement review in the online data reviews into a multi-dimensional score vector;

[0022] Execute the SOM algorithm on the multi-dimensional score vector and analyze to obtain multiple clustering numbers and the neuron weights under each cluster;

[0023] Use the neuron weights as the clustering centers;

[0024] Use the clustering centers as the initial slime mold positions, and perform the slime mold algorithm on the multi-dimensional score vectors to obtain the updated slime mold positions;

[0025] After determining the updated slime mold positions, classify the multi-dimensional score vectors close to the updated slime mold positions according to the nearest neighbor rule;

[0026] Calculate the number of the multi-dimensional score vectors classified into one class;

[0027] Determine the result of the tourism consumer market segmentation analysis according to the clustering of the multi-dimensional score vectors and the quantity under each cluster.

[0028] Optionally, after obtaining the online data comments, clean and trim the online data comments to obtain the preprocessed text.

[0029] Optionally, the new words included in the online data comments are:

[0030] Obtain the combined words that appear in the online data comments;

[0031] Calculate the frequency of the combined words that appear in the online data comments;

[0032] Judge whether the frequency of the combined words that appear in the online data comments is greater than the first threshold;

[0033] If it is judged that the frequency of the combined words that appear in the online data comments is greater than the first threshold, retain the combined words and name them new words.

[0034] Optionally, the method includes:

[0035] After obtaining the new words, retain the new words that conform to the NN, NN NN, JJ NN, NN NN NN, JJ NN NN, JJ JJNN patterns, where NN represents a single noun and JJ represents a single adjective.

[0036] Optionally, the merging of the noun words and the new words into a noun phrase set includes:

[0037] Identify the repeated noun words and new words in the noun phrase set;

[0038] Delete the repeated noun words and new words in the noun phrase set;

[0039] Compare the new words in the noun phrase set with the noun words, and find the new words that coincide with the meaning of the noun words;

[0040] Delete the new words that coincide with the meaning of the noun words;

[0041] Merge the noun words and the remaining new words after deletion into a candidate attribute feature word library.

[0042] Optionally, merging the noun words and the remaining new words after deletion into a candidate attribute feature word library includes:

[0043] Obtain the frequencies of the noun words and the remaining new words after deletion that appear in the online data comments;

[0044] Judge whether the frequencies of the noun words and the remaining new words after deletion that appear in the online data comments are greater than a second threshold;

[0045] In the case where it is judged that the frequencies of the noun words and the remaining new words after deletion that appear in the online data comments are greater than the second threshold, merge the noun words and the remaining new words after deletion with frequencies greater than the second threshold into a candidate attribute feature word library.

[0046] Optionally, incorporating the noun words and the new words in the noun phrase set as attribute feature words under the corresponding themes to obtain a product attribute feature word library includes:

[0047] Determine some of the new words and noun words in the noun phrase set under the corresponding themes, and name them seed attribute words;

[0048] Compare the similarity between the remaining new words and noun words in the candidate attribute feature word library and the seed attribute words;

[0049] Select the remaining new words and noun words with similarity greater than a third threshold and incorporate them into the corresponding themes to obtain a product attribute feature word library.

[0050] Optionally, selecting some sentiment words in the sentiment dictionary as seed sentiment words includes:

[0051] Select some sentiment words in the sentiment dictionary as candidate seed sentiment words;

[0052] Determine the co-occurrence degree of the candidate seed sentiment words and the seed attribute words under each theme according to formula (1);

[0053] (1)

[0054] Wherein, Indicates the number of times the seed attribute word appears alone in the statements of the online data comments. Indicates the number of times the candidate seed sentiment word appears alone in the statements of the online data comments. Indicates the number of times the candidate seed sentiment word and the seed attribute word co - appear in a statement of the online data comments. Indicates the co - occurrence degree of the candidate seed sentiment word and the seed attribute word;

[0055] Determine whether the co - occurrence degree is greater than the fourth threshold;

[0056] When the co - occurrence degree is greater than the fourth threshold, determine that the candidate seed sentiment word is a seed sentiment word under the same theme as the seed attribute word.

[0057] Optionally, screening the adjective words whose matching degree with the seed sentiment word meets the standard includes:

[0058] Match the adjective word with the seed sentiment word;

[0059] Determine whether the matching degree of the adjective word and the seed sentiment word is greater than the fifth threshold;

[0060] When it is determined that the matching degree of the adjective word and the seed sentiment word is greater than the fifth threshold, determine that the matching degree of the adjective word meets the standard.

[0061] Optionally, the method includes:

[0062] Add the general dictionary to the sentiment word library;

[0063] Delete the repeated adjective words in the sentiment word library and the sentiment words in the sentiment dictionary.

[0064] Through the above technical solution, a method for analyzing the segmentation of the tourism consumer market based on online reviews provided by the present invention obtains online data reviews, then processes the online data reviews, obtains noun words from the online data reviews, and then obtains new words from the online data reviews. The noun words and the new words are combined into a noun phrase set, and the noun phrase set is processed to obtain a product attribute feature word library. The sentiment words in the sentiment dictionary and the adjectives obtained from the online data reviews are combined into a sentiment word library. The evaluation values of the attribute feature words that conform to the product attribute feature word library, the sentiment words, the adjective words, and the intensity modifiers that conform to the sentiment word library are calculated respectively to convert each statement review in the online data reviews into a multi-dimensional score vector. The SOM algorithm and the slime mold algorithm are executed on the multi-dimensional vector and according to the nearest neighbor rule to cluster the multi-dimensional score vector into multiple categories, and the quantity under each cluster is obtained. Analysts can obtain a more accurate segmentation result of the tourism consumption market based on the clustering result and the quantity of the online data reviews for subsequent analysis.

[0065] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification, and are used to explain the embodiments of the present invention together with the following specific implementation manners, but do not constitute a limitation to the embodiments of the present invention. In the drawings:

[0067] Figure 1 is the overall flowchart of a method for analyzing the segmentation of the tourism consumer market based on online reviews according to an embodiment of the present invention;

[0068] Figure 2 is the flowchart of new word discovery of a method for analyzing the segmentation of the tourism consumer market based on online reviews according to an embodiment of the present invention;

[0069] Figure 3 is the flowchart of determining a candidate attribute feature word library of a method for analyzing the segmentation of the tourism consumer market based on online reviews according to an embodiment of the present invention;

[0070] Figure 4 is the flowchart of screening the candidate attribute feature word library of a method for analyzing the segmentation of the tourism consumer market based on online reviews according to an embodiment of the present invention;

[0071] Figure 5 is the flowchart of the product attribute feature word library of a method for analyzing the segmentation of the tourism consumer market based on online reviews according to an embodiment of the present invention;

[0072] Figure 6 It is a flow chart for determining seed sentiment words of a method for analyzing and segmenting the tourism consumer market based on online reviews according to an embodiment of the present invention;

[0073] Figure 7 It is a flow chart for screening adjective words of a method for analyzing and segmenting the tourism consumer market based on online reviews according to an embodiment of the present invention;

[0074] Figure 8 It is a flow chart for deleting words from the sentiment word library of a method for analyzing and segmenting the tourism consumer market based on online reviews according to an embodiment of the present invention. Specific embodiments

[0075] The following will describe in detail the specific embodiments of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining and illustrating the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0076] Figure 1 It is a general flow chart of a method for analyzing and segmenting the tourism consumer market based on online reviews according to an embodiment of the present invention. In one embodiment of the invention, the method may include:

[0077] In step S1, online data reviews are obtained.

[0078] In step S2, word segmentation is performed on the online data reviews to divide each sentence review in the online data reviews into multiple words.

[0079] In step S3, noun words in the words are obtained.

[0080] In step S4, new words in the online data reviews are obtained.

[0081] In step S5, the noun words and new words are combined into a noun phrase set.

[0082] In step S6, the online data reviews are analyzed by the LDA model to obtain different topics.

[0083] In step S7, the noun words and new words in the noun phrase set are incorporated as attribute feature words under the corresponding topics to obtain a product attribute feature word library.

[0084] In step S8, some sentiment words in the sentiment dictionary are selected as seed sentiment words;

[0085] In step S9, adjective words in the online data reviews are obtained.

[0086] In step S10, adjective words that match the seed sentiment word to a certain standard are screened out.

[0087] In step S11, all sentiment words in the sentiment dictionary and the adjective words that meet the standard are combined into a sentiment word library.

[0088] In step S12, attribute feature words that match the attribute feature word library in the online data comment are obtained.

[0089] In step S13, sentiment words and adjective words that match the sentiment word library near the attribute feature word in the online data comment are obtained.

[0090] In step S14, intensity modifiers near the sentiment words in the online data comment are obtained.

[0091] In step S15, the evaluation values of the attribute feature words that match the product attribute feature word library, the sentiment words and adjective words that match the sentiment word library, and the intensity modifiers are calculated respectively to convert each statement comment in the online data comment into a multi-dimensional score vector.

[0092] In step S16, the multi-dimensional score vector is analyzed by performing the SOM algorithm to obtain multiple clustering numbers and the neuron weights under each cluster.

[0093] In step S17, the neuron weights are used as the clustering centers.

[0094] In step S18, the clustering centers are used as the initial slime mold positions, and the slime mold algorithm is performed on the multi-dimensional score vector to obtain the updated slime mold positions.

[0095] In step S19, after determining the position of the updated slime mold, the multi-dimensional vectors close to the updated slime mold position are classified into one category according to the nearest neighbor rule.

[0096] In step S20, the number of multi-dimensional vectors classified into one category is calculated.

[0097] In step S21, the results of the tourism consumption market segmentation analysis are determined according to the clustering of the multi-dimensional score vector and the quantity under each cluster.

[0098] In an embodiment of the present invention, after obtaining the online data comment, since there is no gap connection in the Chinese text except for punctuation marks, it is necessary to perform word segmentation processing on it. When performing word segmentation processing on the online data comment, stop word removal processing can be adopted, including common conjunctions and prepositions, etc.

[0099] After dividing the online data comments into multiple words, it is possible to find the noun words among these words and retain them. Then, find the new words among these words and retain them. Combine the retained noun words and new words, and then form a set of noun phrases. After forming a set of noun phrases, the online data comments can be analyzed again. Since there is no unified product design document for online travel products, but there will be several implicit themes in these texts, the LDA model can be used to analyze the texts in the online data comments, so that different themes can be obtained. After obtaining different themes, it can be found that the noun words and new words in the set of noun phrases have different degrees of association with different themes. Therefore, different noun words and new words in the set of noun phrases can be used as attribute feature words and classified according to the degree of association with different themes and assigned to the corresponding themes, so that a product attribute feature word library can be obtained.

[0100] After obtaining the product attribute feature word library, some sentiment words in the existing sentiment dictionary can be used as seed sentiment words, and then the adjective words in the online data comments can be obtained. Match the adjective words with the seed sentiment words, calculate the point mutual information value between the adjective words and the seed sentiment words, and then screen the adjective words that meet the standard of matching with the seed sentiment words, that is, select the adjective words with high point mutual information value. After obtaining the adjective words that meet the standard, all the sentiment words in the sentiment dictionary and the adjective words that meet the standard can be combined into a sentiment word library.

[0101] After obtaining the product attribute feature word library and the sentiment word library, in order to better carry out market segmentation work in the future, it is necessary to convert the unstructured text comment data of the online data comments into structured multi-dimensional vectors, and the online data comments can be further processed. First, it is necessary to find the attribute feature words that meet the product attribute feature word library in each comment sentence of the online data comments, and then retrieve whether there are sentiment words, adjective words and intensity modifiers before and after the attribute feature words. The intensity modifiers can come from the "Word Set for Sentiment Analysis". After obtaining the sentiment words and adjective words that meet the sentiment word library near the attribute feature words in the online data comments, and obtaining the intensity modifiers near the sentiment words and adjective words in the online data comments, the evaluation values of the attribute feature words that meet the product attribute feature word library, the sentiment words and adjective words that meet the sentiment word library, and the intensity modifiers can be calculated respectively to convert each sentence comment in the online data comments into a multi-dimensional score vector. We can confirm the score as 1-5 points according to the sentiment polarity and intensity. 1 point means that the consumer is not satisfied with this attribute, 3 points means that the user has a moderate sentiment towards this attribute, 5 points means very satisfied, and 2 points and 4 points are in the middle respectively, so that each sentence comment in the online data comments can be converted into a multi-dimensional score vector.

[0102] After obtaining the multi-dimensional score vector, the multi-dimensional score vector can be analyzed by performing the SOM algorithm to obtain multiple cluster numbers and the neuron weights under each cluster. Input the structured multi-dimensional score vector into the SOM neural network for clustering, randomize the weights of each neuron in the neural network, and continuously train the model to a certain accuracy, and then output a set of winning neural values, that is, neuron weights. Use the neuron weights as the cluster center, and then use the cluster center as the initial slime mold position to execute the slime mold algorithm. After the multi-dimensional score vector executes the slime mold algorithm, the updated position of the slime mold can be obtained, and this position is the optimal solution to each score vector under each cluster. When performing the slime mold algorithm, it is necessary to calculate and sort the fitness of each slime mold. The fitness of the slime mold can be calculated by formula (2).

[0103] (2)

[0104] is the fitness of the slime mold, is the within-class scatter and, that is, the sum of all multi-dimensional score vectors to the cluster center, is a constant. Under this formula, the fitness of the slime mold is negatively correlated with the scatter sum. The smaller the scatter sum, the greater the fitness. Subsequently, calculate the weights of the slime molds according to the fitness ranking, update the positions of the slime molds, and calculate the new fitness values to update the global optimal solution.

[0105] After obtaining the updated slime mold position, the corresponding slime mold cluster division can be determined according to the nearest neighbor rule. Classify the multi-dimensional score vectors close to the updated slime mold position into one category, and calculate the new cluster center according to the classification, and update the slime mold position according to the new center. After executing the nearest neighbor rule, it can be judged whether the slime mold algorithm converges or reaches the maximum number of iterations. If it does not converge or does not reach the maximum number of iterations, the new cluster center can be used as the slime mold position to execute the slime mold algorithm again. After the slime mold algorithm has converged or reached the maximum number of iterations, the number of multi-dimensional score vectors classified into one category can be calculated, and then the results of the tourism consumer market segmentation analysis can be determined according to the clustering of the multi-dimensional score vectors and the number under each cluster.

[0106] The types divided by this tourism consumer market segmentation analysis method based on online reviews are more detailed, and can more accurately locate the consumption preferences of consumers. Online travel agencies can launch products and strategies more suitable for different segmented markets according to the results obtained by this analysis method.

[0107] In one embodiment of the present invention, after obtaining the online data comments in step S1, the online data comments can be cleaned and trimmed to obtain a preprocessed text. Compared with the online data comments, some comment statements that are too short and some comment statements that are meaningless are deleted, so that the information in the preprocessed text can more effectively reflect the trend of consumers.

[0108] In one embodiment of the present invention, when using the frequently occurring noun words and noun phrases in the online data comments as attribute feature words, relying solely on this method will overly depend on the frequency of the occurrence of noun words or noun phrases, and it is easy to miss uncommon attribute feature words. At the same time, Chinese words often have new meanings after combination. For example, certain scenic spot name word segmentation tools often cannot make correct distinctions. Therefore, an operation of new word discovery is required to obtain a more complete product attribute feature word library. Figure 2 It is a new word discovery flowchart of a tourism consumer market segmentation analysis method based on online reviews according to one embodiment of the present invention. In step S4, new word terms can be obtained. To obtain new word terms, it can include:

[0109] In step S22, the combined words that appear in the online data comments are obtained.

[0110] In step S23, the frequency of the combined words appearing in the online data comments is calculated.

[0111] In step S24, it is judged whether the frequency of the combined words appearing in the online data comments is greater than a first threshold.

[0112] In step S25, when it is judged that the frequency of the combined words appearing in the online data comments is greater than the first threshold, the combined words are retained and named as new word terms.

[0113] After obtaining the combined words in the online data comments, it can be calculated whether the frequency of the combined words appearing in the online data comments is greater than the first threshold. If the frequency of the combined words appears greater than the first threshold, it means that the number of times the combined words appear in the statement comments of the online data comments reaches a certain quantity, then the combined words have a high probability of being a new word term. Therefore, the combined words can be retained and named as new word terms.

[0114] In an embodiment of the present invention, after step S4, after obtaining a new word, the structure of the new word may not conform to the usual language usage habits and is not a noun phrase. Therefore, after obtaining the new word, retain the new words that conform to the patterns of NN, NN NN, JJ NN, NN NN NN, JJ NN NN, and JJ JJ NN, where NN represents a single noun and JJ represents a single adjective. New words that conform to the above structures can all be regarded as a complete noun phrase.

[0115] In an embodiment of the present invention, Figure 3 is a flowchart for determining a candidate attribute feature word library of a tourism consumer market segmentation analysis method according to an embodiment of the present invention. Between steps S5 and S7, after obtaining the noun phrase set, there may be many duplicate noun words and new words in the noun phrase set, and the meanings of some noun words and new words may overlap. Therefore, it is necessary to delete the duplicate and overlapping words. The steps of deleting duplicate words and overlapping words may include:

[0116] In step S26, identify the duplicate noun words and new words in the noun phrase set.

[0117] In step S27, delete the duplicate noun words and new words in the noun phrase set.

[0118] In step S28, compare the noun words and new words in the noun phrase set to find new words that overlap in meaning with the noun words.

[0119] In step S29, delete the new words that overlap in meaning with the noun words.

[0120] In step S30, combine the noun words and the new words after deletion into a candidate attribute feature word library.

[0121] In an embodiment of the present invention, there will be a large number of noun words and new words. Some of the noun words and new words will be repetitive. Therefore, it is necessary to identify the repetitive noun words and new words in the noun phrase set, and then delete the repetitive noun words and new words to streamline the noun phrase set. Some of the noun words and new words in the streamlined noun phrase set may have overlapping meanings. Therefore, the new words and noun words that appear in the streamlined noun phrase set can be compared, and the new words that overlap in meaning with the noun words can be deleted, so that there will be no attribute redundancy among the noun words and new words that appear in the noun phrase set, and it will be more streamlined. The above-mentioned noun words and the new words after deletion can be combined into a candidate attribute feature word library, which is more streamlined than the noun phrase set and there will be no redundancy among the noun words and new words in it.

[0122] In an embodiment of the present invention, Figure 4 is a flowchart for screening a candidate attribute feature word library of a method for segmenting and analyzing the tourism consumer market based on online reviews according to an embodiment of the present invention. When screening the noun words and new words in the candidate attribute feature word library, after step S30, it may include:

[0123] In step S31, obtain the frequencies of the noun words and the new words after deletion that appear in the online data reviews.

[0124] In step S32, determine whether the frequencies of the noun words and the new words after deletion that appear in the online data reviews are greater than a second threshold.

[0125] In step S33, in the case where it is determined that the frequencies of the noun words and the new words after deletion that appear in the online data reviews are greater than the second threshold, the noun words and the new words after deletion with frequencies greater than the second threshold can be combined into a candidate attribute feature word library.

[0126] In an embodiment of the present invention, when streamlining the noun phrase set into a candidate attribute feature word library, the frequencies of the noun words and the new words after deletion that appear in the online data reviews can be obtained, that is, the number of times the noun words and the new words after deletion appear. If the number of times the noun words and new words appear is more, it means that the noun words and the new words after deletion are more likely to be words that completely represent attributes. In the case where it is determined that the frequencies of the noun words and the new words after deletion that appear in the online data reviews are greater than the second threshold, it can be determined that the noun words and the new words after deletion are words that can completely represent attributes, and then the noun words and the new words after deletion are combined into a candidate attribute feature word library.

[0127] In an embodiment of the present invention,Figure 5 It is a flowchart of a product attribute feature word library for a method of analyzing and segmenting the tourism consumer market based on online reviews according to an embodiment of the present invention. In step S7, after incorporating the noun words and new words in the noun phrase set as attribute words under the corresponding themes to obtain the product attribute feature word library, it may include:

[0128] In step S34, determine some of the new words and noun words in the noun phrase set under the corresponding themes and name them as seed attribute words.

[0129] In step S35, compare the similarity between the remaining new words and noun words in the noun phrase set and the seed attribute words.

[0130] In step S36, select the remaining new words and noun words with a similarity greater than the third threshold and incorporate them into the corresponding themes to obtain the product attribute feature word library.

[0131] After determining the theme of the online data review through the LAD mode, some new words and noun words can be selected and determined under the corresponding themes according to their attributes, and they can be named as seed attribute words. After obtaining the seed attribute words, the similarity between the remaining new words and noun words in the noun phrase set and the seed attribute words can be compared. When the similarity is greater than the third threshold, it can indicate that the remaining new words and noun words are in the same theme as the seed attribute words. Therefore, the remaining new words and noun words with a similarity greater than the third threshold can be selected and incorporated into the corresponding themes, thereby obtaining the product attribute feature word library.

[0132] In an embodiment of the present invention, Figure 6 It is a flowchart of determining seed sentiment words for a method of analyzing and segmenting the tourism consumer market based on online reviews according to an embodiment of the present invention. In step S8, incorporating some sentiment words in the sentiment dictionary as seed sentiment words may include:

[0133] In step S37, select some sentiment words in the sentiment dictionary as candidate seed sentiment words.

[0134] In step S38, determine the co-occurrence degree of the candidate seed sentiment words and the seed attribute words under each theme according to formula (1);

[0135] (1)

[0136] Wherein, represents the number of times the seed attribute word appears alone in the sentences of the online data review, represents the number of times the candidate seed sentiment word appears alone in the sentences of the online data review, Indicates the number of times a candidate seed sentiment word and a seed attribute word co - occur in a sentence in the online data comments. Indicates the co - occurrence degree of the candidate seed sentiment word and the seed attribute word.

[0137] In step S39, it is judged whether the co - occurrence degree is greater than the fourth threshold.

[0138] In step S40, when the co - occurrence degree is greater than the fourth threshold, it is determined that the candidate seed sentiment word is a seed sentiment word under the same theme as the seed attribute word.

[0139] Taking the sentiment dictionary as the candidate seed sentiment word, calculate the co - occurrence degree of the candidate seed sentiment word and the seed attribute word, that is, the number of times the candidate seed sentiment word and the seed attribute word appear in a sentence in the online data comments. Calculate its co - occurrence degree through formula (1). When the co - occurrence degree is greater than the fourth threshold, it can be determined that the candidate seed sentiment word is a seed sentiment word under the same theme as the seed attribute word.

[0140] In an embodiment of the present invention, Figure 7 is an adjective word screening flowchart of a tourism consumer market segmentation analysis method based on online reviews according to an embodiment of the present invention. In step S10, the matching of the seed sentiment word and the adjective word may include:

[0141] In step S41, match the adjective word with the seed sentiment word.

[0142] In step S42, judge whether the matching degree of the adjective word and the seed sentiment word is greater than the fifth threshold.

[0143] In step S43, when it is judged that the matching degree of the adjective word and the seed sentiment word is greater than the fifth threshold, it is determined that the matching degree of the adjective word meets the standard.

[0144] When expanding the sentiment word library, it is necessary to select adjective words that meet the standard. The adjective words that meet the standard need to be matched with the seed sentiment word. When the matching degree is greater than the fifth threshold, it can be explained that the attributes shown by the adjective word and the seed sentiment word are similar, so the adjective words with a matching degree greater than the fifth threshold can be incorporated into the sentiment word library.

[0145] In an embodiment of the present invention, Figure 8 is a sentiment word library deletion flowchart of a tourism consumer market segmentation analysis method based on online reviews according to an embodiment of the present invention. Between steps S11 and S12, the sentiment word library can be further expanded, which may include:

[0146] In step S44, add the general dictionary to the sentiment word library.

[0147] In step S45, duplicate adjective words in the sentiment word library and sentiment words in the sentiment dictionary are deleted.

[0148] After expanding the sentiment word library, it is necessary to remove duplicates from the expanded sentiment word library, and delete duplicate adjective words in the sentiment word library and sentiment words in the sentiment dictionary to obtain the final sentiment word library.

[0149] Through the above technical solution, a method for segmenting and analyzing the tourism consumer market based on online reviews provided by the present invention obtains online data reviews, then processes the online data reviews, obtains noun words from the online data reviews, and then obtains new words from the online data reviews. The noun words and new words are combined into a noun phrase set, and the noun phrase set is processed to obtain a product attribute feature word library. The sentiment words in the sentiment dictionary and the adjectives obtained from the online data reviews are combined into a sentiment word library. The evaluation values of the attribute feature words that conform to the product attribute feature word library, the sentiment words and adjective words, and intensity modifiers that conform to the sentiment word library are calculated respectively to convert each statement review in the online data reviews into a multi-dimensional score vector. The SOM algorithm and the slime mold algorithm are executed on the multi-dimensional vector and according to the nearest neighbor rule to cluster the multi-dimensional score vector into multiple categories, and the quantity under each cluster is obtained. Analysts can obtain a more accurate segmentation result of the tourism consumption market based on the clustering result and quantity of the online data reviews for subsequent analysis.

[0150] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0151] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for analyzing the market segmentation of tourism consumers based on online reviews, characterized in that, The method includes: Obtain online data comments; Perform word segmentation on the online data comments to divide each statement comment in the online data comments into multiple words; Obtain the noun words in the words; Obtain the new words in the online data comments; Combine the noun words and the new words into a noun phrase set; Analyze the online data comments through the LDA model to obtain different topics; Incorporate the noun words and the new words in the noun phrase set as attribute feature words under the corresponding topics to obtain a product attribute feature word library; Select some sentiment words in the sentiment dictionary as seed sentiment words; Obtain the adjective words in the online data comments; Screen the adjective words that meet the standard in terms of the matching degree with the seed sentiment words; Combine all the sentiment words in the sentiment dictionary and the qualified adjective words into a sentiment word library; Obtain the attribute feature words in the online data comments that conform to the product attribute feature word library; Obtain the sentiment words and the adjective words that conform to the sentiment word library near the attribute feature words in the online data comments; Obtain the intensity modifiers near the sentiment words and adjective words in the online data comments; Calculate the evaluation values of the attribute feature words that conform to the product attribute feature word library, the sentiment words and the adjective words that conform to the sentiment word library, and the intensity modifiers respectively to convert each statement comment in the online data comments into a multi-dimensional score vector; Execute the SOM algorithm on the multi-dimensional score vector and analyze to obtain multiple clustering numbers and the neuron weights under each cluster; Use the neuron weights as the cluster centers; Use the cluster centers as the initial slime mold positions, and execute the slime mold algorithm on the multi-dimensional score vector to obtain the updated slime mold positions; After determining the updated slime mold positions, classify the multi-dimensional score vectors close to the updated slime mold positions into one category according to the nearest neighbor rule; Calculate the number of the multi-dimensional score vectors classified into one category; Determine the result of the tourism consumer market segmentation analysis according to the clustering of the multi-dimensional score vectors and the quantity under each cluster.

2. The analysis method according to claim 1, characterized in that After obtaining the online data comments, clean and trim the online data comments to obtain a preprocessed text.

3. The analysis method according to claim 1, characterized in that, Obtaining the new words in the online data comments includes: Obtain the combined words that appear in the online data comments; Calculate the frequency of the combined words that appear in the online data comments; Judge whether the frequency of the combined words that appear in the online data comments is greater than a first threshold; In the case where it is judged that the frequency of the combined words that appear in the online data comments is greater than the first threshold, retain the combined words and name them new words.

4. The analysis method according to claim 1, wherein The method includes: After obtaining the new words, retain the new words that conform to the NN, NN NN, JJ NN, NN NN NN, JJ NN NN, JJ JJ NN patterns, where NN represents a single noun and JJ represents a single adjective.

5. The analysis method according to claim 1, wherein Combining the noun words and the new words into a noun phrase set includes: Identify the repeated noun words and new word terms in the set of noun phrases; Delete the repeated noun words and new word terms in the set of noun phrases; Compare the new word terms in the set of noun phrases with the noun words, and find the new word terms that coincide in meaning with the noun words; Delete the new word terms that coincide in meaning with the noun words; Combine the noun words and the remaining new word terms after deletion into a candidate attribute feature word bank.

6. The analysis method according to claim 5, characterized in that Combining the noun words and the remaining new word terms after deletion into a candidate attribute feature word bank includes: Obtain the frequencies of the noun words and the remaining new word terms after deletion that appear in the online data comments; Determine whether the frequencies of the noun words and the remaining new word terms after deletion that appear in the online data comments are greater than a second threshold; In the case where it is determined that the frequencies of the noun words and the remaining new word terms after deletion that appear in the online data comments are greater than the second threshold, combine the noun words and the remaining new word terms with frequencies greater than the second threshold into a candidate attribute feature word bank.

7. The analysis method according to claim 1, wherein Incorporate the noun words and the new word terms in the set of noun phrases as attribute feature words under the corresponding topics to obtain a product attribute feature word bank, including: Determine that some of the new word terms and noun words in the set of noun phrases are under the corresponding topics, and name them seed attribute words; Compare the similarity between the remaining new word terms and noun words in the set of noun phrases and the seed attribute words; Select the remaining new word terms and noun words with a similarity greater than a third threshold and incorporate them into the corresponding topics to obtain a product attribute feature word bank.

8. The analysis method according to claim 7, wherein Selecting some emotion words in the emotion dictionary as seed emotion words includes: Select some emotion words in the emotion dictionary as candidate seed emotion words; Determine the co-occurrence degree of the candidate seed emotion words and the seed attribute words under each topic according to formula (1); (1) Among them, represents the number of times the seed attribute word appears alone in the sentences of the online data comments, represents the number of times the candidate seed sentiment word appears alone in the sentences of the online data comments, represents the number of times the candidate seed sentiment word and the seed attribute word appear together in a sentence of the online data comments, represents the co-occurrence degree of the candidate seed sentiment word and the seed attribute word; Determine whether the co-occurrence degree is greater than a fourth threshold; When the co-occurrence degree is greater than the fourth threshold, determine that the candidate seed emotion word is the seed emotion word under the same topic as the seed attribute word.

9. The analysis method according to claim 1, wherein Screen the adjective words whose matching degree with the seed emotion words meets the standard, including: Match the adjective words with the seed emotion words; Determine whether the matching degree of the adjective words and the seed emotion words is greater than a fifth threshold; In the case where it is determined that the matching degree of the adjective words and the seed emotion words is greater than the fifth threshold, determine that the matching degree of the adjective words meets the standard.

10. The analysis method according to claim 1, characterized in that, The method includes: Add the general dictionary to the emotion word bank; Delete the repeated adjective words in the emotion word bank and the emotion words in the emotion dictionary.

Citation Information

Patent Citations

  • Expert comment induction algorithm based on sentiment classification and SOM clustering

    CN106156184A

  • Method for constructing prediction model based on colibacillus algorithm

    CN110705640A