Tourism comment sorting method based on fuzzy logic

Through the tourism review sorting method based on fuzzy logic, the problem of difficult to efficiently sort massive tourism reviews in the existing technology is solved, and efficient and interpretable comment information sorting is achieved, which significantly improves the accuracy and efficiency of sorting.

CN120106235AInactive Publication Date: 2025-06-06HUNAN XINCHENG YUNZHI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417754.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and interpretably discover high-value tourism review information in massive and complex tourism review information, and the ranking of tourism reviews is uncertain and vague, and the complexity is high.

Method used

A tourism comment sorting method based on fuzzy logic is used to crawl the comment information of the travel website, classify the comment topics and decompose the index data, and use the fuzzy operator to obtain the sorting scores of the comment topics, and finally output the sorting result.

Benefits of technology

It realizes efficient sorting among massive tourism reviews, showing more valuable data, the method is interpretable, has high accuracy and strong applicability, effectively improving the efficiency and accuracy of tourism review sorting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106235A_ABST
    Figure CN120106235A_ABST
Patent Text Reader

Abstract

The invention relates to the field of tourism data analysis, in particular to a tourism comment sorting method based on fuzzy logic, which comprises the following steps: crawling tourism comment information publicly published in a tourism website, performing index data decomposition according to comment subject classification data of the tourism comment information, and performing fuzzy description on each decomposed index data to obtain a fuzzy description result; and obtaining the sorting score of each comment theme by using a fuzzy operator, and outputting a final sorting result based on the score. According to the method, the tourism comment data which is obtained by a crawler and is preprocessed is used as input, sorting of the data is achieved through a sorting method that multi-index calculation, fuzzy operator sorting and tourism comment theme sorting results are used as output, and the tourism comment sorting efficiency and accuracy can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of tourism data analysis, and in particular to a tourism review ranking method based on fuzzy logic. Background Art

[0002] With the development of Internet technology, the Internet has become the core position for the dissemination of tourism reviews. Faced with this massive amount of tourism review information with complex opinions, it is often difficult for tourism review monitors to quickly capture accurate and valuable content. In order to mine high-value tourism review information data from complex data, the current main research points are the review information traceability and content analysis dimensions. There are few studies on how to explain and efficiently discover available tourism review information. The ranking of tourism reviews itself is uncertain and fuzzy, and the complexity of tourism review ranking is greatly increased. Therefore, in order to solve this problem, it is urgent to propose a tourism review ranking method based on fuzzy logic. Summary of the invention

[0003] The present invention aims at the deficiencies in the prior art and provides the following technical solutions:

[0004] A tourism review ranking method based on fuzzy logic, including:

[0005] S10: crawling the tourism review information publicly released on tourism websites (Ctrip, Fliggy, Weibo), decomposing the index data according to the review topic classification data of the tourism review information, and fuzzily describing each index data after decomposition.

[0006] S20: Use the fuzzy operator to obtain the ranking score of each comment topic, and output the final ranking result based on this score.

[0007] As an improvement of the above technical solution, the crawling of publicly released travel review information on travel websites includes the following steps:

[0008] S111: Using Selenium to crawl publicly published travel review information on travel websites;

[0009] S112: Clean the crawled data and obtain tourism-related data information;

[0010] S113: Store the acquired relevant data information into a database to form structured data for subsequent processing.

[0011] As an improvement of the above technical solution, the method for obtaining the comment topic classification data relies on the following steps:

[0012] S121: Using the BERT model to classify the comment topics of the structured data stored in step S113, and aggregating the comment data with the same topic;

[0013] S122: Based on the review topic classification data, the index content in the travel review information with the same topic is stored in the same elasticsearch index table, and the report content is stored in the HBase table.

[0014] S123: The index content and the report content are associated through the RowKey field in the elasticsearch index table, and the id in the structured data is used to complete the construction of the data retrieval set, where the id is an auto-increment primary key.

[0015] As an improvement of the above technical solution, the index data decomposed in step S10 at least include: the information dissemination scope of each comment topic, public attention, negative impact hazard, authority of information source, and event development trend.

[0016] As an improvement of the above technical solution, the fuzzy description of the information dissemination scope of each comment topic depends on the following steps:

[0017] S131: Define the scope of information dissemination as at least a vague description from the three dimensions of traditional media coverage, new media diffusion, and industry media attention; define public attention as at least a vague description from the three dimensions of direct search volume, interactive participation, and topic persistence; define the degree of negative impact as at least a vague description from the three dimensions of reduced tourist proportion, degree of reputation damage, and degree of impact on tourism market confidence; define the authority of information sources as at least a vague description from the two dimensions of media platform level and information content credibility; define the development trend of events as at least a vague description from the two dimensions of change rate of reporting speed and tendency of public attitude evolution.

[0018] S132: According to the comment topic classification data and the dimensions determined in step S131, obtain the traditional media coverage, new media diffusion, industry media attention, direct search volume, interactive participation volume, topic continuity, negative impact, media platform level, information content credibility, reporting speed change rate, and public attitude evolution tendency of a certain comment topic.

[0019] As an improvement of the above technical solution, the acquisition of the traditional media coverage in step S132 depends on the following formula:

[0020] Among them, i represents the media level index, j represents the layout type index, MRC ij represents the number of the jth type of page in the i-th level media reports, w jrepresents the weight value of the j-th layout type, W i represents the weight of the i-th level media, S i represents the scale score of the media coverage at level i, TMC max This is the highest coverage score in the history of traditional media.

[0021] The acquisition of the new media diffusion degree in step S132 depends on the following formula:

[0022] Among them, c represents the media type, MRC c represents the number of reports in the cth media, W c Indicates the preset weight of the cth media, NMC max This is the maximum coverage score in the history of new media.

[0023] The acquisition of the industry media attention in step S132 depends on the following formula:

[0024] Among them, d represents the reporting depth type, DSQ d represents the number of reports in the dth report depth, W d Indicates the preset weight corresponding to the dth reporting depth, IMCA max The highest industry media attention score in history.

[0025] The acquisition of the interactive participation amount in step S132 depends on the following formula:

[0026] Among them, c represents the media type, CC c represents the total number of comments corresponding to the comment topic under the cth media type, CW is the preset comment weight, TUC c represents the total number of likes corresponding to the comment topic under the cth media type, TUW is the preset like weight, SC c represents the total number of shares corresponding to the comment topic under the cth media type, SW is the preset sharing weight, SOW c Indicates the preset weight of the cth media type, IEV max Indicates the maximum historical interactive participation.

[0027] The acquisition of the negative impact degree in step S132 depends on the following formula:

[0028] Among them, PTD represents the score of the reduction ratio of tourists, PTDW represents the weight of the reduction ratio of tourists in NIHD, t represents the number of reports under the comment topic, and DRD trepresents the reputation damage score of the t-th report content, DRDW represents the weight of the reputation damage in NIHD, TMC represents the tourism market information attack score, TMCW represents the weight of the tourism market information attack in NIHD, NIHD max The largest negative impact in history.

[0029] The acquisition of the media platform level in step S132 depends on the following formula:

[0030] Among them, i represents the media level, PRTN i represents the number of reports corresponding to the i-th media level, RPTNW i Indicates the preset weight corresponding to the i-th media level, MPH max Indicates the highest media platform level score in history.

[0031] The acquisition of the credibility of the information content in step S132 depends on the following formula:

[0032] Among them, f represents the fth comment report, DA f represents the data accuracy score of the report in article f, DAW represents the weight of data accuracy in the credibility of information content, LR f represents the logical rigor score of the report in article f, LRW represents the weight of logical rigor in the credibility of information content, ICC max Indicates the historical maximum information content credibility score.

[0033] The acquisition of the reporting speed change rate in step S132 depends on the following formula:

[0034] Among them, tanh is the hyperbolic tangent function, Sum 2 is the number reported in the next time period, Sum 1 The number reported for the previous time period.

[0035] The acquisition of the public attitude evolution tendency in step S132 depends on the following formula:

[0036] Among them, PEC 2 is the proportion of emotional change in the next time period, PEC 1 It is the proportion of emotion change in the previous time period.

[0037] As an improvement of the above technical solution, the acquisition of the direct search volume in step S132 depends on the following steps:

[0038] S1321: Obtain the daily search volume under the comment topic category through Baidu Index, and determine the final DSV value based on the daily search volume.

[0039] The acquisition of the topic duration in step S132 depends on the following steps:

[0040] 1322: Through the elasticsearch index table, filter out the duration of each comment topic category according to the reporting time, and determine the topic duration based on the duration.

[0041] The acquisition of the reduction ratio of tourists, the degree of reputation damage, and the degree of impact on tourism market confidence in step S132 depends on the following steps:

[0042] S1323: Obtain the reduction ratio of tourists in the same period through year-on-year calculation, and classify according to the reduction ratio to obtain the reduction ratio of tourists.

[0043] S1324: Use the XLNet sentiment analysis model to perform sentiment analysis on the report content in step S123, and classify it according to the analysis results to obtain the reputation damage degree.

[0044] S1325: Analyze the travel booking data in the review topic classification data in step S132 through the random forest classification model, and perform classification processing based on the analysis results to obtain the degree of market confidence impact.

[0045] As an improvement of the above technical solution, step S20 includes the following steps:

[0046] S21: Convert the fuzzy description of each indicator into a decision matrix, and convert the decision matrix into a confidence interval to obtain a set of confidence intervals.

[0047] S22: Obtaining intermediate variables based on confidence interval sets And according to the intermediate variable Get the variables in the fuzzy operator

[0048] S23: The comprehensive score of each tourism review topic is calculated using the hesitant fuzzy McLaughlin symmetric average operator under the DST theoretical framework.

[0049] S24: Obtain a final score based on the comprehensive score of each tourism review topic, and sort according to the final score.

[0050] As an improvement of the above technical solution, the conversion method of the decision matrix into the confidence interval depends on the following formula:

[0051] Among them, Bel ij ,Pl ij is a set of confidence intervals, for b ij The envelope of ij ={e 1 ,e 2 ,…,e n}.

[0052] The intermediate variable The way to obtain depends on the following formula:

[0053] Where, i = 1, 2…, m; j, τ = 1, 2,…, η; j ≠ τ, m is the number of tourism review topics to be sorted, η is the number of calculation indicators, For calculation The intermediate variable.

[0054] The variable The way to obtain depends on the following formula:

[0055] Among them, i = 1, 2…, m; j, τ = 1, 2,…, η; j ≠ τ, m is the number of tourism review topics to be sorted, and η is the number of calculation indicators.

[0056] The calculation method of the hesitant fuzzy Maclaurin symmetric average operator under the DST theoretical framework depends on the following formula:

[0057] Among them, Bel ηj ,Pl ηj is a confidence interval set, t is an indicator item, k is a dynamically adjusted value, and its value range is [1, t]. is the number of combinations, where The calculation formula is as follows:

[0058] Among them, ω η is the weight corresponding to different indicators, satisfying Can be adjusted dynamically.

[0059] The final score is obtained by the following formula:

[0060] AF DST (theme) = pl-Bel

[0061] Among them, Bel and pl are the calculation results of the hesitant fuzzy Maclaurin symmetric average operator under the DST theoretical framework.

[0062] As an improvement of the above technical solution, the sorting method of step S24 relies on the following steps:

[0063] If SF DST (theme 1 ) <SF DST (theme 2 ), then theme 1 <theme 2 ;

[0064] In SF DST (theme 1 )=SF DST (theme 2 ) If AF DAT (theme 1 )>AF DST (theme 2 ), then theme 1 <theme 2 If AF DST (theme 1 )=AF DST (theme 1 ), then theme 1 =theme 2 .

[0065] Beneficial effects of the present invention:

[0066] The tourist review data obtained by crawlers and preprocessed is used as input, and the sorting method uses multi-indicator calculation, fuzzy operator sorting, and tourist review topic sorting results as output to achieve data sorting, which can present more valuable data to users. In addition, the indicator combination calculation method used in this method is clear, the sorting process is explainable, the method is simple and effective, with high accuracy and strong applicability, which can effectively improve the efficiency and accuracy of tourist review sorting. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0068] The following describes the embodiments of the present invention through specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention.

[0069] The current main research points are the traceability of review information and content analysis dimensions. There are few studies on how to explain and efficiently discover available tourism review information. The ranking of tourism reviews itself is uncertain and fuzzy, and the complexity of ranking tourism reviews is greatly increased. Therefore, to address this problem, a tourism review ranking method based on fuzzy logic is proposed, including:

[0070] S10: crawling the tourism review information publicly released on the tourism website, and decomposing the indicator data according to the comment topic classification data of the tourism review information, and fuzzy describing each indicator data after decomposition;

[0071] The crawling of publicly released travel review information on the website includes the following steps:

[0072] S111: Use Selenium to crawl publicly published travel review information on travel websites.

[0073] In the process of crawling data, traditional media will be crawled in a hierarchical manner, such as national data, provincial data, and municipal data, while new media will be crawled in a channel-based manner, such as Weibo, WeChat public accounts, Douyin, etc. In addition to these, it can usually also include tourism industry media websites, tourism industry forums, tourism-related websites, etc., and obtain tourism-related comment data through these channels.

[0074] S112: Clean the crawled data and obtain tourism-related data information.

[0075] The crawled data also needs to be cleaned and filtered to delete unimportant data and keep only the data that is useful for calculation. For example, the following types of data are usually saved:

[0076] Media level, page position, report duration / report text length, media name, playback volume, reading volume, report content, report time, number of comments, number of likes, number of shares, media type, number of tourists / day, travel reservation data / day.

[0077] S113: Store the acquired relevant data information into a database to form structured data for subsequent processing.

[0078] The method for obtaining the comment topic classification data depends on the following steps:

[0079] S121: Use the BERT model to classify the comment topics of the structured data stored in step S113, and aggregate the comment data with the same topic.

[0080] S122: Based on the review topic classification data, index contents in the travel review information with the same topic are stored in the same elasticsearch index table, and the report contents are stored in the HBase table;

[0081] S123: The index content and the report content are associated through the RowKey field in the elasticsearch index table, and the id in the structured data is used to complete the construction of the data retrieval set, where the id is an auto-increment primary key.

[0082] In addition, the index data decomposed in step S10 at least include: the information dissemination scope, public attention, negative impact hazard, information source authority, and event development trend of each comment topic. Based on this, fuzzy calculations need to be performed on these index data respectively, as shown below:

[0083] The fuzzy description of the information dissemination scope of each comment topic relies on the following steps:

[0084] S131: Define the scope of information dissemination as at least a vague description from the three dimensions of traditional media coverage, new media diffusion, and industry media attention; define public attention as at least a vague description from the three dimensions of direct search volume, interactive participation, and topic persistence; define the degree of negative impact as at least a vague description from the three dimensions of reduced tourist proportion, degree of reputation damage, and degree of impact on tourism market confidence; define the authority of information sources as at least a vague description from the two dimensions of media platform level and information content credibility; define the development trend of events as at least a vague description from the two dimensions of change rate of reporting speed and tendency of public attitude evolution.

[0085] S132: According to the comment topic classification data and the dimensions determined in step S131, obtain the traditional media coverage, new media diffusion, industry media attention, direct search volume, interactive participation volume, topic continuity, negative impact, media platform level, information content credibility, reporting speed change rate, and public attitude evolution tendency of a certain comment topic.

[0086] In order to calculate the traditional media coverage, it is necessary to comprehensively calculate the media level, the number of media reports at each level, the page position, the scale of the report, and the historical maximum possible coverage score. The media level is divided into national, provincial, and municipal levels, with weights of 0.5, 0.3, and 0.2 respectively. The page position is divided into front page headlines, important pages, and ordinary pages, with scores of 0.5, 0.3, and 0.2 respectively. The report scale is 10 minutes / 1000 words as a unit, and each unit is counted as 0.1 points. The acquisition of the traditional media coverage in step S132 depends on the following formula:

[0087] Among them, i represents the media level index, j represents the layout type index, MRC ij represents the number of the jth type of page in the i-th level media reports, w j represents the weight value of the j-th layout type, W i represents the weight of the i-th level media, S i represents the scale score of the media coverage at level i, TMC max This is the highest coverage score in the history of traditional media.

[0088] In order to calculate the new media diffusion, it is also necessary to obtain the number of Weibo reports, Douyin reports, and WeChat public account reports from the crawled data, and calculate the new media diffusion in step S132 based on the obtained data, specifically as follows:

[0089] Among them, c represents the media type, MRC c represents the number of reports in the cth media, W c Indicates the preset weight of the cth media, NMC max This is the maximum coverage score in the history of new media, among which the weights of Weibo, Tik Tok, and WeChat are 0.4, 0.4, and 0.2 respectively.

[0090] For the report contents with the media type of "tourism industry" screened out, the ERNIE model was used to divide each report into three report depths: "in-depth report", "general report" and "brief report". Finally, the industry media attention was calculated based on the report depth, as shown in the following formula:

[0091] Among them, d represents the reporting depth type, DSQ d represents the number of reports in the dth report depth, W d Indicates the preset weight corresponding to the dth reporting depth, IMCA max The highest industry media attention score in history.

[0092] The acquisition of the direct search volume in step S132 depends on the following steps:

[0093] S1321: Obtain the daily search volume under the comment topic category through Baidu Index, and determine the final DSV value based on the daily search volume.

[0094] The corresponding relationship between the specific daily search volume and DSV value is as follows:

[0095] Daily search volume range DSV value Greater than or equal to 1 million 0.9 500,000-1,000,000 0.7 10,000-100,000 0.5 Less than 100,000 0.1

[0096] Similarly, the data in the elasticsearch index table is used to analyze the number of comments, likes, shares, and sources of reports under the comment topic. Based on these data, a comprehensive calculation is performed. Specifically, the acquisition of the interactive participation amount in step S132 depends on the following formula:

[0097] Among them, c represents the media type, CC c Indicates the total number of comments corresponding to the comment topic under the cth media type, CW is the preset comment weight (the default is 0.5), TUC c represents the total number of likes corresponding to the comment topic under the cth media type, TUW is the preset like weight (the default is 0.3), SC c Indicates the total number of shares corresponding to the comment topic under the cth media type, SW is the preset sharing weight (the default is 0.2), SOW c Indicates the preset weight of the cth media type (default traditional media 0.3, Weibo 0.3, Douyin 0.2, WeChat public account 0.2), IEV max Indicates the maximum historical interactive participation.

[0098] The acquisition of the topic duration in step S132 depends on the following steps:

[0099] 1322: Through the elasticsearch index table, filter out the duration of each comment topic category according to the reporting time, and determine the topic duration based on the duration.

[0100] It is usually divided into four levels. The first level is when the duration is greater than 7 days, and the value of the topic duration is 0.7. 3-7 days is the second level, and the value is 0.5. 1-2 days is the third level, and the value is 0.3. Less than 1 day is the fourth level, and the value is 0.1.

[0101] The acquisition of the reduction ratio of tourists, the degree of reputation damage, and the degree of impact on tourism market confidence in step S132 depends on the following steps:

[0102] S1323: Obtain the reduction ratio of tourists in the same period through year-on-year calculation, and classify according to the reduction ratio to obtain the reduction ratio of tourists.

[0103] The reduction ratio can usually be divided into five levels: more than 50%, 30%-50%, 10%-30%, 5%-10%, and below 5%, with scores of 0.8, 0.5, 0.3, 0.2, and 0.1 respectively.

[0104] S1324: Use the XLNet sentiment analysis model to perform sentiment analysis on the report content in step S123, and classify it according to the analysis results to obtain the reputation damage degree.

[0105] After the degree of reputation damage is graded, the decomposition results are usually divided into four levels: severe damage, major damage, general damage, and minor damage, with scores of 0.9, 0.7, 0.5, and 0.2 respectively.

[0106] S1325: Analyze the travel booking data in the review topic classification data in step S132 through the random forest classification model, and perform classification processing based on the analysis results to obtain the degree of market confidence impact.

[0107] The degree of the impact on market confidence is usually graded according to the analysis results and is divided into four levels: severe impact, major impact, general impact and slight impact, with scores of 0.9, 0.7, 0.5 and 0.2 respectively.

[0108] Specifically, the acquisition of the negative impact degree in step S132 depends on the following formula:

[0109] Among them, PTD represents the score of the reduction ratio of tourists, PTDW represents the weight of the reduction ratio of tourists in NIHD (the default value is 0.4), t represents the number of reports under the comment topic, and DRD t represents the reputation damage score of the t-th report content, DRDW represents the weight of reputation damage in NIHD (the default is 0.3), TMC represents the tourism market information attack score, TMCW represents the weight of tourism market information attack in NIHD (the default is 0.3), NIHD max The largest negative impact in history.

[0110] The acquisition of the media platform level in step S132 depends on the following formula:

[0111] Among them, i represents the media level, PRTN i represents the number of reports corresponding to the i-th media level, RPTNW iIndicates the preset weight corresponding to the i-th media level. In this embodiment, the media levels are divided into national mainstream media platforms, provincial key media platforms, prefecture-level well-known media platforms, county-level and below ordinary media platforms, small self-media or personal accounts, and their weights are 0.9, 0.7, 0.5, 0.3, 0.1, MPH max Indicates the largest media platform level score in history;

[0112] The calculation of the credibility of the information content is obtained by comprehensive calculation of data accuracy and logical rigor, wherein the data accuracy is verified by using the Autoencoder anomaly detection model to verify the text content in step S123, which is divided into three levels: accurate, basically accurate, and partially wrong, with scores of 0.5, 0.3, and 0.2 respectively. The logical rigor uses the Transformer-XL logical analysis model to perform logical analysis on the text content in step S123, which is divided into four levels: strict, relatively strict, general, and loose, with scores of 0.4, 0.3, 0.2, and 0.1 respectively. The data accuracy score and logical rigor are weighted. Specifically, the acquisition of the credibility of the information content in step S132 depends on the following formula:

[0113] Among them, f represents the fth comment report, DA f represents the data accuracy score of the report in the fth article, DAW represents the weight of data accuracy in the credibility of information content (default is 0.6), LR f Indicates the logical rigor score of the report in article f. LRW indicates the weight of logical rigor in the credibility of information content (default is 0.4). ICC max Indicates the historical maximum information content credibility score.

[0114] The reporting speed change rate is calculated by retrieving the reporting time in the index table of step S123 and counting the reporting amount in different time periods. Specifically, the acquisition of the reporting speed change rate in step S132 depends on the following formula:

[0115] Among them, tanh is the hyperbolic tangent function, Sum 2 is the number reported in the next time period, Sum 1 the number reported for the previous time period;

[0116] The sentiment tendency analysis of the comments and reports in different time periods in step S123 is performed by using the RoBERTa sentiment analysis model, and the sentiment change ratio is counted. The acquisition of the public attitude evolution tendency in step S132 depends on the following formula:

[0117] Among them, PEC 2is the proportion of emotional change in the next time period, PEC 1 It is the proportion of emotion change in the previous time period.

[0118] S20: Use the fuzzy operator to obtain the ranking score of each comment topic, and output the final ranking result based on this score.

[0119] Specifically, step S20 includes the following steps:

[0120] S21: Convert the fuzzy description of each of the indicators into a decision matrix, and convert the decision matrix into a confidence interval to obtain a confidence interval set.

[0121] Decision matrix B = (b ij ) m×η , as shown below:

[0122]

[0123] Confidence interval (BI) for b ij The envelope of ij ={e 1 ,e 2 ,…,e n}, the conversion method of the decision matrix into the confidence interval depends on the following formula:

[0124] Among them, Bel ij ,Pl ij is a set of confidence intervals.

[0125] After being converted into confidence intervals, the decision matrix is ​​converted into a confidence interval set Bis, which is expressed as follows:

[0126]

[0127] S22: Obtaining intermediate variables based on confidence interval sets And according to the intermediate variable Get the variables in the fuzzy operator

[0128] The intermediate variable The way to obtain depends on the following formula:

[0129] Where, i = 1, 2…, m; j, τ = 1, 2,…, η; j ≠ τ, m is the number of tourism review topics to be sorted, η is the number of calculation indicators, Representation calculation The intermediate variable It can be calculated using the Jousselme distance formula under the DST framework. The specific calculation formula is as follows:

[0130] In the above formula, is the scalar product, yes The square norm of is as follows:

[0131] Here, X i , X j The address range of i and j is [1,2 n ],in, The calculation formula is as follows:

[0132] The variable The way to obtain depends on the following formula:

[0133] Among them, i = 1, 2…, m; j, τ = 1, 2,…, η; j ≠ τ, m is the number of tourism review topics to be sorted, and η is the number of calculation indicators.

[0134] S23: The comprehensive score of each tourism review topic is calculated using the hesitant fuzzy McLaughlin symmetric average operator under the DST theoretical framework.

[0135] The calculation method of the hesitant fuzzy Maclaurin symmetric average operator under the DST theoretical framework depends on the following formula:

[0136] Among them, Bel ηj ,Pl ηj is a confidence interval set, t is an indicator item, k is a dynamically adjusted value, and its value range is [1, t]. is the number of combinations, where The calculation formula is as follows:

[0137] Among them, ω η is the weight corresponding to different indicators, satisfying Can be adjusted dynamically.

[0138] S24: Obtain a final score based on the comprehensive score of each tourism review topic, and sort according to the final score.

[0139] The final score is obtained by the following formula:

[0140] AF DSTtheme)=pl-Bel

[0141] Among them, Bel and pl are the calculation results of the hesitant fuzzy Maclaurin symmetric average operator under the DST theoretical framework.

[0142] The sorting method of step S24 depends on the following steps:

[0143] If SF DST (theme 1 ) <SF DST (theme 2 ), then theme 1 <theme 2 ;

[0144] In SF DST (theme 1 )=SF DST (theme 2 ) If AF DST (theme 1 )>AF DST (theme 2 ), then theme 1 <theme 2 If AF DST (theme 1 )=AF DST (theme 1 ), then theme 1 =theme 2 .

[0145] The above embodiments are only used to illustrate the technical solutions of the present invention, but not to limit them. Anyone familiar with the technology can modify or change the above embodiments without violating the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A tourism review ranking method based on fuzzy logic, characterized in that: include: S10: crawling the tourism review information publicly released on the tourism website, and decomposing the indicator data according to the comment topic classification data of the tourism review information, and fuzzy describing each indicator data after decomposition; S20: Use the fuzzy operator to obtain the ranking score of each comment topic, and output the final ranking result based on this score.

2. A method for ranking tourist reviews based on fuzzy logic according to claim 1, characterized in that: The crawling of publicly released travel review information on the website includes the following steps: S111: Using Selenium to crawl publicly published travel review information on travel websites; S112: Clean the crawled data and obtain tourism-related data information; S113: Store the acquired relevant data information into a database to form structured data for subsequent processing.

3. A method for ranking tourist reviews based on fuzzy logic according to claim 2, characterized in that: The method for obtaining the comment topic classification data relies on the following steps: S121: Using the BERT model to classify the comment topics of the structured data stored in step S113, and aggregating the comment data with the same topic; S122: Based on the review topic classification data, index contents in the travel review information with the same topic are stored in the same elasticsearch index table, and the report contents are stored in the HBase table; S123: The index content and the report content are associated through the RowKey field in the elasticsearch index table, and the id in the structured data is used to complete the construction of the data retrieval set, where the id is an auto-increment primary key.

4. A method for ranking travel reviews based on fuzzy logic according to claim 3, characterized in that: The index data decomposed in step S10 at least include: the information dissemination scope, public attention, negative impact hazard, authority of information source, and event development trend of each comment topic.

5. A method for ranking travel reviews based on fuzzy logic according to claim 4, characterized in that: The fuzzy description of the information dissemination scope of each comment topic relies on the following steps: S131: Define the scope of information dissemination as at least three dimensions: traditional media coverage, new media diffusion, and industry media attention; define public attention as at least three dimensions: direct search volume, interactive participation, and topic persistence; define the degree of negative impact as at least three dimensions: the proportion of tourists reduced, the degree of reputation damage, and the degree of tourism market confidence; define the authority of information sources as at least two dimensions: the level of media platforms and the credibility of information content; define the development trend of events as at least two dimensions: the rate of change of reporting speed and the tendency of public attitude evolution; S132: According to the comment topic classification data and the dimensions determined in step S131, obtain the traditional media coverage, new media diffusion, industry media attention, direct search volume, interactive participation volume, topic continuity, negative impact, media platform level, information content credibility, reporting speed change rate, and public attitude evolution tendency of a certain comment topic.

6. A method for ranking travel reviews based on fuzzy logic according to claim 5, characterized in that: The acquisition of the traditional media coverage in step S132 depends on the following formula: Among them, i represents the media level index, j represents the layout type index, MRC ij represents the number of the jth type of page in the i-th level media reports, w j represents the weight value of the j-th layout type, W i represents the weight of the i-th level media, S i represents the scale score of the media coverage at level i, TMC max The maximum coverage score in the history of traditional media; The acquisition of the new media diffusion degree in step S132 depends on the following formula: Among them, c represents the media type, MRC c represents the number of reports in the cth media, W c Indicates the preset weight of the cth media, NMC max The maximum diffusion score of new media history; The acquisition of the industry media attention in step S132 depends on the following formula: Among them, d represents the reporting depth type, DSQ d represents the number of reports in the dth report depth, W d Indicates the preset weight corresponding to the dth reporting depth, IMCA max The highest industry media attention score in history; The acquisition of the interactive participation amount in step S132 depends on the following formula: Among them, c represents the media type, CC c represents the total number of comments corresponding to the comment topic under the cth media type, CW is the preset comment weight, TUC c represents the total number of likes corresponding to the comment topic under the cth media type, TUW is the preset like weight, SC c represents the total number of shares corresponding to the comment topic under the cth media type, SW is the preset sharing weight, SOW c Indicates the preset weight of the cth media type, IEV max Indicates the maximum amount of interactive participation in history; The acquisition of the negative impact degree in step S132 depends on the following formula: Among them, PTD represents the score of the reduction ratio of tourists, PTDW represents the weight of the reduction ratio of tourists in NIHD, t represents the tth report content under the comment topic, and DRD t represents the reputation damage score of the t-th report content, DRDW represents the weight of the reputation damage in NIHD, TMC represents the tourism market information attack score, TMCW represents the weight of the tourism market information attack in NIHD, NIHD max The largest negative impact in history; The acquisition of the media platform level in step S132 depends on the following formula: Among them, i represents the media level, RPTN i represents the number of reports corresponding to the i-th media level, RPTNW i Indicates the preset weight corresponding to the i-th media level, MPH max Indicates the largest media platform level score in history; The acquisition of the credibility of the information content in step S132 depends on the following formula: Among them, f represents the fth comment report, DA f represents the data accuracy score of the report in article f, DAW represents the weight of data accuracy in the credibility of information content, LR f represents the logical rigor score of the report in article f, LRW represents the weight of logical rigor in the credibility of information content, ICC max Indicates the historical maximum information content credibility score; The acquisition of the reporting speed change rate in step S132 depends on the following formula: Among them, tanh is the hyperbolic tangent function, Sum2 is the number reported in the latter time period, and Sum1 is the number reported in the former time period; The acquisition of the public attitude evolution tendency in step S132 depends on the following formula: Among them, PEC2 is the proportion of emotional change in the latter time period, and PEC1 is the proportion of emotional change in the previous time period.

7. The method for ranking travel reviews based on fuzzy logic according to claim 5, characterized in that: The acquisition of the direct search volume in step S132 depends on the following steps: S1321: Obtain the daily search volume under the comment topic category through Baidu Index, and determine the final DSV value based on the daily search volume; The acquisition of the topic duration in step S132 depends on the following steps: 1322: Through the elasticsearch index table, filter out the duration of each comment topic classification according to the reporting time, and determine the topic duration according to the duration; The acquisition of the reduction ratio of tourists, the degree of reputation damage, and the degree of impact on tourism market confidence in step S132 depends on the following steps: S1323: Obtain the reduction ratio of tourists in the same period by year-on-year calculation, and classify according to the reduction ratio to obtain the reduction ratio of tourists; S1324: using the XLNet sentiment analysis model to perform sentiment analysis on the report content in step S123, and grading the content according to the analysis results to obtain the reputation damage degree; S1325: Analyze the travel booking data in the review topic classification data in step S132 through the random forest classification model, and perform classification processing based on the analysis results to obtain the degree of market confidence impact.

8. The method for ranking travel reviews based on fuzzy logic according to claim 5, characterized in that: The step S20 comprises the following steps: S21: Convert the fuzzy description of each indicator into a decision matrix, and convert the decision matrix into a confidence interval to obtain a set of confidence intervals; S22: Obtaining intermediate variables based on confidence interval sets And according to the intermediate variable Get the variables in the fuzzy operator S23: Calculate the comprehensive score of each tourism review topic using the hesitant fuzzy MacLaughlin symmetric average operator under the DST theoretical framework; S24: Obtain a final score based on the comprehensive score of each tourism review topic, and sort according to the final score.

9. A method for ranking travel reviews based on fuzzy logic according to claim 8, characterized in that: The conversion method of the decision matrix into the confidence interval depends on the following formula: Among them, Bel ij ,Pl ij is a set of confidence intervals, for b ij The envelope of ij ={e1,e2,…,e n }; The intermediate variable The way to obtain depends on the following formula: Where, i = 1, 2…, m; j, τ = 1, 2,…, η; j ≠ τ, m is the number of tourism review topics to be sorted, η is the number of calculation indicators, For calculation The intermediate variable of The variable The way to obtain depends on the following formula: Where, i = 1, 2…, m; j, τ = 1, 2,…, η; j ≠ τ, m is the number of tourism review topics to be sorted, and η is the number of calculation indicators; The calculation method of the hesitant fuzzy Maclaurin symmetric average operator under the DST theoretical framework depends on the following formula: Among them, Bel ηj ,Pl ηj is a confidence interval set, t is an indicator item, k is a dynamically adjusted value, and its value range is [1, t]. is the number of combinations, where The calculation formula is as follows: Among them, ω η is the weight corresponding to different indicators, satisfying Can be adjusted dynamically; The final score is obtained by the following formula: AF DST (theme)=pl-Bel Among them, Bel and pl are the calculation results of the hesitant fuzzy Maclaurin symmetric average operator under the DST theoretical framework.

10. The method for ranking travel reviews based on fuzzy logic according to claim 9, characterized in that: The sorting method of step S24 depends on the following steps: If SF DST (theme1) <SF DST (theme2), then theme1 <theme2; In SF DST (theme1) = SF DST In the case of AF DST (theme1) > AF DST (theme2), then theme1 < theme2. If AF DST (theme1) = AF DST (theme1), then theme1 = theme2.