A method for collecting and analyzing industry information in real time

By collecting and processing cultural and tourism information through the internet, integrating it into seven sections, extracting keywords for scenic spots, marking the time and location of comments, and assigning ratings and stars, the system solves the problem of users switching between multiple platforms to obtain information and provides an efficient and clear display of scenic spot information.

CN117216191BActive Publication Date: 2026-02-06CHONGQING TOURISM CLOUD INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311234384.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-24
Publication Date
2026-02-06
Estimated Expiration
2043-09-24

AI Technical Summary

Technical Problem

Users have to switch between different platforms to obtain cultural and tourism information. The information is categorized inconsistencies and the descriptions are incomplete. Tourist reviews are inconsistent, making it difficult to judge the value of attractions.

Method used

We collect cultural and tourism information from the internet, preprocess and convert it into a format, divide it into seven sections, extract keywords for scenic spots, mark the time and location of comments, sort the comments, rate and star them, and display keywords and rating results.

Benefits of technology

It achieves efficient integration and filtering of attraction information, providing high-quality information. The comment section is clearly sorted, allowing visitors to quickly obtain the latest information and save search time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117216191B_ABST
    Figure CN117216191B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of tourism information collection, and discloses a method for collecting, analyzing and summarizing industry information in real time, which comprises the following steps: firstly, obtaining tourism industry information and preprocessing; secondly, cutting and integrating the preprocessed scenic spot information to obtain complete scenic spot information; thirdly, extracting multiple scenic spot keywords from the complete scenic spot information through NLP technology; fourthly, marking the scenic spot keywords and collecting and recording for processing; fifthly, binding a comment mark to a browser to obtain the comment time and location of the browser; sixthly, sorting and classifying the comment area according to the comment mark; seventhly, sorting the comment records according to the comment time; eighthly, scoring and marking the stars of the complete information according to the comment area and the number of browses, and displaying the scenic spot keywords and the marking result on the cover; the application solves the problems that the information classification rules of various platforms in the prior art are uneven, the tourism strategies are one-sided, and it is difficult for users to obtain accurate tourism information, and is suitable for accurately collecting tourism information and understanding the real-time situation of scenic spots for tourists.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of tourism information collection, and particularly relates to a method for collecting and analyzing and summarizing industry information in real time. BACKGROUND

[0002] With the increasing improvement of people's living standards, going out for tourism has become the main leisure way of people after work, and with the development of tourism industry and technology, the problems that follow also constantly affect people. When people travel, they usually choose to first check and understand the relevant tourism information on the Internet due to unfamiliarity with the tourism location.

[0003] At present, the information of the cultural and tourism industry is distributed on various platforms, public numbers or websites, and if a user wants to obtain the cultural and tourism information, the user needs to switch back and forth among various platforms, and the information classification rules of various platforms are uneven, which greatly increases the difficulty of the user in obtaining the information of the scenic spots, and because there is too much information, the description of many information about the scenic spots is not necessarily perfect, and the comments of tourists on the scenic spots are always not the same, and the proportion of good comments and bad comments on each platform is different, so that the tourists who are ready to travel are difficult to judge the tourism value of the scenic spots. SUMMARY

[0004] The present application aims to provide a method for collecting and analyzing and summarizing industry information in real time, so as to solve the problems that the user cannot obtain perfect cultural and tourism information and cannot accurately find the information of the scenic area needed by the user, and the comments of tourists are relatively one-sided.

[0005] In order to achieve the above-mentioned purpose, the present application provides the following method:

[0006] The method for collecting and analyzing and summarizing industry information in real time provided by the present application is:

[0007] S1: obtaining the information of the cultural and tourism industry on the Internet and pre-processing to obtain the information of the scenic spots;

[0008] S2: cutting and integrating the information of the scenic spots after pre-processing to obtain seven pieces of complete information of the scenic spots for one scenic spot;

[0009] S3: extracting a plurality of keywords from the complete information of the scenic spots by using the NLP technology to obtain the keywords of the scenic area;

[0010] S4: marking the keywords of the scenic area and collecting and processing the marked records;

[0011] S5: binding a comment mark to the information viewer, and the comment mark can obtain the comment time and comment location of the information viewer;

[0012] S6: sorting the comment area by comment time and comment location;

[0013] S7: sorting the comment content of the existing information comment record according to the comment time, and providing the sorted comment content to the information viewer;

[0014] S8: deriving the scenic information value according to the sorted comment content;

[0015] S9: scoring and star rating the complete information information according to the scenic information value, and displaying the scenic key words and the star rating result on the cover of the complete information information of the scenic spot.

[0016] Preferably, the step of obtaining the internet Chinese travel industry information and preprocessing includes: format conversion: converting the collected travel industry information in various inconsistent file formats into a unified file format through a file format converter; cleaning useless data: removing duplicate data, removing invalid data and short sentences, removing English, numbers and characters, and removing stop words and Chinese word segmentation; uniform template: adjusting the travel industry information to a uniform template of the scenic spot information information according to the order of scenic pictures, scenic introduction, scenic strategy and tourist comments.

[0017] Preferably, the step of segmenting and integrating the preprocessed scenic spot information includes: dividing the preprocessed travel industry information into seven sections of weather, transportation, accommodation, food, shopping, entertainment, and scenic spot guide; randomly dividing the scenic spot information in each section into three groups, randomly selecting one scenic spot information from each group, and comparing two by two through a file similarity comparison method; if the content similarity of two of the three scenic spot information is very high, filtering out the more concise one of the two scenic spot information; if the content similarity of two of the three scenic spot information is relatively high, integrating the two scenic spot information through a file integration method to form a new scenic spot information; if the similarity of the three scenic spot information is not high, introducing the next group of three scenic spot information for similarity comparison, screening out the three scenic spot information with the lowest similarity, and if the similarity still does not meet the screening standard, introducing the next group of three scenic spot information, and screening out the six scenic spot information with the lowest similarity; after each scenic spot information is compared, the remaining scenic spot information is repeatedly grouped and screened; after several rounds of screening, three scenic spot information are screened out, and then the three scenic spot information are integrated two by two to form a complete information; wherein, if the similarity is greater than 85%, it is determined that the similarity is very high, if the similarity is greater than 55%, it is determined that the similarity is relatively high, and if the similarity is less than 55%, it is determined that the similarity is not high.

[0018] Preferably, the step of the file similarity comparison method includes: first calculating the keywords of each scenic spot information, selecting the same number of keywords from each, and combining them into a set; then calculating the word frequency of each scenic spot information for the keywords in the set, and generating a word frequency vector for each; calculating the cosine similarity of the two word frequency vectors of the two scenic spot information through Euclidean distance or cosine distance, and if the threshold is exceeded, the two scenic spot information are similar scenic spot information, otherwise they are considered to be dissimilar.

[0019] Preferably, the step of the file integration method includes: first dividing the two scenic spot information into integrated information and integrated information; dividing the two scenic spot information into a number of words through an article segmentation system; putting the divided words into NLP technology in units of paragraphs, comparing the words, and deleting the integrated information paragraph or adding the integrated information paragraph to the integrated information according to the amount of repeated words.

[0020] Preferably, the step of extracting keywords from the complete information of each of the scenic spots by NLP technology to obtain scenic spot keywords comprises: cutting the information in the seven sections of the scenic spot information package into a plurality of words by NLP technology; packaging the plurality of words and screening through big data to remove common words and general words; and screening keywords with scenic spot characteristics according to the information of the characteristics and history of the scenic spot.

[0021] The keywords are compared and analyzed using a generative large language model, and the most characteristic word is selected as the scenic spot keyword of the section. Finally, a scenic spot has seven scenic spot keywords. The selected keywords can be used to view the corresponding single scenic spot information.

[0022] Preferably, the step of marking the scenic spot keywords and collecting and processing the marking records comprises: first marking the scenic spot keywords; counting once each time a tourist selects a keyword to view the corresponding single scenic spot information; and regularly counting the data and arranging the order of the single scenic spot information according to the statistical data.

[0023] Preferably, the step of classifying the comment area by comment time and comment location comprises: obtaining the weather of the day of the comment area according to the comment date and the address of the corresponding comment scenic spot obtained by the comment marking through artificial intelligence technology; determining whether the commenters are local residents or tourists according to the comment location obtained by the comment marking; marking the obtained comment date, weather of the day of the comment area, and commenters category as corresponding labels on the corresponding comments; and enabling the viewers to obtain different categories of comments according to different labels through the comment location, comment time, and weather of the day of the comment area obtained by the comment label.

[0024] Preferably, the step of sorting the comment content of the existing comment records according to the comment time is: arranging the comment content of the current existing comment records in descending order according to the update time of the current existing comment records, so that the comment content of the existing comment records with the latest update time is located at the top of the comment content part; using artificial intelligence to identify the comment, cutting the comment into a plurality of sentences with punctuation marks as nodes; and arranging the comment in descending order by comparing the number of sentences.

[0025] Preferably, the step of scoring and marking stars comprises: dividing the complete information into four parts, namely comment score, characteristic score, environment score, and heat score, and the total score is the sum of the four parts; determining the comment score by the proportion of good comments and bad comments in the comments; determining the characteristic score by the repetition rate of the scenic spot keywords; determining the environment score by the number of good comments of tourists and the number of bad comments of local residents; and determining the heat score by the number of comments marked by the viewer and the number of comments marked by the viewer.

[0026] The beneficial effects of the present application are embodied in that: the present application collects and summarizes the tourism industry information of various platforms in the Internet, and performs preprocessing processes such as format conversion, useless data cleaning, and unified template, and then performs secondary integration processing on the scenic spot information through a unique classification method, and filters and integrates a large amount of the same scenic spot information into a complete scenic spot information; then the keywords are extracted from the scenic spot information, and the keywords are used to summarize the scenic spot information, so that the viewer can quickly search and browse the scenic spot information he needs through the scenic spot keywords, and the scenic spot information data can be arranged by marking the scenic spot keywords, so that the high-quality scenic spot information can be displayed in front of the viewer first, and the comments in the comment area are designed and planned, so that the comments in the comment area are clearer and clearer, and the viewer can filter the comments according to the holiday, weather condition, local residents' evaluation of the scenic spot, etc. Filtering requirements to filter the desired comments, real-time updating and sorting of the comments, so that the viewer can obtain the latest comments, and thus obtain the latest scenic spot information of the scenic area,

[0027] The entire scenic spot information is divided into four blocks, which can make the scenic area evaluation more comprehensive, and the high-scoring scenic spot information is displayed first, which saves a lot of search time. BRIEF DESCRIPTION OF DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the specific embodiments or prior art of the present application, the drawings needed to be used in the specific embodiments or prior art description will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual proportion.

[0029] Figure 1 The flowchart of the method for real-time collection and analysis of industry information provided by the embodiment of the present application. DETAILED DESCRIPTION

[0030] The embodiments of the technical solutions of the present application will be described in detail below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and therefore only serve as examples, and cannot limit the protection scope of the present application.

[0031] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should be understood as the usual meaning understood by the skilled person in the field to which the present application belongs.

[0032] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. As will be apparent to those of ordinary skill in the art, embodiments described herein can be combined with other embodiments.

[0033] Currently, information of the tourism industry is distributed on various platforms, public numbers or websites. If a user wants to obtain tourism information, the user needs to switch back and forth among various platforms, and the information classification rules of various platforms are uneven, which greatly increases the difficulty of the user to obtain scenic spot information. Moreover, due to too much information, the description of many information about the scenic spot is not necessarily perfect, and the comments of tourists on the scenic spot are always inconsistent, and the proportion of good comments and bad comments on each platform is different, so that tourists preparing to travel are difficult to judge the tourism value of the scenic spot. In order to solve the above problems, it is necessary to develop a method for collecting and analyzing and summarizing industry information in real time.

[0034] The present application aims to provide a method for collecting and analyzing and summarizing industry information in real time, so as to solve the problems that the user obtains imperfect tourism information and it is difficult to accurately find the information of the scenic spot needed by oneself, and the comments of tourists are relatively one-sided. In order to achieve the above purpose, the present application provides the following method:

[0035] The method for collecting and analyzing and summarizing industry information in real time provided by the present application is:

[0036] In step S1, the internet tourism industry information is obtained and pretreated to obtain scenic spot information.

[0037] In the embodiment of the present application, the tourism industry information is first obtained from the internet media platforms such as microblog, WeChat public number, Douyin, Xiaohongshu and news website, which includes a lot of repeated, valueless, format-uniform, and inconsistent layout order tourism information. Therefore, after collecting a large amount of tourism information, the format is first processed, then the repeated data is removed, the missing invalid data is removed, the short sentences are deleted, the English, numbers and characters are deleted, the stop words and Chinese word segmentation are removed, and finally the layout order of the tourism industry information is adjusted. All the tourism information is uniformly laid out according to the order of scenic spot pictures, scenic spot introduction, scenic spot strategy and tourist comments.

[0038] In step S2, the scenic spot information after pretreatment is segmented and integrated to obtain seven pieces of complete scenic spot information of a scenic spot.

[0039] In the embodiment of the present application, the step of segmenting and integrating the preprocessed scenic spot information includes: dividing the preprocessed travel industry information into seven sections of weather, transportation, accommodation, food, shopping, entertainment, and scenic spot guide; randomly dividing the information in each section into three groups, and randomly selecting one piece of information from each group each time for pairwise comparison using a file similarity comparison method, wherein the file similarity comparison method is a cosine text similarity calculation method based on tf-idf, and when the similarity of two files exceeds a predetermined threshold, it is considered that the two news are duplicate files, and only one of them needs to be retained. For example, information 1 and information 2 are N1 and N2. The tf-idf feature vectors of N1 and N2 are f1 and f2. When the text similarity similarity(f1, f2) > a, where a is a threshold value between 0 and 1, it is considered that N1 and N2 are duplicate information. Of course, the deduplication algorithm in this embodiment is only an example, and other algorithms that can calculate text similarity are also applicable; if the similarity of the contents of two of the three information is very high, the more concise of the two information is filtered out, if the similarity of the contents of two of the three information is relatively high, the two information is integrated into a new information by a file integration method, if the similarity of the three information is not high, the next group of three information is introduced, and the similarity is compared, and the three information with the lowest similarity is screened out, if the similarity still does not meet the screening standard, the next group of three information is introduced, and the six information with the lowest similarity is screened out; after each piece of information is compared, the remaining information is repeatedly grouped and screened; after several rounds of screening, three information is screened out, and then the three information is integrated into a complete information information, wherein the similarity greater than 85% is determined as high similarity, the similarity greater than 55% is determined as relatively high similarity, and the similarity less than 55% is determined as low similarity.

[0040] In the S3 step, a plurality of keywords are extracted from each of the complete information information of the scenic spots by NLP technology to obtain scenic area keywords.

[0041] In the embodiment of the present application, the seven section information in the scenic spot information package is divided into a plurality of words by NLP technology; the plurality of words are packaged and screened by big data to remove common words and general words; the keywords with the characteristics of the scenic area are screened out according to the characteristics and historical information of the scenic area; a generative large language model is used to compare and analyze the keywords, and the most characteristic word is selected as the scenic area keyword of the section. Finally, a scenic area has seven scenic area keywords. In the process of extracting keywords, adjectives, quantifiers, prepositions and other meaningless conjunctions in the information are removed first, and then the keywords are extracted from the remaining information content by NLP natural language processing technology.

[0042] In step S4, the scenic spot keywords are marked, and the marking records are collected for processing.

[0043] In the embodiment of the application, the scenic spot keywords are first marked uniformly with implicit data; when a tourist selects a keyword each time to view the corresponding single scenic spot information, the marking is counted once; every 100 times, the marking is upgraded to a secondary marking; the data of the marking is regularly counted in a month, and the single scenic spot information is arranged in order according to the statistical data.

[0044] In step S5, a comment marking is bound to an information viewer, and the comment marking can obtain the comment time and comment location of the information viewer.

[0045] In the embodiment of the application, a comment marking associated with a date and a GPS is bound to a user when the viewer enters the scenic spot information for the first time, and the marking works when the viewer comments on the scenic spot.

[0046] In step S6, the comment area is sorted and classified by the comment time and the comment location.

[0047] In the embodiment of the application, the comment date obtained by the comment marking and the address of the corresponding comment scenic spot are used to obtain the weather condition of the scenic spot on the day by artificial intelligence technology; the comment marking can be associated with time and GPS positioning, and can be used to determine whether the commenter is a local resident or an out-of-town tourist participating in the scenic spot information comment according to the city of the comment location obtained by the comment marking; then the obtained comment date, the weather condition of the scenic spot on the day, and the commenter category are marked as corresponding labels on the corresponding comment; the comment location, the comment time, and the weather condition of the scenic spot obtained by the comment label enable the viewer to obtain different categories of comments according to different labels; the comment area has tabs, and the tab options include the comment time, the comment location, and the weather condition, which can be arbitrarily selected by the viewer.

[0048] In step S7, the comment contents of the existing information comment records are sorted according to the comment time, and the sorted comment contents are provided to the information viewer.

[0049] In the embodiment of the application, the comment contents of the existing comment records are sorted in descending order according to the update time of the existing comment records, so that the comment contents of the existing comment records with the latest update time are located at the top of the comment contents; the comment is cut into a plurality of sentences by using artificial intelligence to identify the comment with punctuation marks as nodes; the comment is sorted in descending order by comparing the number of sentences; the comment with the highest value and the latest update time is preferentially displayed by double screening of the comment time and the comment quality, thereby saving the time of the viewer to view the comment and improving the browsing experience.

[0050] In step S8, the scenic spot information value is obtained according to the sorted comment content.

[0051] In the embodiment of the application, the reason why the scenic spot information viewer judges whether a scenic spot has tourism value is mainly in the comments of other tourists. Since the comments of other tourists on the scenic spot to a large extent reflect the experience in the tourism process, the comments are more convincing. After the comment content is optimized and compared, the comparison of the scenic spot information comments can reflect the scenic spot information value of the scenic spot.

[0052] In step S9, the complete information information is scored and starred according to the scenic spot information value, and the scenic spot keywords and the starred results are displayed on the cover of the complete information information.

[0053] In the embodiment of the application, the complete information information is divided into four parts, namely, comment score, characteristic score, environment score and heat score. The comment score accounts for 30%, the characteristic score accounts for 20%, the environment score accounts for 30%, and the heat score accounts for 20%. The overall score is the sum of the four parts. The comment score is determined by the proportion of good comments and bad comments in the comments. The characteristic score is determined by the repetition rate of the scenic spot keywords. The environment score is determined by the number and proportion of good comments of tourists from other places and the number and proportion of bad comments of local residents. The heat score is determined by the number of comments marked by the viewer and the number of comments marked by the viewer.

[0054] The above is only an embodiment of the application, and well-known specific technical solutions or characteristics in the scheme are not described in detail. It should be noted that for those skilled in the art, without departing from the scheme of the application, some modifications and improvements can be made, which should also be considered as the protection scope of the application. The protection scope of the present application should be subject to the content of its claims, and the specific implementation mode and the like in the specification can be used to explain the content of the claims.

Claims

1. A method for real-time collection, analysis, and summarization of industry information, characterized in that, The method includes the following steps: S1: Obtain and preprocess Chinese travel industry information from the internet to obtain attraction information; S2: The preprocessed scenic spot information is segmented and integrated to obtain seven complete scenic spot information articles for each scenic spot; S3: Extract multiple keywords from the complete information of each attraction using NLP technology to obtain the scenic area keywords; S4: Tag the keywords of the scenic area and collect and process the tagging records; S5: Bind a comment tag to the information viewer, the comment tag being able to obtain the information viewer's comment time and comment location; S6: Organize and categorize the news comment section by comment time and comment location; S7: Sort the comments in the existing news comment records according to the comment time, and display the sorted comments to the news viewers; S8: Determine the information value of the scenic area based on the sorted comments; S9: Based on the information value of the scenic area, the complete information information will be rated and starred, and the scenic area keywords and starred results will be displayed on the cover of the complete information information of the scenic spot.

2. The method for real-time collection, analysis, and summarization of industry information according to claim 1, characterized in that, The steps of acquiring and preprocessing Chinese travel industry information from the internet include: Format conversion: The collected information on the cultural and tourism industry, which has inconsistent file formats, is converted into a unified file format using a file format converter; Clean up useless data: Remove duplicate data from the cultural and tourism industry information, remove missing invalid data, delete short sentences, delete English words, numbers and characters, and remove stop words and Chinese word segments; Unified Template: The cultural and tourism industry information will be arranged in the order of scenic spot pictures, scenic spot introduction, scenic spot guide, and tourist reviews to form a unified template for scenic spot information.

3. The method for real-time collection, analysis, and summarization of industry information according to claim 2, characterized in that, The step of segmenting and integrating the preprocessed attraction information includes: The pre-processed attraction information is divided into seven sections: weather, transportation, accommodation, food, shopping, entertainment, and attraction travel guide. The scenic spot information in each section is randomly divided into three groups. Each time, one piece of scenic spot information is randomly selected from each group and compared pairwise using a file similarity comparison method. If two of the three pieces of information about the scenic spot have a high degree of similarity, then the one with the shorter information will be filtered out. If two of the three pieces of information about tourist attractions have a high degree of similarity, then the two pieces of information will be integrated into a new piece of information about tourist attractions using a file integration method. If the similarity of the three scenic spot information pieces is not high, then the next set of three scenic spot information pieces is introduced, and the similarity is compared. The three scenic spot information pieces with the lowest similarity are filtered out. If the similarity still does not meet the filtering criteria, then the next set of three scenic spot information pieces is introduced, and the six scenic spot information pieces with the lowest similarity are filtered out. After each piece of information about a tourist attraction has been compared, the remaining information about tourist attractions will be grouped and filtered again. After multiple rounds of screening, three pieces of information about the scenic spots will be selected. Then, the three pieces of information about the scenic spots will be combined in pairs to form a complete piece of information. In the similarity comparison, a similarity greater than 85% is considered very high, a similarity greater than 55% is considered relatively high, and a similarity less than 55% is considered not high.

4. The method for real-time collection, analysis, and summarization of industry information according to claim 3, characterized in that, The steps of the file similarity comparison method include: First, calculate the keywords for each piece of information about the scenic spot, then select the same number of keywords from each set and merge them into one set; Then, the word frequency of each scenic spot information item in the set for the keywords is calculated, and word frequency vectors are generated for each item. The cosine similarity between the word frequency vectors of the two scenic spot information documents is obtained by Euclidean distance or cosine distance. If the similarity exceeds a set threshold, the two scenic spot information documents are considered similar; otherwise, they are considered dissimilar.

5. The method for real-time collection, analysis, and summarization of industry information according to claim 3, characterized in that, The steps of the file integration method include: First, the information about the attractions mentioned in the two articles is divided into integrated information and information that is being integrated; The two articles about tourist attractions were segmented into several words using an article segmentation system. Using paragraphs as units, the vocabulary from the two segments is fed into NLP technology for vocabulary comparison. Based on the amount of repeated words, the segment being integrated is either deleted or added to the integrated information.

6. The method for real-time collection, analysis, and summarization of industry information according to claim 3, characterized in that, The step of extracting multiple keywords from the complete information of each of the scenic spots using NLP technology to obtain scenic spot keywords includes: The seven sections of information in the complete information were divided into several words using NLP technology. Package a number of words, use big data to search and filter them, and remove common and general words; Keywords that highlight the unique characteristics of the scenic area were selected based on its features and historical information. Generative large language models are used to compare and analyze keywords, and the most distinctive word is selected as the scenic spot keyword for this section. In the end, each scenic spot has seven scenic spot keywords. Selecting a keyword allows you to view information about the corresponding individual scenic spot.

7. The method for real-time collection, analysis, and summarization of industry information according to claim 6, characterized in that, The steps of tagging scenic area keywords and collecting and processing the tagging records include: First, tag the keywords of the scenic area with data; Each time a tourist selects a keyword to view information about a single attraction, it counts once. Regularly collect statistics and rank the information of individual attractions according to the statistics.

8. The method for real-time collection, analysis, and summarization of industry information according to claim 6, characterized in that, The steps for organizing and categorizing the comment section by comment time and comment location include: Based on the comment date obtained from the comment tags and the address of the corresponding scenic spot, the weather conditions of the scenic spot on that day are obtained through artificial intelligence technology; Based on the location of the comment obtained from the comment marker, it can be determined whether the commenter is a local resident or a tourist from out of town; The date of the review, the weather conditions at the scenic spot that day, and the reviewer's category were used to create corresponding tags and marked on the respective reviews; By using comment tags to obtain information such as the comment location, comment time, and the weather conditions at the scenic spot on that day, visitors can access different categories of comments based on different tags.

9. The method for real-time collection, analysis, and summarization of industry information according to claim 8, characterized in that, The step of sorting the comment content of existing information comment records according to the comment time is as follows: The comment content of the existing comment records is sorted in descending order according to the update time of the existing comment records, so that the comment content of the existing comment record with the latest update time is at the top of the comment content section; Artificial intelligence is used to identify comments and divide them into several sentences using punctuation marks as nodes; Comments are sorted in descending order by the number of sentences.

10. The method for real-time collection, analysis, and summarization of industry information according to claim 8, characterized in that, The rating and star-rating process includes: The complete information is divided into four parts: comment score, feature score, environment score, and popularity score. The overall score is the sum of the four parts. The score for a comment is determined by the ratio of positive to negative comments. The distinctive score is determined by the repetition rate of keywords in the scenic area. The environmental score is determined by the number of positive reviews from tourists and the number of negative reviews from local residents. The popularity score is determined by the number of views of the complete information and the number of comment tags attached by the viewers.

Citation Information

Patent Citations

  • Tourist attraction evaluation information quality effectiveness analysis method based on integrated learning data mining technology

    CN115018255A

  • Comment classification method and device, equipment and storage medium

    CN115168677A