Consumption data analysis system for collecting scenic spot travel information
By introducing a user data purification module into the scenic spot cultural and tourism information consumption data analysis system, the Jaccard similarity is calculated and the consumption characteristics of relatives and friends groups are eliminated, the problem of inaccurate analysis of popular projects in the scenic spot due to the existence of relatives and friends groups is solved, and more accurate analysis of popular projects and resource allocation is achieved, and the overall competitiveness of the scenic spot is enhanced.
Patent Information
- Application Number
- CN202510149265.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When analyzing popular projects in the scenic area, the analysis is inaccurate due to the existence of relatives and friends.
It provides a consumption data analysis system for collecting cultural and tourism information of scenic spots, including user information acquisition module, popular project analysis module and user data purification module. The Jaccard similarity of multiple user play projects is calculated through the user data purification module, and the relatives and friends group is judged based on the time threshold, and the consumption characteristics of relatives and friends group are eliminated in order to accurately analyze popular projects.
Effectively eliminate interference factors from relatives and friends, so that the analysis results of popular projects can more truly reflect tourists' independent choices and the actual attractiveness of the projects, help scenic spots to reasonably allocate resources and enhance overall competitiveness.
Smart Images

Figure CN120070104A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and in particular to a consumption data analysis system for collecting cultural and tourism information of scenic spots. Background Art
[0002] In scenic area operations and management, existing technology primarily analyzes the number of times users visit attractions within the area through consumption data to identify popular attractions. This information is then fed back to staff, who then use this feedback to develop similar attractions.
[0003] However, in actual scenic area visits, it is quite common for tourists to be accompanied by others. This is particularly true for less stimulating attractions that are more suitable for multiple people to experience together, such as sightseeing cable cars and boat tours. Some tourists may not have originally intended to visit these attractions, but participate in them for the purpose of accompanying relatives and friends. When counting the number of visits, the current calculation method will include this type of accompanying participation, which leads to an inflated number of visits to certain attractions. If the data on the projects accompanied by relatives and friends cannot be accurately eliminated, it will affect the analysis of popular projects. In view of this, we propose a consumption data analysis system that collects cultural and tourism information from scenic spots. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem of inaccurate analysis caused by the presence of friends and family groups when analyzing popular projects in a scenic area.
[0005] To achieve the above-mentioned purpose, the present invention provides a consumption data analysis system for collecting cultural and tourism information of scenic spots, including a user information acquisition module, a popular project analysis module and a user data purification module;
[0006] The user information acquisition module is used to establish a connection channel with the scenic area database server through the information acquisition method to obtain the consumption characteristics and promotional materials of each user in the scenic area;
[0007] The popular item analysis module retrieves the number of times an item has been played in the play records of old users, where old users refer to users who have visited the scenic spot ≥1 time, calculates the proportion of the number of times each item has been played, and uses a direct comparison method to traverse the proportion of the number of times all items have been played, sorting them from high to low according to the proportion, and defining the item with the highest proportion as the target item;
[0008] Using a part-of-speech tagging tool, the key words for different ways of playing the attractions in the scenic spot promotional materials are determined and extracted, and the key words for the same attractions are merged into a keyword set. The similarity between the target attraction keyword set and the other attractions keyword sets is calculated, and a similarity threshold is set. If the similarity is greater than the similarity threshold, the target attraction is determined to be the same as the other attractions and is defined as a popular attraction. The ways of playing the popular attractions are output to the scenic spot database server through the connection channel in the user information acquisition module to prompt staff members of the attractions that can be added to the scenic spot.
[0009] The user data purification module is used to treat the play items selected by each user as a set, where the elements in each set are the play items selected by the user. For each pair of user play item sets, the module traverses the elements in the two sets to find the intersection and union of the two sets, and calculates the Jaccard similarity based on the ratio of the number of elements in the intersection to the number of elements in the union. When a user type determination method is used to determine that multiple users are friends and family groups, the consumption characteristics of the friends and family groups in the user information acquisition module are eliminated.
[0010] As a further improvement to the present technical solution, the information acquisition method in the user information acquisition module specifically constructs a data packet based on an established network communication protocol. In addition to the necessary network layer and transport layer header information for initiating a connection request, the data packet also embeds the sender's identity information in a specific format.
[0011] The user information acquisition module sends the constructed data packet to the local area network environment of the scenic spot through the network interface, and the data packet is forwarded by the network device according to the IP address and finally arrives at the network port of the scenic spot database server. After the scenic spot database server receives the data packet sent by the user information acquisition module, the server performs an unpacking operation according to the rules of the network communication protocol and accurately extracts the identity information from the data packet;
[0012] The server compares the extracted identity information with the legitimate user identity records pre-stored in the database. If the identity information is identical and matches completely, the authentication is deemed to be successful. The server establishes a reliable connection channel with the user information acquisition module according to the TCP connection establishment process in the TCP / IP protocol;
[0013] The user information acquisition module sends a data request to the scenic area database server through the established connection channel in accordance with the agreed data format and protocol to obtain the consumption characteristics and promotional materials of each user in the scenic area. After receiving the request, the scenic area database server parses the request statement and retrieves the consumption characteristics and promotional materials of each user in the scenic area according to the content of the request.
[0014] As a further improvement to this technical solution, the formula for calculating the proportion of play times of each game item in the popular game analysis module is as follows:
[0015] Where n represents the total number of attractions in the scenic area, and i represents the i-th attraction;
[0016] The direct comparison method selects the item with the largest proportion from all the items, exchanges it with the item at the current position, and then traverses all the items in turn to sort them from high to low proportion.
[0017] As a further improvement of the present technical solution, the part-of-speech tagging tool determines the part of speech according to grammatical rules, which are a set of detailed part-of-speech tagging rules formulated in advance;
[0018] When tagging the parts of speech for scenic spot promotional materials, the part-of-speech tagging tool matches the words in the text one by one with pre-set rules. If the part of speech is a verb and the verb describes the actual operational behavior during the tour, the word is extracted as a keyword. If the part of speech is a noun, the word is extracted as a keyword. If the part of speech is an adjective and the adjective is closely related to the gameplay and has a modifying effect on the gameplay, the word is extracted as a keyword.
[0019] As a further improvement of this technical solution, the popular item analysis module calculates the similarity between the target item keyword set and other game item keyword sets: in and are the gameplay feature vectors in the keyword sets of the target project and other projects respectively, is the dot product of two vectors, and are the magnitudes of the two vectors respectively.
[0020] As a further improvement of this technical solution, the game item sets in the user data purification module are respectively set C and set D. Each element in the two sets is compared through nested loops. The outer loop traverses the elements in set C. For each element C in set C, i , the inner loop traverses all elements D in the set D i , in each inner loop, if C i =D i , then determine C i and D i Existing in both set C and set D, that is, comparing all combinations of elements and calling out the common elements in the game item sets, that is, the intersection;
[0021] When traversing set C, add each element in set C to the union set in turn, and then traverse set D. If each element in set D does not exist in the union set, add it to the union set.
[0022] As a further improvement of this technical solution, the calculation formula for calculating the Jaccard similarity of the user data purification module is as follows:
[0023] Here, |C∩D| represents the number of elements in the intersection of sets C and B, that is, the number of game items selected by both users, and |C∪D| represents the number of elements in the union of sets A and B, that is, the number of all game items selected by both users (after removing duplicate items).
[0024] As a further improvement of the present technical solution, the user category determination method in the user data purification module calls out the play items corresponding to multiple users when j(C, D)=1, and senses the play time of the play items in the consumption characteristics obtained by the user information acquisition module, and sets a time threshold. If j(C, D)=1 and the interval between the user play times is less than the time threshold, the multiple users are determined to be a group of relatives and friends.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] In the consumption data analysis system for collecting cultural and tourism information of scenic spots, a connection channel is established with a scenic spot database server through a user information acquisition module to obtain the consumption characteristics and promotional materials of each user in the scenic spot, and then the number of times the play items in the play records of old users are called out through a popular project analysis module, the proportion of the number of times each play item is played is calculated, and the proportion reflects the popularity of each project among tourists, and the similarity between the play items and the project with the highest proportion is calculated, and a similarity threshold is set to determine other game projects that are the same as the game project with the highest proportion, and the determined game project and the game project with the highest proportion are defined as popular projects. The popular projects are fed back to the scenic spot database server through the connection channel established by the user information acquisition module, providing strong support for the long-term strategic planning of the scenic spot, helping the scenic spot to reasonably allocate resources, and investing more resources in popular projects and related expansions, thereby improving the overall competitiveness of the scenic spot;
[0027] Before feedback, the user data purification module retrieves each user's play items, calculates the Jaccard similarity of multiple users' play items, and sets a time threshold. When the Jaccard similarity is 1 and the interval between users' play times is less than the time threshold, the multiple users are determined to be a group of friends and family, and the consumption characteristics of the group are removed from the user information acquisition module. During scenic area visits, the presence of friends and family in a group can lead to an inflated number of visits to certain attractions, affecting the accurate identification of popular attractions. By calculating the Jaccard similarity of multiple users' play items, combining the time threshold to identify friends and family groups, and then removing the consumption characteristics of these groups, this interference factor can be effectively eliminated, ensuring that the analysis results of popular attractions more truly reflect tourists' independent choices and the actual appeal of the attractions. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 This is the overall module principle diagram of the present invention.
[0029] The meaning of each number in the figure is:
[0030] 100. User information acquisition module; 200. Popular project analysis module; 300. User data purification module. DETAILED DESCRIPTION
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0032] A consumption data analysis system for collecting scenic spot cultural tourism information includes a user information acquisition module 100, a popular project analysis module 200 and a user data purification module 300;
[0033] The continuous expansion of the tourism market has led to the rapid proliferation of various scenic spots, including traditional natural and historical scenic spots, as well as emerging theme parks and distinctive towns. In this fiercely competitive environment, scenic spots must continuously enhance their competitiveness to attract more tourists and increase their market share. By analyzing visitor spending data, scenic spots can understand their strengths and weaknesses in the market, identify potential market opportunities, and develop more targeted marketing strategies, optimize products and services, and enhance their overall appeal.
[0034] As scenic spots expand in size and their businesses become increasingly complex, the amount of data generated and accumulated by scenic spots has exploded. Ticket sales records, visitor information including name, contact information, ID number, etc., operational data of scenic spot facilities and equipment such as the number of times amusement facilities are used and maintenance records, and various consumption data such as catering, accommodation, shopping, etc. are all stored in the scenic spot database server.
[0035] The information acquisition method in the user information acquisition module 100 specifically constructs a data packet based on an established network communication protocol. In addition to the necessary network layer and transport layer header information for initiating a connection request, such as the source IP address, destination IP address, port number, protocol type, etc., which are used to guide the data packet to be accurately transmitted to the scenic area database server in the network, the sender's identity information is also embedded in it according to a certain format. The identity information may include a user name, a password, and usually the password is encrypted, such as using a hash algorithm to generate a password hash value to ensure security, and the user information acquisition module 100 number and other content, which is used to prove its legitimacy in the subsequent server verification link;
[0036] The user information acquisition module 100 sends the constructed data packet to the local area network environment of the scenic spot through a network interface such as an Ethernet card, and forwards it according to the IP address through network devices such as routers and switches in the network, and finally arrives at the network port of the scenic spot database server;
[0037] The scenic area database server is equipped with a network monitoring program, which always monitors the connection request data packet from the outside on the corresponding port. After receiving the data packet sent by the user information acquisition module 100, the server first performs the unpacking operation according to the rules of the network communication protocol and accurately extracts the identity information part from the data packet.
[0038] The server compares the extracted identity information with the legitimate user identity records pre-stored in the database. If the identity information is identical and completely matches, it is considered that the identity authentication is passed. The server establishes a reliable connection channel with the user information acquisition module 100 according to the TCP connection establishment process three-way handshake in the TCP / IP protocol, effectively preventing illegal users from accessing the scenic area database, protecting the security and privacy of scenic area user data, and preventing sensitive information from being leaked or maliciously tampered with.
[0039] The user information acquisition module 100 sends a data request for each user's consumption characteristics and promotional materials in the scenic area to the scenic area database server through the established connection channel in accordance with the agreed data format and protocol. After receiving the request, the scenic area database server parses the request statement and retrieves the consumption characteristics and promotional materials of each user in the scenic area according to the content of the request.
[0040] The popular project analysis module 200 retrieves the number of times of playing the projects in the play records of old users, where old users refer to those who have visited the scenic area ≥ 1 time, and calculates the percentage of the number of times of playing each project:
[0041]
[0042] where n represents the total number of play projects in the scenic area, and i represents the i-th play project;
[0043] The play records of old users, that is, those who have visited the scenic area ≥ 1 time, contain rich and valuable information. By retrieving the number of times of playing each project by old users, the scenic area can deeply understand the popularity of different projects among repeat visitors. Compared with the play records of new users, since new users are often exposed to the scenic area for the first time, their play choices may be affected by various temporary factors, such as the scenic area's publicity and promotion, recommendations of current popular projects, etc. Their play records more reflect the first-impression attractiveness of the projects and may not necessarily reflect the long-term and in-depth value of the projects;
[0044] From all the play projects, each time select the play project with the largest percentage, exchange its position with the play project at the current position, and traverse all the play projects in turn to achieve sorting from high to low. Select the play project with the highest percentage as the target project. The target project enables the scenic area to clearly understand the most attractive and popular projects of itself and identify the core play projects;
[0045] Extract the keyword vocabulary of the play methods of different play projects in the scenic area's publicity materials, and merge the keyword vocabulary of the same play project into a keyword set. For each vocabulary w in the scenic area's publicity materials, first use a词性标注工具 (lexical category tagging tool) to judge the词性 (lexical category);
[0046] The lexical category tagging tool determines the lexical category according to the grammar rules specified by linguists or professionals. The grammar rules are a set of detailed lexical category tagging rules formulated in advance. For example, in Chinese, it is stipulated that verbs with the endings "着", "了", "过" are generally in the form of verbs attached with dynamic auxiliaries. For example, the "跑", "吃", "看" in "跑着", "吃了", "看过" are verbs; those ending with "子", "儿", "头", etc. are often nouns, such as "桌子", "花儿", "石头", etc.; for English, the rules involve the ending changes, prefixes and suffixes of words. For example, words with suffixes such as "-tion", "-ment" at the end are usually nouns, such as "information", "development", etc.
[0047] Text Matching and Annotation: When annotating the词性 of scenic area promotional materials, the词性 annotation tool matches each word in the text with pre-set rules one by one. For example, when encountering the word "table" in the text, it is found to conform to the ending characteristics and other determination rules of a noun through the matching rules, and it is then annotated as a noun; in this way, all the words in the text are matched and annotated in turn to determine the词性 of each word;
[0048] If the词性 is a verb V, and the verb describes an actual operation behavior during the play process, that is, V ∈ {set of verbs describing play operation behaviors}, then the word w is extracted as a keyword. For example, "jump", in some play items such as trampolines and obstacle crossing, belongs to the verb describing play operation behaviors and meets the extraction conditions;
[0049] If the词性 is a noun N, and the noun belongs to the key elements such as props, scenes, objects, etc. involved in the play project, that is, N ∈ {set of nouns such as props, scenes, objects, etc. related to the play project}, then the word w is extracted as a keyword. For example, "slide", in the slide project in the children's play area, belongs to the key object of the play project and meets the extraction standard;
[0050] If the词性 is an adjective A, and the adjective is closely related to the gameplay and has an important modifying effect on the nature, characteristics, etc. of the gameplay, that is, a ∈ {set of adjectives closely related to the gameplay}, then the word w is extracted as a keyword. For example, "competitive", in some competitive play projects, indicates that the gameplay has a competitive nature and belongs to the adjective closely related to the gameplay and should be extracted;
[0051] Calculate the similarity between the keyword set of the target project and the keyword sets of other play projects: Among them and are the gameplay feature vectors in the keyword sets of the target project and other projects respectively, is the sum of the products of the corresponding elements after multiplying the dot product of the two vectors, and<lable id="0000134">are the norms of the two vectors calculated by taking the square root of the sum of the squares of the elements of the vectors respectively;
[0052] A similarity threshold is set. If the similarity is greater than the similarity threshold, the target project is determined to be the same as other recreational projects and is defined as a popular project. The target project is determined not only based on the proportion of the number of times played, but also by calculating the similarity of the project keyword set, other projects with high similarity to the target project are also included in the category of popular projects, so that the staff can find multiple projects with similar gameplay and high popularity, and more comprehensively identify the truly popular recreational project sets in the scenic area, avoiding the one-sidedness of the single criterion of the number of times played, and more accurately locating the popular projects. The gameplay of the popular projects is output to the scenic area database server through the connection channel in the user information acquisition module 100, providing strong data support for the decision-making of the scenic area. The staff can formulate more reasonable operation strategies based on the determination of the target projects.
[0053] The present invention also takes into account the situation that when visiting scenic spots, friends and family members often accompany each other to play projects. For example, in some projects with low stimulation and suitable for multiple people to experience together, such as sightseeing cable cars and cruise ships, one member may not originally have the intention to play the project, but participates in it to accompany relatives and friends. In this case, when calculating the number of times played, these accompanying participations will also be counted, making the number of times played for certain projects inflated, affecting the accuracy of the calculation of the proportion of the number of times played for the popular project analysis module 200. Therefore, the user data purification module 300 regards the play projects selected by each user as a set, and the elements in each set are the play projects selected by the user. For each pair of user play project sets, each element in the two sets is compared through a nested loop, and the outer loop traverses the elements in the set C. For each element C in the set C, i , the inner loop traverses all elements D in the set D i , in each inner loop, if C i =D i , then determine C i and D i Existing in both set C and set D, that is, comparing all combinations of elements and calling out the common elements in the game item set, that is, the intersection represents the common game options;
[0054] When traversing set C, add each element in set C to the union set in turn. Then traverse set D. If each element in set D does not exist in the union set, add it to the union set. The union is the set containing all the game items selected by the two users, after removing duplicate items. The more elements in the intersection, the more similar the game items of the two users are relative to the union.
[0055]
[0056] Where |C∩D| represents the number of elements in the intersection of sets C and B, that is, the number of game items selected by both users; |C∪D| represents the number of elements in the union of sets A and B, that is, the number of all game items selected by both users after removing duplicate items;
[0057] By calculating the ratio of the number of elements in the intersection to the number of elements in the union, the similarity between the two users' game items is quantified as a value between 0 and 1, which is the Jaccard similarity. It directly reflects the similarity between the two users' game item selections.
[0058] Retrieve the play items corresponding to multiple users with j(C, D) = 1, and sense the play time of the play items in the consumption characteristics obtained by the user information acquisition module 100, set a time threshold, and if j(C, D) = 1, and the interval between the user play times is less than the time threshold, then the multiple users are determined to be a group of friends and family, and the consumption characteristics of the group of friends and family in the user information acquisition module 100 are eliminated. When analyzing tourists' consumption willingness for a single item, if the group purchase consumption data of the group of friends and family is not eliminated, the actual market demand for the item will be overestimated. After elimination, the staff can accurately understand the demand for various resources by different tourist groups.
[0059] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A consumption data analysis system for collecting cultural and tourism information of scenic spots, characterized in that: It includes a user information acquisition module (100), a hot item analysis module (200) and a user data purification module (300); The user information acquisition module (100) is used to establish a connection channel with the scenic area database server through an information acquisition method to obtain the consumption characteristics and promotional materials of each user in the scenic area; The popular item analysis module (200) retrieves the number of times the items are played in the play records of old users, calculates the percentage of the number of times each item is played, and sorts the items from high to low according to the percentage, and defines the item with the highest percentage as the target item; Using part-of-speech tagging tools, determine and extract key words for different ways of playing the attractions in the scenic spot promotional materials, and merge the key words of the same attractions into a keyword set, calculate the similarity between the target item keyword set and the other items keyword set, set a similarity threshold, and if the similarity is greater than the similarity threshold, the target item is determined to be the same as the other attractions, defined as a popular item, and a channel is connected to output the popular item's gameplay to the scenic spot database server; The user data purification module (300) regards the recreation items selected by each user as a set, wherein the elements in each set are the recreation items selected by the user, finds the intersection and union in the set, calculates the Jaccard similarity, and adopts the user type determination method. When it is determined that multiple users are a group of friends and relatives, the consumption characteristics of the group of friends and relatives in the user information acquisition module (100) are removed.
2. The consumption data analysis system for collecting scenic spot cultural and tourism information according to claim 1 is characterized by: The information acquisition method in the user information acquisition module (100) specifically constructs a data packet according to an established network communication protocol, wherein the data packet contains not only the necessary network layer and transport layer related header information for initiating a connection request, but also the identity information of the sender is embedded in the data packet in a certain format; The user information acquisition module (100) sends the constructed data packet to the local area network environment of the scenic spot through the network interface, and the data packet is forwarded by the network device according to the IP address, and finally arrives at the network port of the scenic spot database server. After the scenic spot database server receives the data packet sent by the user information acquisition module (100), the server performs an unpacking operation according to the rules of the network communication protocol, and accurately extracts the identity information from the data packet; The server compares the extracted identity information with the legal user identity records pre-stored in the database. If the identity information is identical and completely matches, the identity authentication is deemed to be successful. The server establishes a reliable connection channel with the user information acquisition module (100) according to the TCP connection establishment process in the TCP / IP protocol. The user information acquisition module (100) sends a data request for acquiring the consumption characteristics and promotional materials of each user in the scenic area to the scenic area database server through the established connection channel in accordance with the agreed data format and protocol. After receiving the request, the scenic area database server parses the request statement and retrieves the consumption characteristics and promotional materials of each user in the scenic area according to the content of the request.
3. The consumption data analysis system for collecting scenic spot cultural and tourism information according to claim 1 is characterized by: The calculation formula for calculating the proportion of the number of times each game item is played in the popular game analysis module (200) is as follows: Where n represents the total number of attractions in the scenic area, and i represents the i-th attraction; The direct comparison method selects the item with the largest proportion from all the items each time, exchanges it with the item at the current position, and then traverses all the items in turn to sort them from high to low proportion.
4. The consumption data analysis system for collecting scenic spot cultural and tourism information according to claim 3 is characterized by: The part-of-speech tagging tool determines the part-of-speech according to grammatical rules, which are a set of detailed part-of-speech tagging rules formulated in advance; When tagging the parts of speech of scenic spot promotional materials, the part-of-speech tagging tool matches the words in the text one by one with the pre-set rules. If the part of speech is a verb and the verb describes the actual operational behavior during the tour, the word is extracted as a keyword. If the part of speech is a noun, the word is extracted as a keyword. If the part of speech is an adjective and the adjective is closely related to the gameplay and has a modifying effect on the gameplay, the word is extracted as a keyword.
5. The consumption data analysis system for collecting scenic spot cultural and tourism information according to claim 4 is characterized by: The hot item analysis module (200) calculates the similarity between the target item keyword set and other game item keyword sets: in and are the gameplay feature vectors in the keyword sets of the target project and other projects, is the dot product of two vectors, and are the magnitudes of the two vectors respectively.
6. The consumption data analysis system for collecting scenic spot cultural and tourism information according to claim 5 is characterized by: The game item sets in the user data purification module (300) are respectively set C and set D. Each element in the two sets is compared by nested loops. The outer loop traverses the elements in set C. For each element C in set C, i , the inner loop traverses all elements D in set D i , in each inner loop, if C i =D i , then determine C i and D i Existing in both set C and set D, that is, comparing all combinations of elements and calling out the common elements in the game item sets, i.e. the intersection; When traversing set C, add each element in set C to the union set in turn, and then traverse set D. If each element in set D does not exist in the union set, add it to the union set.
7. The consumption data analysis system for collecting scenic spot cultural and tourism information according to claim 6 is characterized by: The calculation formula for calculating the Jaccard similarity of the user data purification module (300) is as follows: Among them, |C∩D| represents the number of elements in the intersection of sets C and B, that is, the number of game items selected by both users, and |C∪D| represents the number of elements in the union of sets A and B, that is, the number of all game items selected by the two users (after removing duplicate items).
8. The consumption data analysis system for collecting scenic spot cultural and tourism information according to claim 7 is characterized by: The user category determination method in the user data purification module (300) calls out the play items of multiple users corresponding to j(C, D)=1, and senses the play time of the play items in the consumption characteristics obtained by the user information acquisition module (100), sets a time threshold, and if j(C, D)=1, and the interval between the user play times is less than the time threshold, the multiple users are determined to be a group of relatives and friends.