Personalized tourism lug recommendation method based on behavior data mining

By integrating users' multi-channel behavior data and using differentiated mining time intervals and hierarchical analysis, the cold start and data sparse problems in the tourism recommendation system are solved, and the accuracy and efficiency of personalized recommendations are improved.

CN120561381AActive Publication Date: 2025-08-29HANGZHOU DOUBLE LEAF NETWORK TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511062284.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-08-29
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

The existing travel package recommendation system has difficulties in the cold start of content and sparse data, and cannot effectively integrate users' multi-channel behavior data, resulting in the recommendation results that are not consistent with user needs and affecting the travel social experience.

Method used

By collecting multiple sets of behavior source data of users, it is divided into user basic situation, travel preferences, past interactions and consumption habits data, differentiated mining time intervals and hierarchical analysis methods are used to combine and analyze behavior sample combinations, calculate the degree of recommendation matching evaluation values, and dynamically adjust the recommendation strategy.

Benefits of technology

Multi-channel and multi-dimensional data integration has been achieved, the degree of fit between recommendation results and user real needs has been improved, recommendation accuracy and efficiency has been improved, invalid communication costs have been reduced, and user preferences have been dynamically learned to improve recommendation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561381A_ABST
    Figure CN120561381A_ABST
Patent Text Reader

Abstract

The invention discloses a travel lug personalized recommendation method based on behavior data mining, and relates to the technical field of travel services, the recommendation method comprises the following steps: collecting multiple groups of behavior source data of a user in a travel scene; determining different data mining time interval information according to the behavior source data; multiple batches of behavior sample combinations are obtained according to different data mining time interval information; combining and analyzing the multiple batches of behavior sample combinations to obtain a recommendation matching degree evaluation numerical value; the technical key points are as follows: capturing tourism trends published by a user from a tourism social network platform, extracting reservation records from tourism reservation software, collecting score contents from a tourism evaluation webpage, summarizing consumption details from an offline service system, and displaying the consumption details. Four types of data including basic conditions, travel preferences, past interaction and consumption habits of users are divided through dimensions such as identities and preferences; according to the scheme, the recommendation deviation caused by information fragmentation is thoroughly solved, and the integrating degree of the recommendation result and the real demand of the user is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of tourism service technology, and in particular to a personalized travel partner recommendation method based on behavioral data mining. Background Art

[0002] The tourism market during holidays is always bustling, and many people traveling alone hope to find like-minded companions on their journey. After opening a travel social platform and entering the destination and travel theme in the search box, the user list that pops up on the system only shows age, gender, and a few words of personal signature. Some signatures say "love traveling" and some are marked as "easy-going personality." After flipping through dozens of pages of content, either the time cannot be synchronized, or the other party has shared updates about their stay in high-end hotels. It is speculated that the difference in consumption habits is too great, and ultimately they have to give up looking for companions and embark on the journey alone.

[0003] After entering the destination and time in the "Find a Companion" section of the travel booking app, the system quickly recommended several users. However, further conversation revealed that some preferred shopping over participating in pre-set activities, others wanted a packed itinerary from morning to night, while the user preferred free time for wandering, and some requested a hostel, contradicting their preference for budget hotels. After multiple failed matches, they were unable to find a suitable companion.

[0004] When using the travel package recommendation function for the first time, since there is no historical travel record, the system can only randomly push recommendations to users, leaving them in a dilemma of not knowing how to choose, and even if they choose, it may not be suitable. This is a typical content cold start problem.

[0005] At the same time, users' behavioral data is scattered across different platforms: information such as equipment preferences shared on social platforms, comments about "not liking to rush itineraries" left on review websites, and "preferred homestays" set in booking software have never been systematically integrated and analyzed. When looking for a companion with a compatible pace, the system cannot identify implicit needs such as "no more than 6 hours of activity per day" and "preferring to bring own meals" from multi-channel data, resulting in a significant deviation from the expected recommendation results. This scattered data and inability to effectively integrate it creates a data sparsity problem, making it difficult for many users to find a truly suitable travel partner even if they spend a lot of time screening, seriously affecting the improvement of their travel social experience. Therefore, there is an urgent need for a personalized travel partner recommendation method based on behavioral data mining. Summary of the Invention

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A personalized travel recommendation method based on behavioral data mining, the recommendation method comprising: Collect multiple sets of user behavior source data in travel scenarios. Each set of behavior source data includes user basic information, travel preference data, past interaction data, and consumption habit data. Determine different data mining time interval information based on the above behavioral source data; According to different data mining time interval information, multiple batches of behavior sample combinations are obtained; Combine and analyze multiple batches of behavioral samples to obtain a recommended matching evaluation value; Calculate the recommendation judgment value based on the recommendation matching degree evaluation value; Comparing the recommended judgment value with the pre-set limit value; If the recommendation judgment value is greater than or equal to the preset limit value, the information of the travel companions who meet the conditions is sent to the device used by the user; otherwise, the behavior source data is reselected and the above steps are repeated.

[0007] Furthermore, the steps of collecting multiple sets of user behavior source data in the travel scenario include: Identify multiple ways to obtain user data in tourism scenarios, including at least: tourism social network platforms, travel reservation software, travel review websites, and offline travel service record systems; Collect multiple sets of raw behavioral data through various user data acquisition channels; Analyze the identity characteristics contained in each set of original behavior data, which include the user's name, age, gender, place of birth and identity identification code, and divide the corresponding original behavior data into user basic information data based on these identity characteristics; analyze the preference characteristics contained in each set of original behavior data, which include the user's preferred travel location type, travel activity type, travel time length, travel season and accommodation selection criteria, and divide the corresponding original behavior data into travel preference data based on these preference characteristics; analyze the interaction characteristics contained in each set of original behavior data, which include the number of times the user communicates with other users on travel-related platforms, the topic of the communication content, the number of times they travel together and the scores of each other's evaluations, and divide the corresponding original behavior data into past interaction data based on these interaction characteristics; analyze the consumption characteristics contained in each set of original behavior data, which include the user's dining spending standards, transportation cost range, types of attraction ticket consumption and shopping consumption tendencies during the travel process, and divide the corresponding original behavior data into consumption habit data based on these consumption characteristics.

[0008] Furthermore, the step of determining different data mining time interval information based on the above-mentioned behavioral source data includes: understanding the storage structure of the user basic information in the user information database based on the user basic information; obtaining update information of the key identity identifier based on the storage structure, the update information including the number of updates, the size of the update change, and the update trigger condition; creating an identity identifier index and a user classification tag based on the update information, the user classification tag including an age tag, a gender tag, and a region tag; recording the number of updates of the identity identifier index and the user classification tag, and using the update number as the first type of mining time interval; The user's travel records are obtained based on travel preference data. The travel record content includes information on the type of tourist location, travel time, and travel mode. The correlation between the type of tourist location, travel time, and travel mode information is calculated. The correlation is determined by counting the number and probability of different information appearing at the same time, and a preference-time correlation index is created based on the correlation. The stability of multiple preferences is calculated based on the preference-time correlation index. The preference stability is calculated by the length of time the same preference appears continuously and the size of the change. The second type of mining time interval of travel preference data is determined based on the stability of multiple preferences. Obtain interaction record content based on past interaction data and create an interaction object-frequency association index; analyze the user's dependence level on the interaction relationship based on the interaction object-frequency association index, calculate the dependence level by multiplying the number of interactions and the interaction duration, and determine the third type of mining time interval based on the dependence level; Obtain consumption record content based on consumption habit data; analyze the dynamic changes in consumption record content, and calculate the dynamic changes by the differences in consumption amounts and consumption types in adjacent time periods; calculate the consumption-preference sensitivity based on the dynamic changes, and calculate the consumption-preference sensitivity by the ratio of consumption changes to preference changes, and determine the fourth type of mining time interval based on the consumption-preference sensitivity; The first type of mining time interval, the second type of mining time interval, the third type of mining time interval and the fourth type of mining time interval are aggregated and merged. When merging, a weighted average is performed according to the importance of each time interval on the recommendation results to obtain different data mining time interval information.

[0009] Furthermore, the step of obtaining multiple batches of behavior sample combinations according to different data mining time interval information includes: Determine multiple data mining tasks based on different data mining time interval information, each data mining task corresponds to a specific behavioral data type and mining target; set mining configuration requirements for each data mining task, the mining configuration requirements include data extraction scope, data cleaning rules and data conversion standards; determine the original behavioral data sources corresponding to user basic situation data, travel preference data, past interaction data and consumption habit data based on the mining configuration requirements; formulate a mining sequence arrangement for each original behavioral data source, the mining sequence arrangement includes the mining order, mining times and mining interval time; conduct multiple preliminary mining on each original behavioral data source according to the preset mining sequence arrangement, identify and correct abnormal data during the mining process, judge abnormal data by the deviation from the average value of historical data, and obtain mining information; merge the mining information corresponding to user basic situation data, travel preference data, past interaction data and consumption habit data, and use data alignment when merging to make different types of data consistent in time and content, forming multiple batches of behavioral sample combinations.

[0010] Furthermore, the step of merging and analyzing the combination of multiple batches of behavioral samples to obtain a recommended matching degree evaluation value includes: obtaining the structural form of the multiple batches of data and the content of the multiple batches of data from the combination of multiple batches of behavioral samples; Obtaining structured result content according to the structural forms of the multiple batches of data, the structured result content including normalized data result content, semi-normalized data result content, and non-normalized data result content; The behavior consistency evaluation value is calculated based on the results of normalized data, semi-normalized data, and non-normalized data. The behavior consistency evaluation value is obtained by counting the proportion of matches of the same behavior description in data of different structures. Extracting content features from the content of multiple batches of data, including information completeness, information matching, information timeliness, information accuracy, information relevance, and behavior consistency; The content feature evaluation values ​​are calculated using a hierarchical analysis method, in which the importance of each content feature is determined by constructing a judgment matrix; The recommended matching degree evaluation value is calculated based on the behavior consistency evaluation value and the content feature evaluation value. The recommended matching degree evaluation value is the weighted sum of the behavior consistency evaluation value and the content feature evaluation value calculated in a ratio of 3:7.

[0011] Furthermore, the step of calculating the recommendation judgment value based on the recommendation matching degree evaluation value includes: Determine the importance of user profile data, travel preference data, past interaction data, and consumption habit data in travel partner recommendation. The importance of data is determined by expert ratings combined with user feedback data, with user profile data accounting for 20%, travel preference data for 35%, past interaction data for 25%, and consumption habit data for 20%. Count the actual amount of data corresponding to user basic information data, travel preference data, past interaction data, and consumption habit data, and calculate the actual amount of data by the number of data records; The matching quality evaluation value corresponding to the user's basic information data, travel preference data, past interaction data, and consumption habit data is calculated based on the recommended matching degree evaluation value, data importance ratio, and actual data quantity. The matching quality evaluation value is the product of the recommended matching degree evaluation value, data importance ratio, and actual data quantity; Count the total verified data volume corresponding to user basic information data, travel preference data, past interaction data, and consumption habit data. The total verified data volume is the sum of the number of verified accurate data items in each type of data; The recommended judgment value is calculated based on the total verification data volume and the matching quality evaluation value. The recommended judgment value is the ratio of the matching quality evaluation value to the total verification data volume.

[0012] Furthermore, after sending the information of the qualified travel companions to the device used by the user, the method further includes: Receive user feedback data on recommended travel companion information, including acceptance of recommendations, rejection of recommendations, and reasons for rejection; classify and organize the feedback data, and establish associations between the feedback data and data from each link in the corresponding recommendation process; analyze the influencing factors of recommendation success or failure based on the association results; and adjust the parameter settings of subsequent data mining and the calculation method of recommendation matching degree evaluation based on the influencing factors.

[0013] Furthermore, in the process of reselecting behavioral source data, the following steps are included: Analyze the specific data types in the previously selected behavioral source data that cause the recommended judgment values ​​to not meet the standards; expand the data acquisition scope or improve the data screening standards for specific data types; preprocess the newly acquired behavioral source data, including removing duplicate data and supplementing missing data; and perform data mining and recommendation judgment on the preprocessed behavioral source data according to the original steps.

[0014] Furthermore, when collecting multiple sets of raw behavioral data through various user data acquisition channels, it also includes: The original behavioral data obtained through various channels are subject to data legitimacy verification, including the authorization of the data source and the compliance of the data content; only the original behavioral data that has passed the legitimacy verification is retained for subsequent processing.

[0015] Furthermore, when the hierarchical analysis method is used to calculate the content feature evaluation value, it also includes: Regularly collect user satisfaction evaluations on recommendation results; adjust the importance weights of each content feature in the judgment matrix based on the satisfaction evaluations; use the adjusted judgment matrix to recalculate the content feature evaluation values ​​to optimize the accuracy of the recommendation matching degree evaluation values.

[0016] The present invention provides a personalized travel recommendation method based on behavioral data mining, which has the following beneficial effects: 1. During the data collection phase, the system, in accordance with the technical solution, captures user-posted travel updates and interactive comments from travel social networking platforms, extracts booking records from travel reservation software, collects ratings from travel review webpages, and summarizes consumption details from offline service systems. The system then divides the data into four categories: user basic information, travel preferences, past interactions, and consumption habits, using dimensions such as identity characteristics and preferences. This multi-channel, multi-dimensional data integration approach can fully outline users' underlying needs. For example, data such as frequent bookings of seaside hotels with diving options and high-frequency mentions of seafood restaurants in food reviews can not only clarify users' destination preferences, but also accurately identify their activity tendencies and consumption levels. Compared to traditional recommendations that rely solely on personal signatures, this solution completely eliminates the recommendation bias caused by information fragmentation, significantly improving the alignment of recommendation results with users' real needs.

[0017] 2. When determining the data mining interval, the system develops differentiated strategies for different data types: user basic information data uses a longer mining cycle due to its low update frequency; travel preference data may fluctuate with the seasons, so the mining frequency is adjusted quarterly; past interaction data has a shorter mining cadence due to its high real-time requirements. This dynamic adjustment mechanism has shown obvious advantages in implementation. When users suddenly share hiking-related content intensively on social platforms, travel preference data mining in the corresponding period can capture this change in a timely manner, while traditional fixed-period mining methods may miss such demand shifts, resulting in delayed recommendations. By matching data characteristics with mining frequency, this solution ensures data freshness while avoiding ineffective calculations, significantly improving recommendation efficiency.

[0018] 3. When merging and analyzing behavioral sample combinations, the system first performs a structural consistency check, then combines it with a content feature assessment to ultimately generate a quantitative recommendation match value. In practice, this two-dimensional analysis can effectively avoid "false matches"—for example, for two users who both label themselves as "favoring islands," the system will further calculate the match based on detailed dimensions such as whether they prefer independent travel and whether their dining consumption intervals overlap, rather than relying solely on a single label recommendation. This scientific and quantitative evaluation method significantly improves recommendation accuracy and reduces the cost of ineffective communication.

[0019] 4. The post-recommendation feedback optimization mechanism creates a significant closed-loop advantage in implementation: when a user rejects a recommendation and notes "significant differences in consumption habits," the system automatically traces the mining weight of consumption data and increases the evaluation weight of consumption dimensions such as dining and accommodation in subsequent calculations; if a user accepts the "hiking theme" recommendation multiple times, the system strengthens the matching weight of "activity type" in the travel preference data; this dynamic adjustment allows the system to continuously learn user preferences. The longer the user is used, the higher the recommendation accuracy, and the user satisfaction with the recommendation results is significantly higher than that of the traditional recommendation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flowchart of an embodiment of the present invention; Figure 2 A simplified flowchart of an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0022] Example 1: Please refer to Figure 1 This embodiment provides a personalized travel recommendation method based on behavioral data mining, which includes the following steps: S1. Collect multiple sets of user behavior source data in travel scenarios. Each set of behavior source data includes user basic information data, travel preference data, past interaction data, and consumption habit data; S2. Determine different data mining time interval information based on the above behavioral source data; S3, obtaining multiple batches of behavior sample combinations according to different data mining time interval information; S4. Combine and analyze multiple batches of behavioral sample combinations to obtain a recommended matching degree evaluation value; S5. Calculate the recommended judgment value based on the recommended matching degree evaluation value; S6. Compare the recommended judgment value with the pre-set limit value; If the recommended judgment value is greater than the pre-set threshold value, the information of travel companions who meet the conditions will be sent to the device used by the user; If the recommended judgment value is less than the preset limit value, the behavior source data is reselected and the above steps are repeated.

[0023] The steps for collecting multiple sets of user behavior source data in a travel scenario include: S101. Determine multiple user data acquisition channels in tourism scenarios, including tourism social networking platforms, tourism reservation software, tourism review websites, and offline tourism service record systems; S102. Collect multiple sets of original behavior data through multiple user data acquisition channels; S103, analyzing the identity characteristics contained in each set of raw behavior data, the identity characteristics including the user's name, age, gender, place of birth, and identity identification code, and classifying the corresponding raw behavior data into user basic information data based on these identity characteristics; S104, analyzing the preference characteristics contained in each set of raw behavior data, where the preference characteristics include the user's preferred travel location type, travel activity type, travel duration, travel season, and accommodation selection criteria, and classifying the corresponding raw behavior data into travel preference data based on these preference characteristics; S105. Analyze the interaction characteristics contained in each set of raw behavior data. The interaction characteristics include the number of times the user communicates with other users on the travel-related platform, the topics of the communication content, the number of trips taken together, and the scores of mutual evaluations. Based on these interaction characteristics, the corresponding raw behavior data is divided into past interaction data. S106. Analyze the consumption characteristics contained in each set of raw behavior data. The consumption characteristics cover the user's dining spending standards, transportation cost range, types of attraction ticket consumption, and shopping consumption tendencies during the travel process. Based on these consumption characteristics, the corresponding raw behavior data is divided into consumption habit data.

[0024] When collecting multiple sets of original behavioral data through various user data acquisition channels, it also includes data legitimacy verification of the original behavioral data obtained through each channel. The verification content includes the authorization status of the data source and the compliance of the data content; only the original behavioral data that has passed the legitimacy verification is retained for subsequent processing.

[0025] This embodiment corresponds to technical content related to data collection. Suppose a user frequently shares his travel experiences in coastal cities on a travel social networking platform, expressing his love for diving activities; he has booked many long-distance travel tickets and high-end hotels on a travel reservation app; he has given high scores to some historical and cultural attractions on a travel review webpage and has exchanged travel strategies with other tourists many times; and the offline travel service record system shows that he spends a lot of money on food and beverages during his travels and prefers to taste local specialties.

[0026] During the data collection process, we first identify various user data acquisition channels, including travel social networking platforms, travel reservation software, travel evaluation websites, and offline travel service record systems. When collecting original behavioral data through these channels, we conduct data legitimacy verification on the original behavioral data obtained through each channel to check whether the data source has user authorization and whether the data content complies with relevant regulations. Only data that has passed the verification is retained.

[0027] Afterwards, the collected raw behavioral data is classified; through the travel social network platform, the original behavioral data such as travel updates and comments posted by the user are collected, and the user's basic information data is determined based on the personal information mentioned by the user (such as age, gender, etc.); the travel preference data is determined based on the type of travel destinations (coastal cities) and travel activities (diving) shared by the user; and the past interaction data is determined based on the number of communications with other users and the topics of the communications.

[0028] Among them, the user's basic information data: extracted from the personal information of the travel social platform and the real-name authentication information of the reservation software, including the name Ms. Zhang, age 28, gender female, birthplace Hangzhou, and the identity identification code is the unique user ID generated by the system, namely U20230512001.

[0029] Travel preference data: extracted from social platform updates, reservation records, and review webpages; social platform updates mentioned that they loved beach sunsets and considered diving a must-do activity; reservation records showed that they had booked air tickets to coastal cities such as Sanya and Qingdao every summer for the past three years, choosing four-star and above hotels; review webpages gave high scores to diving clubs and seaside restaurants; these data covered the following content: preferred location type was coastal cities, activity type was diving and watching sunsets, trip duration was 5 to 7 days each time, travel season was summer, and accommodation standard was four-star and above hotels.

[0030] Past interaction data: extracted from social platform interaction records and evaluation web page replies; social platform interaction records show discussions with 10 users on island strategies, with an average of 2 to 3 exchanges per week, focusing on diving equipment and itinerary planning; evaluation web page replies show mutual travel experience reviews with 5 users, with ratings of 4 stars or above; these data include interaction objects of 10 island travel enthusiasts, interaction frequency of 2 to 3 times per week, interaction content themes of diving and itineraries, and cooperation times of 1 trip with 2 users, with mutual ratings of 4 to 5 stars.

[0031] Consumption habit data: extracted from offline records and reservation software consumption; offline records show that the average per capita consumption of dining is 200 to 300 yuan, with a preference for seafood restaurants, and transportation choices include economy class flights and local car rentals; reservation software consumption shows that hotel expenses are 800 to 1,200 yuan per night; these data cover dining standards of 200 to 300 yuan per person, transportation costs of an average air ticket price of 1,500 yuan per trip, and car rentals of 300 yuan per day. Attraction consumption prioritizes scenic spots with diving projects and has a ticket budget of 200 to 500 yuan. Shopping tends to purchase diving related products and local specialties with a single shopping budget of 500 to 1,000 yuan.

[0032] Through travel reservation software, we obtain the user's booking records and determine travel preference data based on the length of the travel time and accommodation selection criteria (high-end hotels); at the same time, we obtain some interaction information through the user's communication records with customer service and determine past interaction data.

[0033] Through the travel evaluation webpage, the user's evaluation content of the scenic spots, the interaction with other tourists, and the travel preference data and past interaction data are collected.

[0034] Through the offline travel service record system, the user's consumption records are obtained, and consumption habit data is determined based on dining cost standards, transportation cost range, etc.

[0035] In a specific implementation example, the step of determining different data mining time interval information based on the above-mentioned behavioral source data includes: S201, understanding the storage structure of the user in the user information database based on the user's basic information data; S202: Obtain update information of the key identity according to the storage structure, the update information including the number of updates, the size of the update change, and the update triggering condition; S203: Create an identity index and a user classification tag based on the update information, where the user classification tag includes an age tag, a gender tag, and a region tag; S204, recording the number of updates of the identity index and the user classification mark, and using the number of updates as the first type of mining time interval; S205: Acquire the user's travel record content based on the travel preference data, where the travel record content includes information on the type of travel location, travel time, and travel mode; S206, calculating the correlation between the tourist location type information, the travel time information, and the travel mode information, determining the correlation by counting the number and probability of different information appearing at the same time, and creating a preference-time correlation index based on the correlation; S207, calculating the stability of multiple preferences based on the preference-time association index, calculating the stability of preferences based on the length of time the same preference appears continuously and the magnitude of the change, and determining the second type of mining time interval for travel preference data based on the multiple preference stability conditions; S208: Obtain interaction record content based on past interaction data and create an interaction object-number association index; S209, analyzing the user's dependence level on the interaction relationship based on the interaction object-number association index, calculating the dependence level by multiplying the number of interactions by the interaction duration, and determining the third type of mining time interval based on the dependence level; S210, obtaining consumption record content based on consumption habit data; S211. Analyze the dynamic changes in consumption record content and calculate the dynamic changes based on the differences in consumption amounts and consumption types in adjacent time periods; S212, calculating the consumption-preference sensitivity based on the dynamic change, calculating the consumption-preference sensitivity by the ratio of the consumption change to the preference change, and determining the fourth type of mining time interval based on the consumption-preference sensitivity; S213: Summarize and merge the first type of mining time interval, the second type of mining time interval, the third type of mining time interval, and the fourth type of mining time interval. When merging, perform weighted averaging based on the importance of each time interval on the recommendation result to obtain different data mining time interval information.

[0036] This embodiment corresponds to the relevant technical content of determining data mining time interval information. For the user's basic information data, in the user information database, it is found that the user's age tag is updated every six months (because the user's age increases once a year, the system performs data update synchronization every six months). First, understand its storage structure, obtain the update information of key identity identifiers, create an identity identifier index and user classification tag, and then use six months as the first type of mining time interval.

[0037] For travel preference data, we analyzed the user's travel records and found that his preference for seaside tourism is relatively stable. He has taken a seaside trip every year in the past three years, and the types of travel activities are mainly concentrated in diving and beach leisure. First, we obtain the content of the user's travel records, calculate the correlation between information such as the type of travel location and travel time, create a preference-time correlation index, and then calculate the stability of the preference based on the length of time and change size of the continuous appearance of the same preference, and determine each year as the second type of mining time interval for travel preference data.

[0038] For past interaction data, we observed their interaction records with other users on tourism-related platforms and found that they interacted frequently with some users who frequently exchanged travel experiences, with an average of 3-4 interactions per month. Based on past interaction data, we obtained the content of the interaction records and created an interaction object-number association index. We calculated the user's dependence level on the interaction relationship by multiplying the number of interactions by the duration of the interaction, and determined each month as the third type of mining time interval.

[0039] For consumption habit data, we analyzed their consumption records and found that their dining consumption habits fluctuated with seasonal changes. For example, the consumption of seafood increased in summer. We obtained the consumption record content from the consumption habit data, analyzed the differences in consumption amounts and types in adjacent time periods, and obtained the dynamic changes. We calculated the consumption-preference sensitivity based on the ratio of consumption changes and preference changes, and determined each quarter as the fourth type of mining time interval.

[0040] Finally, the first type of mining time interval (half a year), the second type of mining time interval (one year), the third type of mining time interval (one month) and the fourth type of mining time interval (one quarter) are summarized and merged, and a weighted average is taken according to the importance of each time interval on the recommendation results (assuming that the influence weight of user basic information data is 0.2, the influence weight of travel preference data is 0.3, the influence weight of past interaction data is 0.3, and the influence weight of consumption habit data is 0.2), to obtain different data mining time interval information.

[0041] For example, the data mining time interval needs to be dynamically set according to the update frequency and stability of different data types to ensure a balance between data timeliness and mining efficiency.

[0042] Mining interval for basic user information data: By analyzing the storage structure of the user information database, we found that the identity identification code remains unchanged throughout life, age is updated once a year, and the system synchronizes basic information once every six months. By calculating the update frequency, age is updated once every 180 days, so the first type of mining interval is set to 6 months.

[0043] Travel preference data mining interval: extract travel records, including a trip to Sanya in July 2020, a trip to Qingdao in August 2021, and a trip to Xiamen in June 2022; calculate the correlation between location type, time, and method, where the location type is all seaside, the time is all summer, and the method is all flight plus car rental. The probability of the same features appearing at the same time is 100%, there is no change for three consecutive years, and the preference stability is 90% (the full score is 100%); the calculation formula for preference stability is: preference stability = (duration of continuous appearance of the same preference ÷ total observation time) × (1-preference change range) × 100%; in this example, the duration of continuous appearance of the same preference is 3 years, the total observation time is 3 years, and the preference change range is 0, so the preference stability = (3 ÷ 3) × (1-0) × 100% = 100%, but considering other potential factors, the actual setting is 90%; accordingly, the second type of mining interval is set to 12 months, that is, mining is carried out in the first month of summer every year.

[0044] Mining intervals for past interaction data: Construct interaction object and frequency indexes and calculate the dependency level. The number of interactions with the core interaction objects, i.e., three high-frequency users, is 50 times per year on average, and the duration is two years. Multiplying the two together gives 100, indicating a high dependency level. The dependency level calculation formula is: dependency level = number of interactions × duration of interaction. Therefore, the third type of mining interval is set to 30 days, i.e., mining is conducted once a month.

[0045] Consumption habit data mining interval: Analyze the dynamic changes in consumption. Catering consumption in 2022 increased by 15% compared with 2021 due to the increase in seafood prices and the unchanged accommodation standards. Calculate the sensitivity of consumption and preferences. The change in consumption is 15% and the change in preferences is 0%. The calculation formula for consumption-preference sensitivity is: consumption-preference sensitivity = change in consumption ÷ change in preference. Since the change in preference is 0, the result is meaningless. Therefore, the fourth type of mining interval is set to 90 days according to seasonal fluctuations, that is, mining once per quarter.

[0046] Finally, the four types of intervals are integrated and weighted according to the influence weight; the basic situation accounts for 10%, preferences account for 40%, interactions account for 30%, and consumption accounts for 20%; the calculation formula for the total mining interval is: total mining interval = (first type of mining interval × basic situation weight) + (second type of mining interval × preference weight) + (third type of mining interval × interaction weight) + (fourth type of mining interval × consumption weight); substitute the values ​​into the formula, 6 months multiplied by 10% plus 12 months multiplied by 40% plus 30 days multiplied by 30% plus 90 days multiplied by 20%, the result is 0.6 months plus 4.8 months plus 9 days plus 18 days, which is 32.4 days after conversion, that is, about 32 days for a full data mining.

[0047] In a specific implementation example, the step of obtaining multiple batches of behavior sample combinations according to different data mining time interval information includes: S301, determining multiple data mining tasks based on different data mining time interval information, each data mining task corresponding to a specific behavior data type and mining target; S302, setting mining configuration requirements for each data mining task, wherein the mining configuration requirements include data extraction scope, data cleaning rules, and data conversion standards; S303: Determine the source of original behavior data corresponding to the user's basic information data, travel preference data, past interaction data, and consumption habit data based on the mining configuration requirements; S304: Develop a mining sequence arrangement for each source of original behavior data, the mining sequence arrangement including the mining order, mining times, and mining interval time; S305: Perform multiple preliminary mining operations on each source of original behavior data according to a preset mining sequence, identify and correct abnormal data during the mining process, and determine abnormal data by the deviation from the average value of historical data to obtain mining information; S306. Merge the mining information corresponding to the user's basic information data, travel preference data, past interaction data, and consumption habit data. Use data alignment when merging to keep different types of data consistent in time and content, thereby forming multiple batches of behavioral sample combinations.

[0048] This embodiment corresponds to the relevant technical content of obtaining a combination of behavioral samples. Based on the data mining time interval information determined above, multiple data mining tasks are determined. For example, for travel preference data, data mining is performed one month before the peak tourist season (such as summer) each year. The mining goal is to obtain the latest changes in users' travel preferences.

[0049] Mining configuration requirements are set for this data mining work, including that the data extraction scope is data related to travel preferences on travel social networking platforms and travel reservation software in the past year, the data cleaning rules are to remove duplicate records and invalid data (such as incorrect travel location names), and the data conversion standard is to convert travel location names into a unified geocoding format.

[0050] According to the mining configuration requirements, the original behavioral data source corresponding to the user's basic information data, travel preference data, past interaction data and consumption habit data is determined. Here, it is determined to be the user's travel dynamic records on the travel social network platform and the booking records of the travel reservation software.

[0051] A mining sequence is developed for each source of raw behavioral data. First, data is extracted from the tourism social network platform once a week for a total of four times; then data is extracted from the travel reservation software once every two weeks for a total of two times.

[0052] Multiple preliminary mining operations are performed on each source of original behavioral data according to the preset mining sequence. During the mining process, abnormal data is judged by the deviation from the average value of historical data. For example, if it is found that the type of accommodation booked by the user in a certain travel reservation is too different from the previous preferences, that is, the price is much lower than the standard of high-end hotels that the user usually chooses, after verification, it is found that it is because the travel destination is special and there is a lack of hotels that meet its standards in the local area. This data will be corrected to a reasonable accommodation option, such as choosing a local characteristic homestay of the same level.

[0053] Finally, the mined information obtained from different sources is merged, and the data obtained from the tourism social network platform and travel reservation software are integrated in chronological order using data alignment to form a behavioral sample combination.

[0054] For example, based on a 32-day mining interval, data needs to be extracted and processed in batches to form a sample combination.

[0055] Determine the mining tasks: Execute 4 tasks every 32 days, targeting four types of data respectively.

[0056] Task 1 extracts fields such as user ID, age, and gender based on the basic situation, with the goal of updating the user's basic tags.

[0057] Task 2 focuses on travel preferences and extracts destination search and collection records from the past 32 days, with the goal of capturing changes in preferences.

[0058] Task 3 extracts chat records and mutual comments from the past 32 days based on past interactions, with the goal of updating the strength of the interactive relationship.

[0059] Task 4 focuses on consumption habits and extracts consumption bills from the past 32 days. The goal is to track consumption trends.

[0060] Set mining parameters: In terms of extraction scope, Task 1 extracts data for the past 180 days, as the basic information is updated every six months; Tasks 2 to 4 extract data for the past 32 days.

[0061] In terms of cleaning rules, Task 1 requires deleting duplicate ID records; Task 2 requires filtering out invalid search terms, such as "just take a look"; Task 3 requires removing advertising content; and Task 4 requires eliminating refund orders.

[0062] In terms of conversion standards, the place names will be unified into the format of city plus region, such as Yalong Bay, Sanya; the consumption amount will be unified into RMB.

[0063] Execute the mining sequence: In terms of sequence, first execute Task 1, which is the mining of basic data, and then execute Tasks 2 to 4, which are the mining of dynamic data.

[0064] In terms of frequency, Task 1 occurs once every 180 days, and Tasks 2 to 4 occur once every 32 days.

[0065] In terms of intervals, tasks 2 to 4 are performed in the order of preference, interaction, and consumption, one item per day, and a round is completed in 3 days.

[0066] Handling abnormal data: A search record for Harbin was found in Task 2, which conflicted with the Haibin preference. After verification, it was found to be an accidental touch by the user, so it was marked as an exception and deleted.

[0067] In Task 3, an interaction record is missing a timestamp. Based on the adjacent record times of 9:15 and 9:30 on June 10, 2023, it is completed to 9:22.

[0068] Merge samples: Align the four types of data by user ID and timestamp to form a sample combination; for example, the sample of user U20230512001 on June 10, 2023 includes basic information of a 28-year-old female, a hobby of searching for Qingdao diving, an interaction of discussing diving suits with user U20230401002, and a consumption of booking a Qingdao hotel for 800 yuan.

[0069] In a specific implementation example, the steps of merging and analyzing multiple batches of behavioral sample combinations to obtain a recommended matching evaluation value include: S401, obtaining the structural form and content of multiple batches of data from the combination of multiple batches of behavior samples; S402. Obtaining structure result content according to the structural forms of the multiple batches of data, the structure result content including normalized data result content, semi-normalized data result content, and non-normalized data result content; S403: Calculate a behavior consistency evaluation value based on the results of the normalized data, the results of the semi-normalized data, and the results of the non-normalized data. The behavior consistency evaluation value is obtained by counting the percentage of matches of the same behavior description in data of different structures. S404. Extracting content characteristics from the contents of the multiple batches of data. The content characteristics include information completeness, information matching, information timeliness, information accuracy, information relevance, and behavior consistency. S405, calculating the content feature evaluation value using a hierarchical analysis method, wherein the importance of each content feature is determined by constructing a judgment matrix in the hierarchical analysis method; S406. Calculate a recommended matching degree evaluation value based on the behavior consistency evaluation value and the content feature evaluation value. The recommended matching degree evaluation value is a weighted sum of the behavior consistency evaluation value and the content feature evaluation value calculated in a ratio of 3:7.

[0070] When using the hierarchical analysis method to calculate the content feature evaluation value, regularly collect users' satisfaction evaluation of the recommendation results; adjust the importance weight of each content feature in the judgment matrix based on the satisfaction evaluation; use the adjusted judgment matrix to recalculate the content feature evaluation value to optimize the accuracy of the recommendation matching degree evaluation value.

[0071] This embodiment corresponds to the relevant technical content of merging and analyzing to obtain the recommended matching degree evaluation value. From a group of behavioral sample combinations, the structural form of the data is analyzed, and it is found that the data of the travel social network platform is semi-normalized data, the data of the travel reservation software is normalized data, and the data of the travel review webpage is non-normalized data.

[0072] The behavioral consistency evaluation value is calculated by counting the proportion of matches of the same behavioral description in data of different structures. For example, when describing travel preferences, a user on a travel social network platform mentions "liking seaside cities." The booking records of the travel reservation software show that the user has booked multiple trips to seaside cities, and there are also a large number of positive reviews of seaside attractions on the travel review webpage. According to statistics, the number of matches of the same behavioral description is 8, and the total number of behavioral descriptions is 10. Therefore, the behavioral consistency evaluation value is 8 ÷ 10 × 100% = 80%. The calculation formula for the behavioral consistency evaluation value is: Behavior consistency evaluation value = (number of matches of the same behavioral description ÷ total number of behavioral descriptions) × 100%.

[0073] Content features are extracted from the content of this batch of data, including the degree of information completeness and information matching. A hierarchical analysis method is used to determine the importance of each content feature by constructing a judgment matrix. It is assumed that the weight of information completeness is 0.2, the weight of information matching is 0.3, the weight of information timeliness is 0.1, the weight of information accuracy is 0.2, the weight of information relevance is 0.1, and the weight of behavior consistency is 0.1.

[0074] Each content feature was scored with a full score of 10; in terms of information completeness, the data covered many aspects such as user basic information and preferences, scoring 8 points; in terms of information matching, the descriptions of user preferences from data from different sources were basically consistent, scoring 9 points; in terms of information timeliness, most of the data was within the past 32 days, scoring 7 points; in terms of information accuracy, after verification, the data was consistent with the actual situation, scoring 8 points; in terms of information relevance, the data all revolved around travel-related content, scoring 8 points; in terms of behavioral consistency, user behavior was uniform in different scenarios, scoring 8 points.

[0075] The calculation formula for the content feature evaluation value is: Content feature evaluation value = (information completeness score × information completeness weight) + (information matching score × information matching weight) + (information timeliness score × information timeliness weight) + (information accuracy score × information accuracy weight) + (information relevance score × information relevance weight) + (behavior consistency score × behavior consistency weight); Substituting the values ​​into the formula, we can obtain: (8×0.2) + (9×0.3) + (7×0.1) + (8×0.2) + (8×0.1) + (8×0.1) = 1.6 + 2.7 + 0.7 + 1.6 + 0.8 + 0.8 = 8.2.

[0076] In order to continuously optimize the evaluation results, it is necessary to regularly collect user satisfaction evaluations on the recommendation results; assuming that the user satisfaction with the previous recommendation results is 85%, adjust the importance weights of each content feature based on this satisfaction; increase the weight of information matching to 0.35, the weight of information relevance to 0.15, the weight of information completeness to 0.15, the weight of information accuracy to 0.15, the weight of information timeliness to 0.08, and the weight of behavioral consistency to 0.07; use the adjusted weights to recalculate the content feature evaluation value: (8×0.15)+(9×0.35)+(7×0.08)+(8×0.15)+(8×0.15)+(8×0.07)=1.2+3.15+0.56+1.2+1.2+0.56=7.87, so as to improve the accuracy of the recommendation matching evaluation value.

[0077] Finally, the weighted sum is calculated according to the ratio of 3:7 between the behavior consistency evaluation value and the content feature evaluation value to obtain the recommended matching degree evaluation value; the calculation formula for the recommended matching degree evaluation value is: recommended matching degree evaluation value = (behavior consistency evaluation value × 0.3) + (content feature evaluation value × 0.7); substituting the data into the formula, we can get: (80% × 0.3) + (8.2 × 0.7) = 0.24 + 5.74 = 5.98.

[0078] In a specific implementation example, the step of calculating a recommendation judgment value based on a recommendation matching degree evaluation value includes: S501, determining the importance of user basic information data, travel preference data, past interaction data, and consumption habit data in the travel pair recommendation process, wherein the importance of the data is determined by adjusting expert scores and user feedback data, wherein the user basic information data accounts for 20%, the travel preference data accounts for 35%, the past interaction data accounts for 25%, and the consumption habit data accounts for 20%; S502: Counting the actual amount of data corresponding to the user's basic information data, travel preference data, past interaction data, and consumption habit data, and counting the actual amount of data based on the number of data records; S503. Calculate the matching quality evaluation value corresponding to the user's basic information data, travel preference data, past interaction data, and consumption habit data based on the recommended matching evaluation value, the data importance ratio, and the actual data quantity. The matching quality evaluation value is the product of the recommended matching evaluation value, the data importance ratio, and the actual data quantity. S504. Count the total verified data volume corresponding to the user's basic information data, travel preference data, past interaction data, and consumption habit data. The total verified data volume is the sum of the number of data items that have been verified to be accurate in each type of data. S505 , calculating a recommended judgment value based on the total verification data volume and the matching quality evaluation value. The recommended judgment value is the ratio of the matching quality evaluation value to the total verification data volume.

[0079] This embodiment corresponds to the relevant technical content of calculating the recommendation judgment value. The recommendation judgment value is calculated to ultimately determine whether to recommend a travel package to the user. This value comprehensively considers the importance, actual quantity and matching quality of various types of data. First, the importance ratio of user basic information data, travel preference data, past interaction data and consumption habit data in the travel package recommendation work is determined; among them, user basic information data accounts for 20%, travel preference data accounts for 35%, past interaction data accounts for 25%, and consumption habit data accounts for 20%. These ratios are determined by adjusting expert scores and a large amount of user feedback data, which can better reflect the impact of various types of data on the recommendation results.

[0080] The actual data quantity corresponding to these four types of data is counted. Through statistics on the number of data records, there are 100 items of user basic information data, 200 items of travel preference data, 150 items of past interaction data, and 120 items of consumption habit data.

[0081] Based on the recommended matching degree evaluation value, the proportion of data importance and the actual number of data, the matching quality evaluation value corresponding to each type of data is calculated; the calculation formula for the matching quality evaluation value is: matching quality evaluation value of a certain type of data = recommended matching degree evaluation value × proportion of importance of this type of data × actual number of this type of data.

[0082] User basic information data matching quality assessment value = 5.98 × 20% × 100 = 5.98 × 0.2 × 100 = 119.6; Travel preference data matching quality assessment value = 5.98 × 35% × 200 = 5.98 × 0.35 × 200 = 418.6; Past interaction data matching quality assessment value = 5.98 × 25% × 150 = 5.98 × 0.25 × 150 = 224.25; The consumption habit data matching quality assessment value = 5.98×20%×120=5.98×0.2×120=143.52.

[0083] The total amount of verified data corresponding to these four types of data is counted, that is, the total number of data items that have been verified to be accurate in each type of data, which is 500; the recommended judgment value is calculated based on the total amount of verified data and the matching quality assessment value. The calculation formula for the recommended judgment value is: Recommended judgment value = (user basic information data matching quality assessment value + travel preference data matching quality assessment value + past interaction data matching quality assessment value + consumption habit data matching quality assessment value) ÷ total amount of verified data; substituting the data into the calculation, we can get: (119.6+418.6+224.25+143.52) ÷ 500 = 895.97 ÷ 500 = 1.79194.

[0084] In a specific implementation example, after sending qualified travel companion information to the device used by the user, it also includes receiving user feedback data on the recommended travel companion information, the feedback data including acceptance of the recommendation, rejection of the recommendation and the reason for rejection; classifying and organizing the feedback data, and establishing an association between the feedback data and the data of each link in the corresponding recommendation process; analyzing the influencing factors of recommendation success or failure based on the association results; and adjusting the parameter settings of subsequent data mining and the calculation method of recommendation matching degree evaluation based on the influencing factors.

[0085] The process of reselecting behavioral source data includes analyzing the specific data types in the previously selected behavioral source data that cause the recommended judgment values ​​to fail to meet the standards; expanding the data acquisition scope or improving the data screening standards for specific data types; preprocessing the newly acquired behavioral source data, including removing duplicate data and supplementing missing data; and performing data mining and recommendation judgment on the preprocessed behavioral source data according to the original steps.

[0086] After the recommendation judgment value is calculated in this embodiment, it is compared with a pre-set threshold value to determine whether to recommend and subsequent actions. Assuming the pre-set threshold value is 1.5, since the calculated recommendation judgment value 1.79194 is greater than 1.5, information about travel companions that meet the requirements is sent to the user's device. This information includes the travel companions' basic information, travel preferences, past interaction records, and consumption habits, so that the user can fully understand potential partners.

[0087] After receiving the recommendation information, the user will give feedback to the recommended travel companion; the feedback data may include acceptance of the recommendation, rejection of the recommendation and the reason for rejection, etc.; for example, the user refuses the recommendation because the other party's consumption habits are too different from his own, and he prefers high-end accommodation, while the other party prefers economy hotels.

[0088] The feedback data was classified and organized, and associated with the data of each link in the corresponding recommendation process; through analysis, it was found that although consumption habit data accounted for 20% in the recommendation matching evaluation, the user's sensitivity to consumption habits was high, which was a key factor affecting his decision-making, which led to the failure of this recommendation; based on this influencing factor, the parameter settings of subsequent data mining were adjusted, such as increasing the proportion of consumption habit data in the data importance to 25%, and at the same time adjusting the calculation method of the recommendation matching degree evaluation, adjusting the ratio of the behavior consistency evaluation value to the content feature evaluation value to 2:8, and increasing the weight of consumption habit matching in the content feature evaluation.

[0089] If the recommendation judgment value is lower than the pre-set limit value, it means that the current behavioral source data is insufficient to generate appropriate recommendations. At this time, it is necessary to analyze the specific data type in the previously selected behavioral source data that causes the recommendation judgment value to fall short of the standard. It is assumed that there is insufficient past interaction data to accurately assess the user's interaction preferences.

[0090] Expand the scope of data acquisition for this type of data and collect the user's interaction records from more travel social platforms, such as travel forums, travel enthusiasts groups, etc.; or improve data screening standards to only retain data with high interaction frequency and meaningful interaction content, such as interaction frequency of no less than 3 times a week, and interaction content centered around travel planning, attraction recommendations, etc.

[0091] Preprocess the newly acquired behavioral source data to remove duplicate interaction records, such as repeated postings of the same interactive content on different platforms; supplement missing information such as interaction time and interaction objects, such as determining the specific time of the interaction based on the time description in the interactive content; perform data mining and recommendation judgment on the preprocessed behavioral source data according to the previous steps, and recalculate the recommendation judgment value; for example, the actual number of newly acquired past interaction data is 200; its matching quality evaluation value = 5.98×25%×200=299, and the total matching quality evaluation value becomes 119.6+418.6+299+143.52=980.72. The new recommendation judgment value = 980.72÷500=1.96144, which is greater than the limit value of 1.5, and a suitable travel trip can be recommended to the user.

[0092] Example 2: Based on Example 1, this embodiment further provides a personalized travel recommendation system based on behavioral data mining, which specifically includes: The data collection module is used to collect multiple sets of behavioral source data of users in travel scenarios. Each set of behavioral source data includes basic information of users, travel preference data, past interaction data and consumption habit data; A time interval determination module, configured to determine different data mining time interval information based on the above-mentioned behavioral source data; A sample combination acquisition module is used to obtain multiple batches of behavior sample combinations according to different data mining time interval information; The analysis and evaluation module is used to combine and analyze multiple batches of behavioral samples to obtain a recommended matching degree evaluation value; A judgment value calculation module is used to calculate the recommendation judgment value based on the recommendation matching degree evaluation value; A recommendation judgment module is used to compare the recommended judgment value with a pre-set limit value; If the recommended judgment value is greater than the pre-set threshold value, the information of travel companions who meet the conditions will be sent to the device used by the user; If the recommended judgment value is less than the preset limit value, the behavior source data is reselected and the data collection module is triggered to perform related operations.

[0093] In a specific implementation example, the data collection module includes: a pathway determination unit, used to determine multiple pathways for obtaining user data in tourism scenarios, including tourism social networking platforms, tourism reservation software, tourism evaluation web pages, and offline tourism service record systems; A raw data collection unit, used to collect multiple sets of raw behavior data through multiple user data acquisition channels; The basic data segmentation unit is used to analyze the identity characteristics contained in each set of raw behavior data. The identity characteristics include the user's name, age, gender, birthplace, and identity identification code, and to segment the corresponding raw behavior data into user basic information data based on these identity characteristics. The preference data classification unit is used to analyze the preference characteristics contained in each set of raw behavior data. The preference characteristics include the user's preferred travel location type, travel activity type, travel duration, travel season, and accommodation selection criteria, and classify the corresponding raw behavior data into travel preference data based on these preference characteristics; The interactive data segmentation unit is used to analyze the interactive characteristics contained in each set of raw behavioral data. The interactive characteristics include the number of times a user communicates with other users on travel-related platforms, the topics of the communication content, the number of trips taken together, and the scores of mutual evaluations. Based on these interactive characteristics, the corresponding raw behavioral data is divided into past interactive data; The consumption data segmentation unit is used to analyze the consumption characteristics contained in each set of original behavior data. The consumption characteristics cover the user's dining spending standards, transportation cost range, types of attraction ticket consumption, and shopping consumption tendencies during the travel process, and based on these consumption characteristics, the corresponding original behavior data is divided into consumption habit data.

[0094] In a specific implementation example, the time interval determination module includes: A structure understanding unit is used to understand the storage structure of the user in the user information database based on the user's basic information data; An update information acquisition unit, configured to obtain update information of the key identity identifier according to the storage structure, the update information including the number of updates, the size of the update change, and the update triggering condition; An index mark creation unit, configured to create an identity index and a user classification mark based on the update information, wherein the user classification mark includes an age mark, a gender mark, and a region mark; A first time interval unit is used to record the number of updates of the identity index and the user classification mark, and use the number of updates as the first type of mining time interval; A travel record acquisition unit is used to acquire the user's travel record content based on the travel preference data. The travel record content includes information on the type of travel location, travel time, and travel mode; A correlation index creation unit is used to calculate the correlation between the tourist location type information, travel time information, and travel mode information, determine the correlation by counting the number and probability of different information appearing at the same time, and create a preference-time correlation index based on the correlation; The second time interval unit is used to calculate the stability of multiple preferences based on the preference-time association index, calculate the stability of preferences by the length of time and change of the continuous appearance of the same preference, and determine the second type of mining time interval of travel preference data based on the multiple preference stability conditions; An interaction association unit, used to obtain interaction record content based on past interaction data and create an interaction object-frequency association index; A third time interval unit is used to analyze the user's dependence level on the interaction relationship based on the interaction object-number association index, calculate the dependence level by multiplying the number of interactions by the interaction duration, and determine the third type of mining time interval based on the dependence level; A sensitivity calculation unit is used to obtain consumption-preference sensitivity based on consumption habit data, and calculate the consumption-preference sensitivity by the ratio of consumption change to preference change; The fourth time interval unit is used to determine the fourth type of mining time interval according to the consumption-preference sensitivity; The time interval merging unit is used to aggregate and merge the first type of mining time interval, the second type of mining time interval, the third type of mining time interval and the fourth type of mining time interval. When merging, a weighted average is performed according to the importance of each time interval on the recommendation result to obtain different data mining time interval information.

[0095] In a specific implementation example, the sample combination acquisition module includes: a task determination unit, configured to determine a plurality of data mining tasks according to different data mining time interval information, each data mining task corresponding to a specific behavior data type and mining target; Configuration setting unit, used to set mining configuration requirements for each data mining task, including data extraction scope, data cleaning rules and data conversion standards; A data source determination unit, configured to determine the source of original behavioral data corresponding to user basic information data, travel preference data, past interaction data, and consumption habit data based on mining configuration requirements; A sequence setting unit is used to set a mining sequence for each source of original behavioral data. The mining sequence includes the mining order, mining times, and mining interval time. A mining execution unit is used to perform multiple preliminary mining operations on each source of original behavioral data according to a preset mining sequence, identify and correct abnormal data during the mining process, and determine abnormal data by the deviation from the average value of historical data to obtain mining information; The sample merging unit is used to merge the mining information corresponding to the user's basic information data, travel preference data, past interaction data and consumption habit data. The data alignment method is used during the merging to make different types of data consistent in time and content, forming multiple batches of behavioral sample combinations.

[0096] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.

[0097] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the objectives of this embodiment based on actual needs.

[0098] The above is only a specific implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.

Claims

1. A personalized travel recommendation method based on behavioral data mining, characterized by: The recommendation method comprises the following steps: collecting multiple sets of behavioral source data of users in a travel scenario, each set of behavioral source data including basic information data of the user, travel preference data, past interaction data and consumption habit data; Determine different data mining time interval information based on behavioral source data; According to different data mining time interval information, multiple batches of behavior sample combinations are obtained; Combine and analyze multiple batches of behavioral samples to obtain a recommended matching evaluation value; Calculate the recommendation judgment value based on the recommendation matching degree evaluation value; Comparing the recommended judgment value with the pre-set limit value; If the recommendation judgment value is greater than or equal to the preset limit value, the information of the travel companions who meet the conditions is sent to the device used by the user; otherwise, the behavior source data is reselected and the above steps are repeated.

2. A personalized travel recommendation method based on behavioral data mining according to claim 1, characterized in that: The steps for collecting multiple sets of user behavior source data in a travel scenario include: Identify multiple user data acquisition methods in tourism scenarios; collect multiple sets of raw behavioral data through various user data acquisition methods; verify the legitimacy of the raw behavioral data obtained through each method; and retain only the raw behavioral data that has passed the legitimacy verification for subsequent processing; Analyze the identity information contained in each set of raw behavior data, and divide the corresponding raw behavior data into basic user information data; analyze the preference characteristics contained in each set of raw behavior data, and divide the corresponding raw behavior data into travel preference data; analyze the interaction characteristics contained in each set of raw behavior data, which include the number of times users communicate with other users on travel-related platforms, the topics of the communication content, the number of trips together, and the scores of mutual evaluations, and divide the corresponding raw behavior data into past interaction data; analyze the consumption characteristics contained in each set of raw behavior data, and divide the corresponding raw behavior data into consumption habit data.

3. A personalized travel recommendation method based on behavioral data mining according to claim 1, characterized in that: The steps of determining different data mining time interval information based on behavioral source data include: Based on the basic user data, the storage structure of the user information database is understood to obtain the update information of the key identity identifier; based on the update information, the identity identifier index and user classification tag are created, and the update number is recorded, which is used as the first type of mining time interval; The content of user travel records is obtained based on travel preference data, and the travel record content includes travel location type information, travel time information, and travel mode information; the correlation between travel location type information, travel time information, and travel mode information is calculated, and the correlation is determined by counting the number and probability of different information appearing at the same time, and a preference-time correlation index is created based on the correlation; the stability of multiple preferences is calculated based on the preference-time correlation index, and the preference stability is calculated by the length of time and change size of the continuous appearance of the same preference, and the second type of mining time interval of travel preference data is determined based on the stability of multiple preferences.

4. A personalized travel recommendation method based on behavioral data mining according to claim 3, characterized in that: The step of determining different data mining time interval information based on the behavioral source data also includes: Create an interaction object-number association index based on the interaction record content obtained from past interaction data, analyze the user's dependence level on the interaction relationship, calculate the dependence level by multiplying the number of interactions and the interaction duration, and determine the third type of mining time interval based on the dependence level; Obtain consumption record content based on consumption habit data; analyze the dynamic changes in consumption record content, and calculate the dynamic changes by the differences in consumption amounts and consumption types in adjacent time periods; calculate the consumption-preference sensitivity based on the dynamic changes, and calculate the consumption-preference sensitivity based on the ratio of consumption changes to preference changes, and determine the fourth type of mining time interval based on the consumption-preference sensitivity; The first to fourth categories of mining time intervals are aggregated and merged, and when merging, weighted average is performed according to the importance of each time interval in affecting the recommendation results to obtain different data mining time interval information.

5. A personalized travel recommendation method based on behavioral data mining according to claim 4, characterized in that: The importance of each time interval on the recommendation results is determined by the following method: Collect the recommendation success rate corresponding to each time interval in the historical recommendation results; Calculate the ratio of the recommendation success rate at each time interval to the average recommendation success rate; This ratio is used as the importance of each time interval on the recommendation result; When weighted averaging is performed according to importance, the weight value is equal to the ratio of the importance to the sum of all importances.

6. A personalized travel recommendation method based on behavioral data mining according to claim 1, characterized in that: The steps of obtaining multiple batches of behavior sample combinations according to different data mining time interval information include: Determine multiple data mining tasks based on different data mining time interval information, and set mining configuration parameters for each data mining task. The mining configuration parameters include data extraction scope, data cleaning rules, and data conversion standards. Obtain the original behavioral data source corresponding to the basic data based on the mining configuration parameters. The basic data includes user basic information data, travel preference data, historical interaction data, and consumption behavior data. Obtain a mining sequence based on each original behavioral data source. Perform high-frequency preliminary mining on each original behavioral data source according to the preset mining sequence. Identify and correct abnormal data during the mining process. The abnormal data is judged by the deviation from the historical data average to obtain mining information. Integrate the mining information corresponding to the basic data, and use data alignment during integration to keep different types of data consistent in time and content dimensions, forming multiple batches of behavioral sample combinations.

7. A personalized travel recommendation method based on behavioral data mining according to claim 1, characterized in that: The step of merging and analyzing the combination of multiple batches of behavioral samples to obtain a recommended matching degree evaluation value includes: obtaining the structural form of the multiple batches of data and the content of the multiple batches of data from the combination of multiple batches of behavioral samples; The structural result content is obtained based on the structural form of multiple batches of data; the behavior consistency evaluation value is calculated based on the normalized data result content, semi-normalized data result content, and non-normalized data result content, and the behavior consistency evaluation value is obtained by counting the proportion of matches of the same behavior description in different structural data; Extract content features from the content of multiple batches of data; calculate content feature evaluation values ​​using a hierarchical analysis method, in which the importance of each content feature is determined by constructing a judgment matrix; The recommended matching degree evaluation value is calculated based on the behavior consistency evaluation value and the content feature evaluation value. The recommended matching degree evaluation value is the weighted sum of the behavior consistency evaluation value and the content feature evaluation value calculated according to a preset ratio.

8. A personalized travel recommendation method based on behavioral data mining according to claim 6, characterized in that: The step of calculating the recommendation judgment value based on the recommendation matching degree evaluation value includes: determining the importance ratio of the basic data in the travel partner recommendation work, and adjusting the data importance ratio based on expert scores and user feedback data; Count the actual data quantity corresponding to the basic data, and count the actual data quantity based on the number of data records; The matching quality evaluation value corresponding to the basic data is calculated based on the recommended matching evaluation value, the data importance ratio, and the actual data quantity. The matching quality evaluation value is the product of the recommended matching evaluation value, the data importance ratio, and the actual data quantity. The total amount of verified data corresponding to the statistical basic data is the sum of the number of data items that have been verified to be accurate in each type of data; The recommended judgment value is calculated based on the total verification data volume and the matching quality evaluation value. The recommended judgment value is the ratio of the matching quality evaluation value to the total verification data volume.

9. A personalized travel recommendation method based on behavioral data mining according to claim 1, characterized in that: After sending the eligible travel companion information to the device used by the user, it also includes: Receive user feedback data on recommended travel companion information, including acceptance of recommendations, rejection of recommendations, and reasons for rejection; classify and organize the feedback data, and group them according to feedback type and feedback time; establish an association between the feedback data and the data of each link in the corresponding recommendation process, and the association method includes data identifier correspondence and time node correspondence; analyze the influencing factors of recommendation success or failure based on the association results, and the influencing factors include insufficient data completeness and matching algorithm deviation; adjust the parameter settings of subsequent data mining and the calculation method of recommendation matching degree evaluation based on the influencing factors, and the parameter setting adjustment includes adjusting the data extraction scope and data cleaning rules.

10. A personalized travel recommendation method based on behavioral data mining according to claim 1, characterized in that: The process of reselecting behavioral source data includes: Analyze the specific data types in the previously selected behavioral source data that cause the recommendation judgment values ​​to not meet the standards. The specific data types include user basic information data, travel preference data, past interaction data, and consumption habit data; adjust the data acquisition scope or data screening criteria for specific data types. Adjusting the data acquisition scope includes increasing data acquisition channels, and adjusting the data screening criteria includes setting data integrity rules; preprocess the newly acquired behavioral source data. Preprocessing includes removing duplicate data and supplementing missing data; and perform data mining and recommendation judgment on the preprocessed behavioral source data according to the original steps.

Citation Information

Patent Citations

  • Regional document travel intelligent recommendation method, system and device and storage medium

    CN119311949A

  • Urban operation index monitoring method based on multi-source heterogeneity

    CN119494463A

  • Data acquisition method and device, electronic equipment and storage medium

    CN119718637A

  • Tourism destination recommendation method and system based on big data and computer equipment

    CN119807537A

  • Tourism information recommendation method and system based on big data platform

    CN120030233A