A method for recommending a travel companion based on behavior data mining

By collecting data from multiple channels and conducting differentiated analysis, user behavior is quantitatively analyzed, solving the problems of content cold start and data sparsity in the travel buddy recommendation system, and improving the accuracy and efficiency of personalized recommendations.

CN120561381BActive Publication Date: 2025-10-24HANGZHOU DOUBLE LEAF NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511062284.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-24
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing travel buddy recommendation systems, when faced with cold start content and sparse data, cannot accurately match users' deeper needs, resulting in a large discrepancy between the recommendation results and user expectations, thus affecting the travel social experience.

Method used

By collecting user behavior data from multiple channels, the data is divided into four categories: basic user information, travel preferences, past interactions, and consumption habits. A differentiated data mining time interval strategy is adopted to merge and analyze behavioral sample combinations, quantify recommendation matching degree, and continuously adjust the recommendation strategy through a feedback optimization mechanism.

Benefits of technology

It improves the alignment between recommended results and users' actual needs, reduces ineffective communication costs, increases recommendation accuracy and efficiency, and dynamically learns user preferences to improve satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561381B_ABST
    Figure CN120561381B_ABST
Patent Text Reader

Abstract

The application discloses a kind of tourism buddy personalized recommendation method based on behavior data mining, it is related to tourism service technical field, the recommendation method includes collecting the multiple groups of behavior source data of user in tourism scene;Different data mining time interval information is determined according to the behavior source data;According to different data mining time interval information, obtain multiple batches of behavior sample combination;Multiple batches of behavior sample combination are merged and analyzed, and recommended matching degree evaluation value is obtained;Its technical key points are: tourism social network platform is grabbed from user publishing tourism dynamic, from tourism reservation software extraction booking record, from tourism evaluation webpage collection score content, from offline service system summary consumption details, by identity, like and so on Dimension divides out user basic situation, travel like, past interaction and consumption habit four kinds of data;The scheme completely solves the recommendation deviation caused by information fragmentation, and the compatibility of recommended result and user real demand is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of tourism services, in particular to a personalized recommendation method for a travel companion based on behavior data mining. BACKGROUND

[0002] The holiday tourism market is always lively, and many people traveling alone hope to find like-minded companions on their journey; open a tourism social platform, enter the destination and travel theme in the search box, and in the user list popped up by the system, only the age, gender and personal signature of a few words can be seen; some signatures read "love travel", and some are marked "easy-going personality"; after flipping through dozens of pages, either the time cannot be synchronized or the consumer habits are too different after seeing the other party's sharing of high-end hotel check-in dynamics, so the user eventually gives up finding a companion and embarks on the journey alone.

[0003] After filling in the destination and time in the "find a companion" section of the tourism reservation software, the system quickly recommends several users; after in-depth communication, it is found that some prefer shopping rather than participating in the preset theme activities, some hope that the itinerary is full from morning to night, while the user wants to reserve time for free wandering, and some require hostel accommodation, which is contrary to the user's idea of staying in an economy hotel. After several unsuccessful matches, the user still cannot find a suitable companion.

[0004] When using the travel companion recommendation function for the first time, since there is no historical travel record, the system can only randomly push users, leaving the user in a dilemma of not knowing how to choose, and even if they choose, it may not be suitable, which is a typical content cold start problem.

[0005] At the same time, the user's behavior data is scattered in different platforms: equipment preferences shared on social platforms, "don't like to rush the itinerary" comments left on review websites, and "prefer to stay in a hostel" settings in reservation software, etc. Information has never been integrated and analyzed by the system; when looking for a companion with the same rhythm, the system cannot identify implicit needs such as "no more than 6 hours of activity per day" and "tendency to bring their own food" from multi-channel data, resulting in a large deviation between the recommended results and expectations; this scattered data and ineffective integration situation forms a data sparsity problem, making it difficult for many users to find a truly compatible travel companion even after spending a lot of time filtering, which seriously affects the improvement of the tourism social experience, therefore, there is an urgent need for a personalized recommendation method for a travel companion based on behavior data mining. SUMMARY

[0006] To achieve the above purpose, the present application is implemented by the following technical solutions:

[0007] A personalized recommendation method for a travel companion based on behavior data mining, the recommendation method comprising:

[0008] Collecting multiple sets of behavior source data of users in a travel scenario, each set of behavior source data containing user basic condition data, travel preference data, past interaction data and consumption habit data;

[0009] Determining different data mining time interval information according to the behavior source data;

[0010] Obtaining multiple batches of behavior sample combinations according to the different data mining time interval information;

[0011] Performing merging analysis on the multiple batches of behavior sample combinations to obtain a recommendation matching degree evaluation value;

[0012] Calculating a recommendation judgment value according to the recommendation matching degree evaluation value;

[0013] Comparing the recommendation judgment value with a pre-set limit value;

[0014] If the recommendation judgment value is greater than or equal to the pre-set limit value, sending travel companion information meeting the conditions to a device used by the user; otherwise, reselecting behavior source data and repeating the above steps.

[0015] Further, the step of collecting multiple sets of behavior source data of users in a travel scenario comprises:

[0016] Determining multiple user data acquisition approaches in the travel scenario, at least including a travel social network platform, a travel reservation software, a travel evaluation webpage and an offline travel service record system;

[0017] Collecting multiple sets of original behavior data through the multiple user data acquisition approaches;

[0018] Analyzing identity features contained in each set of original behavior data, the identity features covering a user's name, age, gender, birthplace and identity identification code, and dividing the corresponding original behavior data into user basic condition data according to the identity features; analyzing preference features contained in each set of original behavior data, the preference features covering a user's preferred travel location type, travel activity type, travel time length, travel season and accommodation selection criterion, and dividing the corresponding original behavior data into travel preference data according to the preference features; analyzing interaction features contained in each set of original behavior data, the interaction features covering a user's communication times with other users on a travel-related platform, communication content theme, travel times together and mutual evaluation scores, and dividing the corresponding original behavior data into past interaction data according to the interaction features; analyzing consumption features contained in each set of original behavior data, the consumption features covering a user's dining expense standard, transportation expense range, scenic spot ticket consumption category and shopping consumption tendency in a travel process, and dividing the corresponding original behavior data into consumption habit data according to the consumption features.

[0019] Further, the step of determining different data mining time interval information according to the above behavior source data comprises: understanding the storage structure of the user in the user information database according to the user basic situation data; obtaining the update situation information of the key identity according to the storage structure, the update situation information comprising the update times, the update change size and the update trigger condition; creating the identity index and the user classification mark according to the update situation information, the user classification mark comprising the age mark, the gender mark and the region mark; recording the update times of the identity index and the user classification mark, and taking the update times as the first type of mining time interval;

[0020] obtaining the travel record content according to the travel preference data, the travel record content comprising the travel location type information, the travel time information and the travel mode information; calculating the association between the travel location type information, the travel time information and the travel mode information, determining the association by counting the number of times and the possibility of different information appearing at the same time, and creating the preference-time association index according to the association; calculating a plurality of preference stability situations according to the preference-time association index, calculating the preference stability situation by the time length and the change size of the same preference continuously appearing, and determining the second type of mining time interval of the travel preference data according to the plurality of preference stability situations;

[0021] obtaining the interaction record content according to the past interaction data, and creating the interaction object-number association index; analyzing the dependence level of the user on the interaction relationship according to the interaction object-number association index, calculating the dependence level by the product of the interaction times and the interaction duration, and determining the third type of mining time interval according to the dependence level;

[0022] obtaining the consumption record content according to the consumption habit data; analyzing the dynamic change situation of the consumption record content, calculating the dynamic change situation by the difference between the consumption amount and the consumption type of adjacent time periods; calculating the consumption-preference sensitivity degree according to the dynamic change situation, calculating the consumption-preference sensitivity degree by the ratio of the consumption change size and the preference change size, and determining the fourth type of mining time interval according to the consumption-preference sensitivity degree;

[0023] The first type of mining time interval, the second type of mining time interval, the third type of mining time interval and the fourth type of mining time interval are summarized and combined, and when combined, weighted average is performed according to the importance of the influence of each time interval on the recommendation result, to obtain different data mining time interval information.

[0024] Further, the step of obtaining a plurality of batches of behavior sample combinations according to different data mining time interval information comprises:

[0025] Determine a plurality of data mining jobs according to different data mining time interval information, each data mining job corresponding to a specific behavior data type and a mining target; set mining configuration requirements for each data mining job, the mining configuration requirements including data extraction range, data cleaning rules and data conversion standards; determine the original behavior data sources corresponding to the user basic situation data, travel preference data, past interaction data and consumption habit data according to the mining configuration requirements; develop a mining sequence arrangement for each original behavior data source, the mining sequence arrangement including mining sequence, mining times and mining interval time; perform multiple preliminary mining on each original behavior data source according to the preset mining sequence arrangement, identify and correct abnormal data in the mining process, judge the abnormal data by the deviation size from the historical data average value, and obtain mining information; merge the mining information corresponding to the user basic situation data, travel preference data, past interaction data and consumption habit data, and adopt data alignment when merging to make different types of data consistent in time and content, and form a plurality of batch behavior sample combinations.

[0026] Further, the step of performing merging analysis on the plurality of batch behavior sample combinations to obtain a recommendation matching degree evaluation value, comprising:

[0027] Obtaining a structure result content according to the structure form of the plurality of batch data, the structure result content including standardized data result content, semi-standardized data result content and non-standardized data result content;

[0028] Calculating a behavior consistency evaluation value according to the standardized data result content, the semi-standardized data result content and the non-standardized data result content, and obtaining the behavior consistency evaluation value by counting the matching number proportion of the same behavior description in different structure data;

[0029] Extracting content feature content from the content of the plurality of batch data, the content feature content including information completeness degree, information matching degree, information timeliness degree, information accuracy degree, information relevance degree and behavior consistency degree;

[0030] Calculating a content feature evaluation value by using an analytic hierarchy process, and determining the importance of each content feature by constructing a judgment matrix in the analytic hierarchy process;

[0031] Calculating a recommendation matching degree evaluation value according to the behavior consistency evaluation value and the content feature evaluation value, the recommendation matching degree evaluation value being a weighted sum of the behavior consistency evaluation value and the content feature evaluation value according to a 3:7 ratio.

[0032] Further, the step of calculating a recommendation judgment value according to the recommendation matching degree evaluation value, comprising:

[0033] Determine the data importance proportion of the user basic situation data, travel preference data, past interaction data and consumption habit data in the travel companion recommendation work, and the data importance proportion is determined through expert scoring combined with user feedback data adjustment, wherein the user basic situation data accounts for 20%, the travel preference data accounts for 35%, the past interaction data accounts for 25%, and the consumption habit data accounts for 20%;

[0034] Statistically count the actual data quantity corresponding to the user basic situation data, travel preference data, past interaction data and consumption habit data, and the actual data quantity is counted by the number of data records;

[0035] According to the recommendation matching degree evaluation value, the data importance proportion and the actual data quantity, calculate the matching quality evaluation value corresponding to the user basic situation data, travel preference data, past interaction data and consumption habit data, and the matching quality evaluation value is the product of the recommendation matching degree evaluation value, the data importance proportion and the actual data quantity;

[0036] Statistically count the total verification data quantity corresponding to the user basic situation data, travel preference data, past interaction data and consumption habit data, and the total verification data quantity is the sum of the number of verified accurate data in each type of data;

[0037] According to the total verification data quantity and the matching quality evaluation value, calculate the recommendation judgment value, and the recommendation judgment value is the ratio of the matching quality evaluation value to the total verification data quantity.

[0038] Further, after sending the qualified travel companion information to the device used by the user, it further includes:

[0039] Receive user feedback data on the recommended travel companion information, and the feedback data includes accepting the recommendation, rejecting the recommendation and the reason for rejecting; classify and arrange the feedback data, establish the association between the feedback data and the data in each link in the corresponding recommendation process; analyze the influencing factors of successful or failed recommendation according to the association result; adjust the parameter setting of subsequent data mining and the calculation method of recommendation matching degree evaluation according to the influencing factors.

[0040] Further, in the process of reselecting the behavior source data, it includes:

[0041] Analyze the specific data types in the previously selected behavior source data that cause the recommendation judgment value to be substandard; expand the data acquisition range or improve the data screening standard for the specific data types; pre-process the newly acquired behavior source data, which includes removing duplicate data and supplementing missing data; and perform data mining and recommendation judgment on the pre-processed behavior source data according to the original steps.

[0042] Further, when collecting multiple groups of original behavior data through multiple user data acquisition approaches, it further includes:

[0043] The original behavior data obtained by each path is subjected to data legality verification, and the verification content includes the authorization of the data source and the compliance of the data content.

[0044] Further, when the analytic hierarchy process is used to calculate the content feature evaluation value, the following is further included:

[0045] Periodically collect user satisfaction evaluation of the recommendation result, adjust the importance weight of each content feature in the judgment matrix according to the satisfaction evaluation, and recalculate the content feature evaluation value using the adjusted judgment matrix to optimize the accuracy of the recommendation matching degree evaluation value.

[0046] The application provides a travel companion personalized recommendation method based on behavior data mining, which has the following beneficial effects:

[0047] 1. In the data collection stage, the system according to the technical scheme extracts the travel dynamics and interactive comments published by the user from the travel social networking platform, extracts the booking records from the travel reservation software, collects the rating content from the travel evaluation webpage, and summarizes the consumption details from the offline service system. Four types of data are divided according to the dimensions of identity characteristics and preference characteristics, including user basic information, travel preferences, past interactions and consumption habits. This multi-channel and multi-dimensional data integration method can fully outline the user's deep needs, such as frequent booking of seaside hotels with diving projects and high-frequency mention of seafood restaurants in food evaluation data. Not only can the user's destination preference be determined, but also the activity tendency and consumption level can be accurately positioned. Compared with the traditional recommendation which only relies on personal signature, this scheme completely solves the recommendation deviation caused by information fragmentation, and significantly improves the compatibility of the recommendation result with the user's real needs.

[0048] 2. When determining the data mining time interval, the system will develop differentiated strategies for different data types: user basic information data has a low update frequency, so a longer mining period is used; travel preference data may fluctuate with the seasons, so the mining frequency is adjusted quarterly; past interaction data requires high real-time performance, so a shorter mining rhythm is set. This dynamic adjustment mechanism has obvious advantages in implementation: when a user suddenly shares a lot of hiking-related content on a social platform, the corresponding travel preference data mining can capture this change in a timely manner, while the traditional fixed-period mining method may miss such demand changes, resulting in delayed recommendations. By matching the data characteristics and mining frequency, this scheme not only ensures the freshness of the data, but also avoids unnecessary calculations, significantly improving the recommendation efficiency.

[0049] 3. When merging and analyzing behavioral sample combinations, the system first performs a structural consistency check, then combines it with a content feature assessment to ultimately generate a quantitative recommendation match value. In practice, this two-dimensional analysis can effectively avoid "false matches"—for example, for two users who both label themselves as "favoring islands," the system will further calculate the match based on detailed dimensions such as whether they prefer independent travel and whether their dining consumption intervals overlap, rather than relying solely on a single label recommendation. This scientific and quantitative evaluation method significantly improves recommendation accuracy and reduces the cost of ineffective communication.

[0050] 4. The post-recommendation feedback optimization mechanism creates a significant closed-loop advantage in implementation: when a user rejects a recommendation and notes "significant differences in consumption habits," the system automatically traces the mining weight of consumption data and increases the evaluation weight of consumption dimensions such as dining and accommodation in subsequent calculations; if a user accepts the "hiking theme" recommendation multiple times, the system strengthens the matching weight of "activity type" in the travel preference data; this dynamic adjustment allows the system to continuously learn user preferences. The longer the user is used, the higher the recommendation accuracy, and the user satisfaction with the recommendation results is significantly higher than that of the traditional recommendation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flowchart of an embodiment of the present invention;

[0052] Figure 2 A simplified flowchart of an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0054] Example 1: Please refer to Figure 1 This embodiment provides a personalized travel recommendation method based on behavioral data mining, which includes the following steps:

[0055] S1. Collect multiple sets of user behavior source data in travel scenarios. Each set of behavior source data includes user basic information data, travel preference data, past interaction data, and consumption habit data;

[0056] S2. Determine different data mining time interval information based on the above behavioral source data;

[0057] S3, obtaining multiple batches of behavior sample combinations according to different data mining time interval information;

[0058] S4, performing a combined analysis on the multiple batches of behavior sample groups to obtain a recommended matching degree evaluation value;

[0059] S5, calculating a recommended judgment value according to the recommended matching degree evaluation value;

[0060] S6, comparing the recommended judgment value with a pre-set boundary value;

[0061] If the recommended judgment value is greater than the pre-set boundary value, sending the qualified travel companion information to the device used by the user;

[0062] If the recommended judgment value is less than the pre-set boundary value, re-selecting the behavior source data and repeating the above steps.

[0063] The step of collecting multiple sets of behavior source data of the user in the travel scene comprises:

[0064] S101, determining multiple user data acquisition approaches in the travel scene, which include a travel social network platform, a travel reservation software, a travel evaluation webpage and an offline travel service record system;

[0065] S102, collecting multiple sets of original behavior data through the multiple user data acquisition approaches;

[0066] S103, analyzing identity features contained in each set of original behavior data, the identity features covering a name, an age, a gender, a birthplace and an identity identification code of the user, and dividing corresponding original behavior data into user basic situation data according to the identity features;

[0067] S104, analyzing preference features contained in each set of original behavior data, the preference features covering a preferred travel location type, a preferred travel activity type, a travel time length, a travel season and a lodging selection criterion of the user, and dividing corresponding original behavior data into travel preference data according to the preference features;

[0068] S105, analyzing interaction features contained in each set of original behavior data, the interaction features covering a communication frequency, a communication content theme, a number of times of traveling together and a mutual evaluation score of the user with other users on a travel-related platform, and dividing corresponding original behavior data into past interaction data according to the interaction features;

[0069] S106, analyzing consumption features contained in each set of original behavior data, the consumption features covering a dining expense standard, a transportation expense range, a type of scenic spot ticket consumption and a shopping consumption tendency of the user in a travel process, and dividing corresponding original behavior data into consumption habit data according to the consumption features.

[0070] When collecting multiple sets of original behavior data through multiple user data acquisition approaches, data legality verification is also performed on the original behavior data acquired by each approach, and the verification content includes the authorization of the data source and the compliance of the data content; only the original behavior data that passes the legality verification is retained for subsequent processing.

[0071] The present embodiment corresponds to the technical content related to data collection. Assume that a user frequently shares his / her travel experiences in a coastal city on a tourism social networking platform, expressing his / her love for diving activities; has booked multiple long-distance travel tickets and high-end hotels on a travel reservation software; has given high ratings to some historical and cultural scenic spots on a travel review webpage and has exchanged travel strategies with other tourists multiple times; and shows high dining consumption during travel and prefers to taste local specialties in the offline tourism service record system.

[0072] During data collection, first, determine the multiple user data acquisition approaches of the tourism social networking platform, the travel reservation software, the travel review webpage, and the offline tourism service record system; when collecting original behavior data through these approaches, perform data legality verification on the original behavior data acquired by each approach, check whether the data source is authorized by the user and whether the data content complies with relevant regulations, and only retain the data that passes the verification.

[0073] Then, classify the collected original behavior data; through the tourism social networking platform, collect the original behavior data such as the user's travel updates and comments, determine the user's basic situation data according to the personal information (such as age, gender, etc.) mentioned by the user, determine the travel preference data according to the types of travel locations (coastal cities) and travel activities (diving) shared by the user, and determine the past interaction data according to the number of exchanges with other users and the topics of the exchanges.

[0074] Among them, the user's basic situation data is extracted from the personal information on the tourism social platform and the real-name authentication information on the reservation software, including the user's name, age, gender, birthplace, and unique user ID generated by the system, i.e., U20230512001.

[0075] The travel preference data is extracted from the social platform updates, reservation records, and review webpages; the social platform updates mention that the user loves sunsets by the sea and diving is a must-play activity; the reservation records show that the user has booked tickets to coastal cities such as Sanya and Qingdao every summer for the past three years and has chosen four-star and above hotels; the review webpage gives high ratings to diving clubs and seaside restaurants; these data cover the following content: preferred location type is coastal city, activity type is diving and watching sunset, travel duration is 5 to 7 days each time, travel season is summer, and accommodation standard is four-star and above hotel.

[0076] Past interaction data: extracted from social platform interaction records and review webpage replies; social platform interaction records show that the user discussed island strategies with 10 users, with an average of 2-3 exchanges per week, focusing on diving equipment and trip planning; review webpage replies show that the user reviewed travel experiences with 5 users, with ratings above 4 stars; these data include 10 island tourism enthusiasts as interaction objects, 2-3 exchanges per week as interaction frequency, diving and trip as interaction content theme, 1 trip with 2 users as cooperation times, and 4-5 stars as mutual evaluation score.

[0077] Consumption habit data: extracted from offline records and reservation software consumption; offline records show that the average consumption per person in a restaurant is 200-300 yuan, with a preference for seafood restaurants and transportation options of economy class and local car rental; reservation software consumption shows that the hotel costs 800-1200 yuan per night; these data cover a restaurant standard of 200-300 yuan per person, transportation costs of 1500 yuan per flight and 300 yuan per day of car rental, and a preference for scenic spots with diving projects with a ticket budget of 200-500 yuan, and a shopping tendency of purchasing diving peripherals and local specialties with a single shopping budget of 500-1000 yuan.

[0078] Through the tourism reservation software, the user's reservation records are obtained, and the travel preference data is determined according to the length of the travel time, the selection criteria of the accommodation (high-end hotel). At the same time, through the communication records with the customer service, some interaction information is obtained to determine the past interaction data.

[0079] Through the tourism evaluation webpage, the user's evaluation content of the scenic spots and the interaction with other tourists are collected to determine the travel preference data and the past interaction data.

[0080] Through the offline tourism service record system, the user's consumption records are obtained, and the consumption habit data is determined according to the dining expense standard, transportation cost range, etc.

[0081] In a specific implementation example, the step of determining different data mining time interval information according to the above behavior source data includes:

[0082] S201, according to the user basic situation data, understand the storage structure of the user in the user information database;

[0083] S202, obtain the update situation information of the key identity according to the storage structure, the update situation information including the update times, the update change size and the update trigger condition;

[0084] S203, create an identity index and a user classification mark according to the update situation information, the user classification mark including an age mark, a gender mark and a region mark;

[0085] S204, record the update times of the identity index and the user classification mark, and take the update times as the first type of mining time interval;

[0086] S205, obtain the user travel record content according to the travel preference data, and the travel record content includes the travel location type information, the travel time information and the travel mode information;

[0087] S206, calculate the association between the travel location type information, the travel time information and the travel mode information, determine the association by counting the number of times and the possibility of different information appearing at the same time, and create a preference-time association index according to the association;

[0088] S207, calculate a plurality of preference stability conditions according to the preference-time association index, calculate the preference stability conditions by the length of time of continuous appearance of the same preference and the change size, and determine the second type of mining time interval of the travel preference data according to the plurality of preference stability conditions;

[0089] S208, obtain the interaction record content according to the past interaction data, and create an interaction object-number association index;

[0090] S209, analyze the dependence level of the user on the interaction relationship according to the interaction object-number association index, calculate the dependence level by the product of the number of interactions and the duration of interaction, and determine the third type of mining time interval according to the dependence level;

[0091] S210, obtain the consumption record content according to the consumption habit data;

[0092] S211, analyze the dynamic change of the consumption record content, and calculate the dynamic change by the difference between the consumption amount and the consumption type of adjacent time periods;

[0093] S212, calculate the consumption-preference sensitivity degree according to the dynamic change, calculate the consumption-preference sensitivity degree by the ratio of the consumption change size and the preference change size, and determine the fourth type of mining time interval according to the consumption-preference sensitivity degree;

[0094] S213, combine the first type of mining time interval, the second type of mining time interval, the third type of mining time interval and the fourth type of mining time interval, and take the weighted average according to the importance of the influence of each time interval on the recommendation result when combining, to obtain different data mining time interval information.

[0095] The embodiment corresponds to the related technical content of determining data mining time interval information. For user basic situation data, in the user information database, it is found that the age label of the user is updated every half year (because the age of the user increases every year, and the system performs data update synchronization every half year), the storage structure is understood first, the update situation information of the key identity label is obtained, the identity label index and the user classification label are created, and half a year is taken as the first type of mining time interval.

[0096] For travel preference data, the travel record of the user is analyzed, it is found that the user has a stable preference for beach tourism, and will perform beach tourism once a year in the past three years, and the tourism activity type is mainly concentrated in diving and beach leisure. The travel record content of the user is obtained first, the association between the tourism location type and the travel time information is calculated, the preference-time association index is created, the preference stability is calculated according to the time length and the change size of the same preference, and each year is determined as the second type of mining time interval of travel preference data.

[0097] For past interaction data, the interaction record of the user with other users on the tourism related platform is observed, it is found that the user has a high interaction frequency with some users who frequently exchange tourism experience, and the average interaction frequency is 3-4 times per month. The interaction record content is obtained from the past interaction data, the interaction object-time association index is created, the dependence level of the user on the interaction relationship is calculated by the product of the interaction frequency and the interaction duration, and each month is determined as the third type of mining time interval.

[0098] For consumption habit data, the consumption record of the user is analyzed, it is found that the dining consumption habit of the user has certain fluctuation when the season changes, such as the increase of seafood dining consumption in summer. The consumption record content is obtained from the consumption habit data, the dynamic change is obtained by analyzing the difference between the consumption amount and the consumption type of adjacent time periods, the consumption-preference sensitivity is calculated according to the proportion of the consumption change size and the preference change size, and each quarter is determined as the fourth type of mining time interval.

[0099] Finally, the first type of mining time interval (half a year), the second type of mining time interval (one year), the third type of mining time interval (one month) and the fourth type of mining time interval (one quarter) are summarized and combined, and weighted average is performed according to the importance of the influence of each time interval on the recommendation result (assuming that the influence weight of the user basic situation data is 0.2, the influence weight of the travel preference data is 0.3, the influence weight of the past interaction data is 0.3, and the influence weight of the consumption habit data is 0.2), and different data mining time interval information is obtained.

[0100] For example, the data mining time interval needs to be dynamically set according to the update frequency and stability of different data types, so as to balance the data timeliness and mining efficiency.

[0101] User basic situation data mining interval: Analyze the user information database storage structure, find that the identity recognition code is lifelong unchanged, the age is updated once a year, and the system synchronizes the basic information once every 6 months; through the calculation of the update frequency, the age is updated once every 180 days, so the first type of mining interval is set to 6 months.

[0102] Travel preference data mining interval: Extract travel records, including a trip to Sanya in July 2020, a trip to Qingdao in August 2021, and a trip to Xiamen in June 2022; calculate the correlation of location type, time and mode, where the location type is all beach, the time is all summer, and the mode is all airplane plus rental car, the probability of the same characteristics appearing at the same time is 100%, there is no change for 3 consecutive years, the preference stability degree is 90% (full score is 100%); the calculation formula of preference stability degree is: preference stability degree = (the same preference continuously appears for a period of time ÷ total observation time) × (1-preference change amplitude) × 100%; in this example, the same preference continuously appears for a period of time of 3 years, the total observation time is 3 years, and the preference change amplitude is 0, so the preference stability degree = (3 ÷ 3) × (1-0) × 100% = 100%, but considering other potential factors, the actual setting is 90%; accordingly, the second type of mining interval is set to 12 months, i.e. mining once a month before summer every year.

[0103] Past interaction data mining interval: Build an index of interaction objects and frequency, and calculate the dependence level; the interaction frequency with the core interaction object, i.e. 3 high-frequency users, is 50 times per year, the duration is 2 years, and the product of the two is 100, the dependence level is high; the calculation formula of dependence level is: dependence level = interaction frequency × interaction duration; therefore, the third type of mining interval is set to 30 days, i.e. mining once a month.

[0104] Consumption habit data mining interval: Analyze the dynamic changes of consumption, the dining consumption in 2022 increased by 15% compared with 2021, the reason is that the price of seafood has risen, and the accommodation standard has not changed; calculate the consumption and preference sensitivity, the consumption change amplitude is 15%, and the preference change amplitude is 0%; the calculation formula of consumption-preference sensitivity is: consumption-preference sensitivity = consumption change amplitude ÷ preference change amplitude; since the preference change amplitude is 0, the result is meaningless; therefore, according to the seasonal fluctuation, the fourth type of mining interval is set to 90 days, i.e. mining once every quarter.

[0105] Finally, the four types of intervals are integrated, and the weighted calculation is performed according to the influence weight; wherein the basic situation accounts for 10%, the preference accounts for 40%, the interaction accounts for 30%, and the consumption accounts for 20%; the calculation formula of the total mining interval is: total mining interval = (first type mining interval x basic situation weight) + (second type mining interval x preference weight) + (third type mining interval x interaction weight) + (fourth type mining interval x consumption weight); the numerical value is substituted into the formula, 6 months multiplied by 10% plus 12 months multiplied by 40% plus 30 days multiplied by 30% plus 90 days multiplied by 20%, the result is 0.6 months plus 4.8 months plus 9 days plus 18 days, and after conversion, it is 32.4 days, that is, about 32 days of full-amount data mining is performed once.

[0106] In a specific implementation example, the step of obtaining a plurality of batch behavior sample combinations according to different data mining time interval information includes:

[0107] S301, determining a plurality of data mining works according to different data mining time interval information, each data mining work corresponding to a specific behavior data type and a mining target;

[0108] S302, setting a mining configuration requirement for each data mining work, the mining configuration requirement including a data extraction range, a data cleaning rule and a data conversion standard;

[0109] S303, determining a raw behavior data source corresponding to user basic situation data, travel preference data, past interaction data and consumption habit data according to the mining configuration requirement;

[0110] S304, formulating a mining sequence arrangement for each raw behavior data source, the mining sequence arrangement including a mining sequence, a mining frequency and a mining interval time;

[0111] S305, performing a plurality of preliminary mining on each raw behavior data source according to the preset mining sequence arrangement, identifying and correcting abnormal data in the mining process, judging the abnormal data by the deviation size from the average value of historical data, and obtaining mining information;

[0112] S306, merging the mining information corresponding to the user basic situation data, the travel preference data, the past interaction data and the consumption habit data, and adopting a data alignment manner to make different types of data consistent in time and content, and forming a plurality of batch behavior sample combinations.

[0113] The embodiment corresponds to the related technical content of obtaining a behavior sample combination, and a plurality of data mining works are determined according to the above determined data mining time interval information, for example, for travel preference data, data mining is performed one month before the tourism peak season (such as summer) of each year, and the mining target is to obtain the latest travel preference change of the user.

[0114] Set the mining configuration requirements for this data mining work, including the data extraction range for the past year on the tourism social network platform, the tourism reservation software, and the travel preferences related data, the data cleaning rules for removing duplicate records and invalid data (such as incorrect travel location names), and the data conversion standard for converting travel location names to a unified geographic coding format.

[0115] Determine the source of the original behavior data corresponding to the user basic situation data, travel preference data, past interaction data, and consumption habit data according to the mining configuration requirements, which is determined as the user travel dynamic record of the tourism social network platform and the reservation record of the tourism reservation software.

[0116] Develop a mining sequence arrangement for each original behavior data source, first extract data from the tourism social network platform, extract once a week, a total of four times; then extract data from the tourism reservation software, extract once every two weeks, a total of twice.

[0117] According to the preset mining sequence arrangement, multiple preliminary mining is carried out on each original behavior data source, and in the mining process, the abnormal data is judged by the deviation size from the average value of the historical data, such as finding that the user's reservation type of accommodation in a certain tourism reservation is too different from the past preference, that is, the price is much lower than the standard of the high-end hotel he usually chooses, after verification, it is found that it is because the destination of this tourism is special, and there is a lack of hotels that meet his standards in the local, so this data is corrected to a reasonable accommodation choice, such as choosing a local characteristic homestay of the same level.

[0118] Finally, the mining information obtained from different sources is combined, and the data alignment method is adopted to integrate the data obtained from the tourism social network platform and the tourism reservation software in chronological order to form a behavior sample combination.

[0119] For example, based on a 32-day mining interval, data needs to be extracted and processed in batches to form a sample combination.

[0120] Determine the mining task: execute 4 tasks every 32 days, respectively for the four types of data.

[0121] Task 1 is for basic situation, extracting user ID, age, gender and other fields, the goal is to update the user basic label.

[0122] Task 2 is for travel preferences, extracting destination search and collection records in the past 32 days, the goal is to capture preference changes.

[0123] Task 3 is for past interaction, extracting chat records and mutual evaluation content in the past 32 days, the goal is to update the interaction relationship strength.

[0124] Task 4 is for consumption habits, extracting consumption bills in the past 32 days, the goal is to track consumption trends.

[0125] Setting up mining parameters:

[0126] In terms of extraction range, Task 1 extracts data for nearly 180 days because the basic information is updated every six months; Tasks 2 to 4 extract data for nearly 32 days.

[0127] In terms of cleaning rules, Task 1 needs to delete duplicate ID records; Task 2 needs to filter invalid search terms such as "just look around"; Task 3 needs to remove ad screen content; and Task 4 needs to exclude refund orders.

[0128] In terms of conversion standards, unify the location name to the format of city plus area, such as Sanya Yalongwan; and unify the consumption amount to RMB yuan.

[0129] Execute the mining sequence:

[0130] In terms of sequence, first execute Task 1, which is the mining of basic data, and then execute Tasks 2 to 4, which are the mining of dynamic data.

[0131] In terms of frequency, Task 1 is executed once every 180 days, and Tasks 2 to 4 are executed once every 32 days.

[0132] In terms of interval, Tasks 2 to 4 are executed in the order of preference, interaction, and consumption, with one item executed per day and three days completed in a round.

[0133] Handle abnormal data:

[0134] In Task 2, a search record for Harbin was found, which conflicts with the seaside preference. After investigation, it was found that the user had made a mistake, so it was marked as abnormal and deleted.

[0135] In Task 3, a certain interaction record was missing a timestamp, so it was supplemented to 9:22 based on the adjacent record times of June 10, 2023, 9:15 and 9:30.

[0136] Merge samples: Align the four types of data by user ID and timestamp to form a sample combination; for example, the sample of user U20230512001 on June 10, 2023, contains basic information of a 28-year-old female, preference of searching for diving in Qingdao, interaction of discussing diving suits with user U20230401002, and consumption of booking a hotel in Qingdao for 800 yuan.

[0137] In a specific implementation example, the step of obtaining the recommended matching degree evaluation value by merging and analyzing multiple batches of behavior sample combinations includes:

[0138] S401, obtaining the structure form of multiple batches of data and the content of multiple batches of data from multiple batches of behavior sample combinations;

[0139] S402, obtaining a structure result content according to a structure form of the multiple batches of data, the structure result content including a normalized data result content, a semi-normalized data result content and a non-normalized data result content;

[0140] S403, calculating a behavior consistency evaluation value according to the normalized data result content, the semi-normalized data result content and the non-normalized data result content, the behavior consistency evaluation value being obtained by counting a matching number proportion of same behavior descriptions in different structure data;

[0141] S404, extracting a content feature content from the multiple batches of data, the content feature content including an information completeness degree, an information matching degree, an information timeliness degree, an information accuracy degree, an information correlation degree and a behavior consistency degree;

[0142] S405, calculating a content feature evaluation value by using an analytic hierarchy process, the analytic hierarchy process being used to determine an importance degree of each content feature by constructing a judgment matrix;

[0143] S406, calculating a recommendation matching degree evaluation value according to the behavior consistency evaluation value and the content feature evaluation value, the recommendation matching degree evaluation value being a weighted sum of the behavior consistency evaluation value and the content feature evaluation value according to a 3:7 ratio.

[0144] When the content feature evaluation value is calculated by using the analytic hierarchy process, a user's satisfaction evaluation on a recommendation result is collected regularly; the importance degree weight of each content feature in the judgment matrix is adjusted according to the satisfaction evaluation; and the content feature evaluation value is recalculated by using the adjusted judgment matrix, so as to optimize the accuracy of the recommendation matching degree evaluation value.

[0145] The embodiment corresponds to the related technical content of the recommendation matching degree evaluation value obtained by merging and analyzing, from a combination of a batch of behavior samples, the structure form of data, it is found that the data of a tourism social network platform is semi-normalized data, the data of a tourism reservation software is normalized data, and the data of a tourism evaluation webpage is non-normalized data.

[0146] The behavior consistency evaluation value is calculated by counting a matching number proportion of same behavior descriptions in different structure data; for example, when describing a tourism preference, a user mentions “like a coastal city” in the tourism social network platform, the reservation record of the tourism reservation software shows that he has booked a trip to a coastal city many times, and there are a large number of positive evaluations on coastal scenic spots in the tourism evaluation webpage; by statistics, the matching number of same behavior descriptions is 8, and the total number of behavior descriptions is 10, so the behavior consistency evaluation value is 8 ÷ 10 × 100% = 80%; the calculation formula of the behavior consistency evaluation value is: behavior consistency evaluation value = (matching number of same behavior descriptions ÷ total number of behavior descriptions) × 100%.

[0147] Content features are extracted from the content of the batch data, including information completeness, information matching degree, etc. The importance of each content feature is determined by constructing a judgment matrix using the analytic hierarchy process. Assume that the weight of information completeness is 0.2, the weight of information matching degree is 0.3, the weight of information timeliness is 0.1, the weight of information accuracy is 0.2, the weight of information relevance is 0.1, and the weight of behavior consistency is 0.1.

[0148] Each content feature is scored, with a full score of 10 points. In terms of information completeness, the data covers user basic information, preferences and other aspects, scoring 8 points. In terms of information matching degree, different sources of data basically agree on the user's preferences, scoring 9 points. In terms of information timeliness, the data is mostly within 32 days, scoring 7 points. In terms of information accuracy, after verification, the data is consistent with the actual situation, scoring 8 points. In terms of information relevance, the data is all around tourism-related content, scoring 8 points. In terms of behavior consistency, the user's behavior is consistent in different scenarios, scoring 8 points.

[0149] The calculation formula of the content feature evaluation value is: content feature evaluation value = (information completeness score x information completeness weight) + (information matching degree score x information matching degree weight) + (information timeliness score x information timeliness weight) + (information accuracy score x information accuracy weight) + (information relevance score x information relevance weight) + (behavior consistency score x behavior consistency weight). Substituting the value into the formula gives: (8x0.2) + (9x0.3) + (7x0.1) + (8x0.2) + (8x0.1) + (8x0.1) = 1.6 + 2.7 + 0.7 + 1.6 + 0.8 + 0.8 = 8.2.

[0150] To continuously optimize the evaluation results, user satisfaction evaluations of the recommended results need to be collected regularly. Assume that the user's satisfaction with the early recommended results is 85%. According to this satisfaction, the importance of each content feature is adjusted. The weight of information matching degree is increased to 0.35, the weight of information relevance is increased to 0.15, the weight of information completeness is reduced to 0.15, the weight of information accuracy is reduced to 0.15, the weight of information timeliness is reduced to 0.08, and the weight of behavior consistency is reduced to 0.07. Using the adjusted weights, the content feature evaluation value is recalculated: (8x0.15) + (9x0.35) + (7x0.08) + (8x0.15) + (8x0.15) + (8x0.07) = 1.2 + 3.15 + 0.56 + 1.2 + 1.2 + 0.56 = 7.87, which improves the accuracy of the recommended matching degree evaluation value.

[0151] Finally, the weighted total sum is calculated according to the ratio of the behavior consistency evaluation value and the content feature evaluation value 3:7, and the recommendation matching degree evaluation value is obtained. The calculation formula of the recommendation matching degree evaluation value is: recommendation matching degree evaluation value = (behavior consistency evaluation value x 0.3) + (content feature evaluation value x 0.7). Substituting the data, we get: (80% x 0.3) + (8.2 x 0.7) = 0.24 + 5.74 = 5.98.

[0152] In a specific implementation example, the step of calculating the recommendation judgment value according to the recommendation matching degree evaluation value includes: S501, determining the data importance proportion of user basic situation data, travel preference data, past interaction data and consumption habit data in the travel partner recommendation work, and the data importance proportion is determined by expert scoring combined with user feedback data adjustment, wherein the user basic situation data accounts for 20%, the travel preference data accounts for 35%, the past interaction data accounts for 25%, and the consumption habit data accounts for 20%;

[0153] S502, statistics of actual data quantity corresponding to user basic situation data, travel preference data, past interaction data and consumption habit data, and the actual data quantity is counted by the number of data records;

[0154] S503, calculating the matching quality evaluation value corresponding to the user basic situation data, travel preference data, past interaction data and consumption habit data according to the recommendation matching degree evaluation value, data importance proportion and actual data quantity, and the matching quality evaluation value is the product of the recommendation matching degree evaluation value, data importance proportion and actual data quantity;

[0155] S504, statistics of total verification data quantity corresponding to user basic situation data, travel preference data, past interaction data and consumption habit data, and the total verification data quantity is the sum of the number of verified accurate data in each type of data;

[0156] S505, calculating the recommendation judgment value according to the total verification data quantity and the matching quality evaluation value, and the recommendation judgment value is the ratio of the matching quality evaluation value and the total verification data quantity.

[0157] The embodiment corresponds to the related technical content of calculating the recommended judgment value. The recommended judgment value is calculated to finally determine whether to recommend a certain travel partner to the user. The value comprehensively considers the importance of various data, the actual quantity and the matching quality. First, the data importance proportion of the user basic situation data, the travel preference data, the past interaction data and the consumption habit data in the travel partner recommendation work is determined. The user basic situation data accounts for 20%, the travel preference data accounts for 35%, the past interaction data accounts for 25%, and the consumption habit data accounts for 20%. These proportions are determined by expert scoring combined with a large amount of user feedback data adjustment, which can better reflect the influence of various data on the recommended result.

[0158] The actual data quantity corresponding to the four types of data is counted. Through the statistics of the number of data records, the user basic situation data is 100, the travel preference data is 200, the past interaction data is 150, and the consumption habit data is 120.

[0159] According to the recommended matching degree evaluation value, the data importance proportion and the actual data quantity, the matching quality evaluation value corresponding to each type of data is calculated. The calculation formula of the matching quality evaluation value is: the matching quality evaluation value of a certain type of data = recommended matching degree evaluation value × importance proportion of the data × actual data quantity of the data.

[0160] The matching quality evaluation value of the user basic situation data = 5.98 × 20% × 100 = 5.98 × 0.2 × 100 = 119.6;

[0161] The matching quality evaluation value of the travel preference data = 5.98 × 35% × 200 = 5.98 × 0.35 × 200 = 418.6;

[0162] The matching quality evaluation value of the past interaction data = 5.98 × 25% × 150 = 5.98 × 0.25 × 150 = 224.25;

[0163] The matching quality evaluation value of the consumption habit data = 5.98 × 20% × 120 = 5.98 × 0.2 × 120 = 143.52.

[0164] The total amount of verification data corresponding to the four types of data, i.e., the total number of data items in each type of data that are verified to be accurate, is 500; the recommended judgment value is calculated based on the total amount of verification data and the matching quality evaluation value, and the calculation formula of the recommended judgment value is: recommended judgment value = (user basic situation data matching quality evaluation value + travel preference data matching quality evaluation value + past interaction data matching quality evaluation value + consumption habit data matching quality evaluation value) ÷ total amount of verification data; by substituting the data, it can be calculated that: (119.6 + 418.6 + 224.25 + 143.52) ÷ 500 = 895.97 ÷ 500 = 1.79194.

[0165] In a specific implementation example, after sending the qualified travel companion information to the device used by the user, it further includes receiving feedback data of the user on the recommended travel companion information, the feedback data including accepting the recommendation, rejecting the recommendation and the reason for rejection; the feedback data is classified and arranged, and the association between the feedback data and the data in each link in the corresponding recommendation process is established; the influencing factors of the success or failure of the recommendation are analyzed according to the association result; the parameter settings of subsequent data mining and the calculation method of the recommendation matching degree evaluation are adjusted according to the influencing factors.

[0166] In the process of reselecting the behavior source data, the specific data types that cause the recommended judgment value to be substandard in the previously selected behavior source data are analyzed; the data acquisition range is expanded or the data screening standard is improved for the specific data types; the pre-processing of the newly acquired behavior source data includes removing duplicate data and supplementing missing data; the pre-processed behavior source data is subjected to data mining and recommended judgment according to the original steps.

[0167] After the recommended judgment value of the present embodiment is calculated, it is compared with the pre-set limit value to determine whether to recommend and subsequent operation, assuming that the pre-set limit value is 1.5, since the calculated recommended judgment value 1.79194 is greater than 1.5, the qualified travel companion information is sent to the device used by the user; these information includes the basic situation, travel preference, past interaction record and consumption habit of the travel companion, so that the user can fully understand the potential partner.

[0168] After receiving the recommended information, the user will give feedback on the recommended travel companion; the feedback data may include accepting the recommendation, rejecting the recommendation and the reason for rejection; for example, the user indicates that the recommendation is rejected, and the reason for rejection is that the consumption habit of the other party is too different from oneself, and oneself prefers high-end accommodation, while the other party prefers economy hotel.

[0169] The feedback data is classified and sorted, and is associated with the data of each link in the corresponding recommendation process. Through analysis, it is found that although the consumption habit data accounts for 20% in the recommendation matching evaluation, the user has a high sensitivity to consumption habits, which is a key factor affecting his decision-making, which leads to the failure of this recommendation. According to this influencing factor, the parameter settings of subsequent data mining are adjusted, such as increasing the proportion of consumption habit data in the importance of data to 25%, and adjusting the calculation method of the recommendation matching degree evaluation, adjusting the proportion of behavior consistency evaluation value and content feature evaluation value to 2:8, increasing the weight of consumption habit matching degree in content feature evaluation.

[0170] If the recommendation judgment value is less than the pre-set limit value, it means that the current behavior source data is not enough to generate a suitable recommendation; at this time, the specific data type in the previously selected behavior source data that leads to the substandard recommendation judgment value needs to be analyzed, assuming that the past interaction data is insufficient to accurately evaluate the user's interaction preferences.

[0171] For this data type, expand the data acquisition range, collect the user's interaction records from more tourism social platforms, such as tourism forums, tourism enthusiast groups, etc.; or improve the data screening standard, only keep the data with high interaction frequency and substantial interaction content, such as interaction frequency less than 3 times a week, interaction content around travel planning, scenic spot recommendation, etc.

[0172] Preprocess the newly obtained behavior source data to remove duplicate interaction records, such as duplicate publication of the same interaction content on different platforms; supplement missing interaction time, interaction object, etc., such as determining the specific time of interaction according to the time description in the interaction content; perform data mining and recommendation judgment on the preprocessed behavior source data according to the previous steps, and recalculate the recommendation judgment value; for example, the actual data quantity of the newly obtained past interaction data is 200; its matching quality evaluation value = 5.98 x 25% x 200 = 299, the total matching quality evaluation value becomes 119.6 + 418.6 + 299 + 143.52 = 980.72, the new recommendation judgment value = 980.72 ÷ 500 = 1.96144, which is greater than the limit value 1.5, so a suitable travel partner can be recommended to the user.

[0173] Embodiment 2:

[0174] Based on embodiment 1, the embodiment also provides a travel partner personalized recommendation system based on behavior data mining, specifically comprising:

[0175] A data collection module for collecting multiple sets of behavior source data of users in a tourism scene, each set of behavior source data containing user basic situation data, travel preference data, past interaction data and consumption habit data;

[0176] a time interval determination module configured to determine different data mining time interval information according to the behavior source data;

[0177] a sample combination obtaining module configured to obtain multiple batches of behavior sample combinations according to the different data mining time interval information;

[0178] an analysis and evaluation module configured to perform a merging analysis on the multiple batches of behavior sample combinations to obtain a recommended matching degree evaluation value;

[0179] a judgment value calculation module configured to calculate a recommended judgment value according to the recommended matching degree evaluation value;

[0180] a recommendation judgment module configured to compare the recommended judgment value with a pre-set threshold value;

[0181] if the recommended judgment value is greater than the pre-set threshold value, sending the qualified travel companion information to a device used by a user;

[0182] if the recommended judgment value is less than the pre-set threshold value, re-selecting the behavior source data and triggering the data collection module to perform a related operation.

[0183] In a specific implementation example, the data collection module comprises:

[0184] a path determination unit configured to determine multiple user data obtaining paths in a travel scenario, the paths comprising a travel social network platform, a travel reservation software, a travel review webpage and an offline travel service record system;

[0185] a raw data collection unit configured to collect multiple groups of raw behavior data through the multiple user data obtaining paths;

[0186] a basic condition data division unit configured to analyze identity features contained in each group of raw behavior data, the identity features covering a user's name, age, gender, birthplace and identity identification code, and divide corresponding raw behavior data into user basic condition data according to the identity features;

[0187] a preference data division unit configured to analyze preference features contained in each group of raw behavior data, the preference features covering a user's preferred travel location type, travel activity type, travel time length, travel season and accommodation selection criterion, and divide corresponding raw behavior data into travel preference data according to the preference features;

[0188] An interaction data division unit is configured to analyze interaction features contained in each set of original behavior data, the interaction features including the number of communications, communication content topics, the number of joint travels, and mutual evaluation scores of the user with other users on the travel-related platform, and to divide the corresponding original behavior data into past interaction data according to the interaction features;

[0189] A consumption data division unit is configured to analyze consumption features contained in each set of original behavior data, the consumption features including dining expense standards, transportation expense ranges, types of scenic spot ticket consumption, and shopping consumption tendencies of the user during travel, and to divide the corresponding original behavior data into consumption habit data according to the consumption features.

[0190] In an embodiment, the time interval determination module includes:

[0191] A structure understanding unit is configured to understand the storage structure of the user in the user information database according to the user basic situation data;

[0192] An update information acquisition unit is configured to obtain update situation information of the key identity according to the storage structure, the update situation information including the number of updates, update change sizes, and update triggering conditions;

[0193] An index mark creation unit is configured to create an identity index and user classification marks including age marks, gender marks, and region marks according to the update situation information;

[0194] A first time interval unit is configured to record the number of updates of the identity index and the user classification marks, and to take the number of updates as a first type of mining time interval;

[0195] A travel record acquisition unit is configured to obtain travel record content including travel location type information, travel time information, and travel mode information according to the travel preference data;

[0196] An association index creation unit is configured to calculate the association between the travel location type information, the travel time information, and the travel mode information, to determine the association by counting the number of times and the likelihood of simultaneous occurrence of different information, and to create a preference-time association index according to the association;

[0197] A second time interval unit is configured to calculate a plurality of preference stability situations according to the preference-time association index, to calculate the preference stability situations by the length of time and the change size of continuous occurrence of the same preference, and to determine a second type of mining time interval of the travel preference data according to the plurality of preference stability situations;

[0198] An interaction association unit is configured to obtain interaction record content according to the past interaction data and to create an interaction object-number of times association index;

[0199] a third time interval unit configured to analyze the level of dependence of the user on the interactive relationship according to the interactive object-number of times association index, calculate the level of dependence by the product of the number of interactions and the duration of the interaction, and determine a third type of mining time interval according to the level of dependence;

[0200] a sensitivity calculation unit configured to obtain a consumption-preference sensitivity degree according to the consumption habit data, and calculate the consumption-preference sensitivity degree by the ratio of the size of the consumption change and the size of the preference change;

[0201] a fourth time interval unit configured to determine a fourth type of mining time interval according to the consumption-preference sensitivity degree;

[0202] a time interval merging unit configured to merge the first type of mining time interval, the second type of mining time interval, the third type of mining time interval and the fourth type of mining time interval, and perform weighted averaging according to the importance of the influence of each time interval on the recommendation result to obtain different data mining time interval information.

[0203] In an embodiment, the sample combination obtaining module comprises:

[0204] a work determination unit configured to determine a plurality of data mining works according to the different data mining time interval information, each data mining work corresponding to a specific behavior data type and a mining target;

[0205] a configuration setting unit configured to set mining configuration requirements for each data mining work, the mining configuration requirements including a data extraction range, a data cleaning rule and a data conversion standard;

[0206] a data source determination unit configured to determine original behavior data sources corresponding to the user basic situation data, the travel preference data, the past interaction data and the consumption habit data according to the mining configuration requirements;

[0207] a sequence making unit configured to make a mining sequence arrangement for each original behavior data source, the mining sequence arrangement including a mining sequence, a mining number of times and a mining interval time;

[0208] a mining execution unit configured to perform a plurality of preliminary mining on each original behavior data source according to the preset mining sequence arrangement, identify and correct abnormal data in the mining process, and obtain mining information by judging the abnormal data according to the deviation size from the average value of historical data;

[0209] a sample merging unit configured to merge the mining information corresponding to the user basic situation data, the travel preference data, the past interaction data and the consumption habit data, and adopt a data alignment method to make different types of data consistent in time and content to form a plurality of batches of behavior sample combinations.

[0210] The above-described embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented by software, the above-described embodiments can be implemented in whole or in part in the form of a computer program product. A person of ordinary skill in the art can be aware that units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed by hardware or software depends on the specific application and design constraints of the technical solutions.

[0211] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, and can be located in one place or distributed on multiple network units. Part or all of the units can be selected to achieve the purpose of the embodiments according to actual needs.

[0212] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application.

Claims

1.A method for recommending a travel companion based on behavior data mining, the method comprising: The steps of the recommendation method include: collecting multiple sets of behavior source data of a user in a travel scene, each set of behavior source data containing user basic condition data, travel preference data, past interaction data, and consumption habit data; According to the behavior source data, different data mining time interval information is determined, including: According to the user basic condition data, the storage structure of the user in the user information database is understood, and the update information of the key identity is obtained; according to the update information, an identity index and a user classification mark are created, and the number of updates is recorded, which is used as the first type of mining time interval; According to the travel preference data, the travel record content of the user is obtained, including travel location type information, travel time information, and travel mode information; the association between the travel location type information, the travel time information, and the travel mode information is calculated, the association is determined by counting the number of times and the possibility of different information appearing at the same time, and a preference-time association index is created according to the association; according to the preference-time association index, multiple preference stability conditions are calculated, the preference stability condition is calculated by the length of time of the same preference continuously appearing and the change size, and the second type of mining time interval of the travel preference data is determined according to the multiple preference stability conditions; Multiple batch behavior sample combinations are obtained according to different data mining time interval information; The multiple batch behavior sample combinations are merged and analyzed to obtain a recommendation matching degree evaluation value; A recommendation judgment value is calculated according to the recommendation matching degree evaluation value; The recommendation judgment value is compared with the pre-set limit value; If the recommendation judgment value is greater than or equal to the pre-set limit value, the user's device is sent the travel companion information that meets the conditions; otherwise, the behavior source data is reselected and the above steps are repeated. 2.The method of claim 1, wherein, The step of collecting multiple sets of behavior source data of a user in a travel scene includes: Determine multiple user data acquisition approaches in the travel scene; collect multiple sets of original behavior data through the multiple user data acquisition approaches; verify the data legality of the original behavior data obtained by each approach; only keep the original behavior data that passes the legality verification for subsequent processing; Analyze the identity information contained in each set of original behavior data, and divide the corresponding original behavior data into user basic condition data; analyze the preference characteristics contained in each set of original behavior data, and divide the corresponding original behavior data into travel preference data; analyze the interaction characteristics contained in each set of original behavior data, which include the number of exchanges, exchange content topics, number of joint travels, and mutual evaluation scores of the user and other users on a travel-related platform, and divide the corresponding original behavior data into past interaction data; analyze the consumption characteristics contained in each set of original behavior data, and divide the corresponding original behavior data into consumption habit data. 3.The method of claim 1, wherein, The step of determining different data mining time interval information according to the behavior source data also includes: According to the interaction record content obtained from the past interaction data, an interaction object-number association index is created, the dependence level of the user on the interaction relationship is analyzed, the dependence level is calculated by the product of the number of interactions and the duration of the interaction, and the third type of mining time interval is determined according to the dependence level; According to the consumption habit data, the consumption record content is obtained; the dynamic change of the consumption record content is analyzed, the dynamic change is calculated through the difference of the consumption amount and the consumption type of adjacent time periods; the consumption-preference sensitivity degree is calculated according to the dynamic change, the consumption-preference sensitivity degree is calculated through the proportion of the consumption change size and the preference change size, and the fourth type of mining time interval is determined according to the consumption-preference sensitivity degree; The first type to the fourth type of mining time interval are summarized and merged, and when merging, weighted average is performed according to the importance of each time interval on the recommendation result, and different data mining time interval information is obtained. 4.The method of claim 3, wherein, The importance of each time interval on the recommendation result is determined by the following method: The recommendation success rate corresponding to each time interval in the historical recommendation result is collected; The ratio of the recommendation success rate of each time interval to the average recommendation success rate is calculated; The ratio is taken as the importance of each time interval on the recommendation result; When weighted average is performed according to the importance, the weight value is equal to the ratio of the importance to the sum of all importance. 5.The method of claim 1, wherein, The steps of obtaining multiple batch behavior sample combinations according to different data mining time interval information include: According to different data mining time interval information, multiple data mining jobs are determined, and mining configuration parameters are set for each data mining job, the mining configuration parameters include data extraction range, data cleaning rule and data conversion standard; the original behavior data source corresponding to the basic data is obtained according to the mining configuration parameters, the basic data includes user basic situation data, travel preference data, historical interaction data and consumption behavior data; the mining sequence is obtained according to each original behavior data source; high-frequency preliminary mining is performed on each original behavior data source according to the preset mining sequence, and abnormal data is identified and corrected in the mining process, the abnormal data is judged by the deviation size from the average value of historical data, and mining information is obtained; The mining information corresponding to the basic data is integrated, and when integrating, the data alignment method is adopted to make different types of data consistent in time and content dimensions, and multiple batch behavior sample combinations are formed. 6.The method of claim 1, wherein, The steps of obtaining the recommendation matching degree evaluation value by merging and analyzing the multiple batch behavior sample combinations include: obtaining the structure form of multiple batches of data and the content of multiple batches of data from the multiple batch behavior sample combinations; The structure result content is obtained according to the structure form of multiple batches of data; the behavior consistency evaluation value is calculated according to the normalized data result content, the semi-normalized data result content and the non-normalized data result content, and the behavior consistency evaluation value is obtained by counting the matching number proportion of the same behavior description in different structure data; The content feature content is extracted from the content of multiple batches of data; the content feature evaluation value is calculated by using the analytic hierarchy process, and the importance of each content feature is determined by constructing a judgment matrix in the analytic hierarchy process; The recommendation matching degree evaluation value is calculated according to the behavior consistency evaluation value and the content feature evaluation value, and the recommendation matching degree evaluation value is the weighted sum of the behavior consistency evaluation value and the content feature evaluation value according to a preset proportion. 7.The method of claim 5, wherein, The step of calculating the recommendation judgment value according to the recommendation matching degree evaluation value comprises: determining the proportion of the data importance degree of the basic data in the travel companion recommendation work, and the proportion of the data importance degree is determined by expert scoring combined with user feedback data adjustment; The actual data quantity corresponding to the basic data is counted, and the actual data quantity is counted by the number of data records; The matching quality evaluation value corresponding to the basic data is calculated according to the recommendation matching degree evaluation value, the proportion of the data importance degree and the actual data quantity, and the matching quality evaluation value is the product of the recommendation matching degree evaluation value, the proportion of the data importance degree and the actual data quantity; The total verification data quantity corresponding to the basic data is counted, and the total verification data quantity is the sum of the verified accurate data quantity in each type of data; The recommendation judgment value is calculated according to the total verification data quantity and the matching quality evaluation value, and the recommendation judgment value is the ratio of the matching quality evaluation value to the total verification data quantity. 8.The method of claim 1, wherein, After sending the qualified travel companion information to the device used by the user, it further comprises: Receiving user feedback data on the recommended travel companion information, the feedback data including accepting the recommendation, rejecting the recommendation and the reason for rejection; classifying and arranging the feedback data, grouping according to the feedback type and feedback time during classification and arrangement; establishing the association between the feedback data and the data in each link of the corresponding recommendation process, the association mode including data identification correspondence and time node correspondence; analyzing the influencing factors of the success or failure of the recommendation according to the association result, the influencing factors including insufficient data completeness and matching algorithm deviation; adjusting the parameter setting of subsequent data mining and the calculation method of recommendation matching degree evaluation according to the influencing factors, the parameter setting adjustment including adjusting the data extraction range and data cleaning rules. 9.The method of claim 1, wherein, In the process of reselecting the behavior source data, it comprises: Analyzing the specific data types in the previously selected behavior source data that cause the recommendation judgment value to be substandard, the specific data types including user basic situation data, travel preference data, past interaction data and consumption habit data; adjusting the data acquisition range or data screening standard according to the specific data types, the adjustment of the data acquisition range including increasing the data acquisition channel, and the adjustment of the data screening standard including setting the data completeness rule; preprocessing the newly acquired behavior source data, the preprocessing including removing duplicate data and supplementing missing data; performing data mining and recommendation judgment on the preprocessed behavior source data according to the original steps.

Citation Information

Patent Citations

  • Urban operation index monitoring method based on multi-source heterogeneity

    CN119494463A

  • Tourism information recommendation method and system based on big data platform

    CN120030233A