Big data-based document travel consumption data management method and system
Through the big data-based cultural and tourism consumption data management method, the missing data in cultural and tourism consumption data is identified and filled, the problem of insufficient data integrity is solved and the accuracy and reliability of data analysis is improved.
Patent Information
- Application Number
- CN202510092147.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When obtaining cultural and tourism consumption data in the prior art, the data of some users cannot be fully obtained, resulting in poor data integrity and affecting the accuracy of subsequent data analysis.
The cultural and tourism consumption data management method based on big data is adopted to obtain transportation, accommodation, and food consumption data through feature identification, generate consumption nodes, calculate consumption similarity, and fill in missing data based on similarity.
It improves the integrity and accuracy of data, can more objectively reflect users' cultural and tourism consumption behavior, and enhances the reliability of data analysis.
Smart Images

Figure CN120013709A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data management technology, and in particular to a cultural and tourism consumption data management method and system based on big data. Background Art
[0002] The cultural tourism industry is an important part of the tourism industry. The general trend of the development of the cultural tourism industry is high-quality development. Cultural tourism consumption data plays an important role in promoting the development of the cultural tourism industry, enhancing tourists' experience, and assisting policy formulation.
[0003] At present, local cultural and tourism departments obtain cultural and tourism consumption data to analyze consumer preferences, needs and behavior patterns, so as to effectively support the subsequent product recommendations and package design. At present, the acquisition of cultural and tourism consumption data is generally determined by crawling data from various platforms through web crawler technology.
[0004] In the above-mentioned related technologies, in the process of data acquisition, all cultural and tourism consumption data of some users cannot be fully obtained, resulting in poor data integrity. At this time, it is impossible to objectively reflect cultural and tourism consumption behavior based on the user data, which brings inconvenience to subsequent data analysis. There is still room for improvement. Summary of the invention
[0005] In order to improve data integrity and facilitate data analysis, the present application provides a cultural and tourism consumption data management method and system based on big data.
[0006] In the first aspect, the present application provides a method for managing cultural tourism consumption data based on big data, which adopts the following technical solutions:
[0007] A method for managing cultural tourism consumption data based on big data, comprising:
[0008] Obtain user consumption data;
[0009] Perform feature recognition in the user consumption data to determine transportation consumption data, accommodation consumption data, and food consumption data respectively;
[0010] Determine the time period for cultural and tourism consumption based on transportation consumption data, and determine the duration of cultural and tourism consumption based on the time period for cultural and tourism consumption;
[0011] Generate consumption nodes according to preset consumption rules during the cultural and tourism consumption period, and add each accommodation consumption data and food consumption data to the consumption nodes according to the corresponding time points;
[0012] In each consumption node, a consumption node that does not contain accommodation consumption data or food consumption data is defined as a missing node, and the remaining consumption nodes are defined as valid nodes;
[0013] Determine candidate personnel based on the duration of cultural travel and effective nodes, and calculate the consumption similarity based on the accommodation consumption data and food consumption data in the effective nodes of the candidate personnel and the accommodation consumption data and food consumption data in the effective nodes of the current personnel;
[0014] Alternative persons whose consumption similarity is greater than the preset demand similarity are defined as similar persons, and predicted consumption data are generated based on the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar persons, and the predicted consumption data are filled into the missing nodes.
[0015] Optionally, the steps for determining candidate personnel based on the duration of the cultural tour and effective nodes include:
[0016] Define the people whose cultural travel time is the same as that of the current user as the preliminary candidates;
[0017] Count the valid nodes of the current user to determine the valid quantity, and calculate the required quantity based on the valid quantity and the preset proportion parameters;
[0018] The consumption nodes corresponding to the valid nodes of the current user are defined as analysis nodes among the preliminary candidates, and the analysis nodes that are valid nodes are counted to determine the analysis quantity;
[0019] The candidates whose analysis number is not less than the effective number are defined as candidates.
[0020] Optionally, after the candidate is determined, the cultural tourism consumption data management method based on big data also includes:
[0021] Counting the candidates to determine the number of candidates;
[0022] Determine whether the number of alternatives is greater than the preset clear number;
[0023] If the number of candidates is not greater than the clear number, the currently determined candidates are maintained;
[0024] If the number of candidates is greater than the clear number, count the valid nodes in each candidate to determine the number of valid candidates;
[0025] According to the number of valid candidates, the candidate personnel are sorted from large to small to output the personnel selection order, and according to the clear number, the candidate personnel at the front of the personnel selection order are maintained, and the remaining candidate personnel are defined and cancelled.
[0026] Optionally, the step of calculating the consumption similarity based on the accommodation consumption data and the food consumption data in the valid node of the candidate personnel and the accommodation consumption data and the food consumption data in the valid node of the current personnel includes:
[0027] The analysis node that is a valid node among the candidate personnel and the valid node of the current user are defined as binding nodes to each other;
[0028] In the binding node, a difference calculation is performed based on the accommodation consumption data to determine the accommodation consumption difference, and a difference calculation is performed based on the food consumption data to determine the food consumption difference;
[0029] Calculate the node similarity based on the difference in accommodation consumption, the difference in food consumption and the preset calculation parameters;
[0030] A node similarity is randomly selected from all node similarities as the central similarity, and the remaining node similarities are defined as comparative similarities;
[0031] The similarity interval is determined by calculating the center similarity and all the comparison similarities, and the center similarity corresponding to the similarity interval with the smallest value is defined as the standard similarity;
[0032] The effective interval is constructed by calculation based on the standard similarity and the preset proximity, and the consumption similarity is determined by average calculation based on the similarities of all nodes in the effective interval.
[0033] Optionally, the step of generating predicted consumption data according to the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar persons includes:
[0034] Define the missing data type in the missing node of the current user as the missing type, and define the associated data in the valid nodes corresponding to similar persons according to the missing type;
[0035] A simulated consumption data is randomly generated, and a difference calculation is performed based on the simulated consumption data and the associated data to determine the simulated consumption difference;
[0036] Determine the correlation value corresponding to the consumption similarity according to the preset similarity matching relationship;
[0037] Calculate the deviation degree value based on all simulated consumption differences and the corresponding correlation degree values;
[0038] The deviation degree value with the smallest numerical value is determined according to a preset sorting rule, and the simulated consumption data corresponding to the deviation degree value is defined as the predicted consumption data.
[0039] Optionally, after the deviation degree value is determined, the cultural tourism consumption data management method based on big data also includes:
[0040] Determine whether there are at least two simulated consumption data with the same and minimum deviation degree values;
[0041] If there are not at least two simulated consumption data with the same and smallest deviation degree values, the simulated consumption data corresponding to the smallest deviation degree value is determined as the predicted consumption data;
[0042] If there are at least two simulated consumption data with the same and smallest deviation degree values, the simulated consumption data corresponding to the smallest deviation degree value is determined as the candidate consumption data;
[0043] According to the missing type, the data corresponding to the missing type in the remaining consumption nodes of the current user is defined as the relevant consumption data;
[0044] Calculate according to each relevant consumption data to determine the representative coefficient of each relevant consumption data, and define the relevant consumption data corresponding to the largest representative coefficient as the control consumption data;
[0045] The difference calculation is performed based on the alternative consumption data and the control consumption data to determine the interval consumption data, and the alternative consumption data corresponding to the interval consumption data with the smallest value is defined as the predicted consumption data.
[0046] Optionally, after the predicted consumption data is determined, the cultural tourism consumption data management method based on big data also includes:
[0047] Randomly select one associated data from all associated data to exclude and combine them according to the remaining associated data to determine an associated set, and define the associated set that is not in the associated set as excluded data;
[0048] Analyze the association set to determine the predicted consumption data corresponding to the excluded data, and define the predicted consumption data as virtual consumption data;
[0049] Calculate the data deviation ratio based on the virtual consumption data and the corresponding excluded data;
[0050] A calculation is performed based on all data deviation ratios to determine a reasonable deviation ratio, and the currently determined predicted consumption data is updated based on the reasonable deviation ratio.
[0051] In the second aspect, the present application provides a cultural tourism consumption data management system based on big data, which adopts the following technical solutions:
[0052] A cultural tourism consumption data management system based on big data, comprising:
[0053] Acquisition module, used to obtain user consumption data;
[0054] A processing module, connected to the acquisition module and the judgment module, for storing and processing information;
[0055] A judgment module, connected with the acquisition module and the processing module, for judging the information;
[0056] The processing module performs feature recognition in the user consumption data to respectively determine the transportation consumption data, the accommodation consumption data and the food consumption data;
[0057] The processing module determines the cultural travel consumption time period based on the transportation consumption data, and determines the cultural travel duration based on the cultural travel consumption time period;
[0058] The processing module generates consumption nodes according to preset consumption rules during the cultural and tourism consumption period, and adds each accommodation consumption data and food consumption data to the consumption node according to the corresponding time point;
[0059] The processing module defines the consumption nodes that do not contain accommodation consumption data or food consumption data in each consumption node as missing nodes, and defines the remaining consumption nodes as valid nodes;
[0060] The processing module determines candidate personnel according to the duration of cultural travel and effective nodes, and calculates the accommodation consumption data and food consumption data in the effective nodes of the candidate personnel and the accommodation consumption data and food consumption data in the effective nodes of the current personnel to determine the consumption similarity;
[0061] The processing module defines the candidate persons whose consumption similarity determined by the judgment module is greater than the preset demand similarity as similar persons, and generates predicted consumption data based on the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar persons, and fills the predicted consumption data into the missing nodes.
[0062] In summary, the present application includes at least one of the following beneficial technical effects:
[0063] 1. After the consumption data is obtained, consumption habits analysis can be performed for users with missing consumption data to identify people with similar consumption habits, thereby filling in the missing data to improve the integrity of the consumption data;
[0064] 2. For people with different similar consumption habits, consumption data can be referenced according to the degree of similarity, thereby improving the accuracy of the filled data;
[0065] 3. Existing data can be used to verify the predictions to determine the prediction deviations that may occur in the prediction situation, so that the data filled in the prediction can be corrected to further improve the accuracy of the filled data. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 This is a flow chart of the cultural and tourism consumption data management method based on big data.
[0067] Figure 2 It is a module flow chart of the cultural and tourism consumption data management method based on big data. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solutions and advantages of this application more clear, the following Figure 1-Figure 2 It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0069] The embodiments of the present application are further described in detail below in conjunction with the drawings in the specification.
[0070] The present application embodiment discloses a method for managing cultural tourism consumption data based on big data, referring to Figure 1 The method flow of the cultural tourism consumption data management method based on big data includes the following steps:
[0071] Step S100: Obtain user consumption data.
[0072] User consumption data refers to the data on local consumption of users obtained through crawler technology and social media platform data identification.
[0073] Step S101: performing feature recognition in the user consumption data to respectively determine the transportation consumption data, the accommodation consumption data and the food consumption data.
[0074] Transportation consumption data refers to the consumption data used for transportation in the user consumption data, accommodation consumption data refers to the consumption data used for accommodation in the user consumption data, and food consumption data refers to the consumption data used for food in the user consumption data. It can be obtained through the registration information of the payment recipient, such as determining the tourist's reservation information from the platform.
[0075] Step S102: Determine the cultural travel consumption time period based on the transportation consumption data, and determine the cultural travel duration based on the cultural travel consumption time period.
[0076] The cultural and tourism consumption time period is the time point when the current user travels in the current city. The transportation consumption data can be used to determine when the user arrives in the city and when the user leaves the city. The cultural and tourism consumption time period can be determined based on these two data. When the time points of arrival in the city and departure from the city are not fully collected, it can be determined based on the earliest and latest remaining consumption data; the cultural and tourism duration is the duration corresponding to the cultural and tourism consumption time period, which is recorded in integer days, such as 5 days and 4 nights, 7 days and 6 nights, etc.
[0077] Step S103: Generate consumption nodes according to preset consumption rules during the cultural and tourism consumption time period, and add each accommodation consumption data and food consumption data to the consumption node according to the corresponding time point.
[0078] Consumption rules are rules set by staff for statistical analysis of consumption data. Consumption nodes are time nodes used to count consumption data. Generally, a consumption node is designed every 6 hours, that is, a consumption node is set at 0:00, 6:00, 12:00, and 18:00 every day. For example, the cultural and tourism consumption period is from 11:00 on October 1 to 11:00 on October 6. Then, two consumption nodes at 12:00 and 18:00 need to be set on October 1, and two consumption nodes at 0:00 and 6:00 need to be set on October 6 at 11:00. Four consumption nodes need to be set on the remaining dates. The time period for consumption can be known based on the food consumption data, and it can be added to the corresponding consumption nodes. The consumption nodes can be added to complete the consumption statistics, and the adding method is to add backward, for example, the consumption data from 12:00 to 18:00 is added to the consumption node corresponding to 18:00; the adding method for accommodation consumption data is as follows: for example, on September 28, an order for accommodation rooms for three days and two nights from October 1 to October 3 has been placed on the platform. At this time, the accommodation consumption data of the corresponding date can be added to the corresponding date. Taking October 2 as an example, the accommodation expenses on October 2 are recorded in the consumption nodes corresponding to 0:00, 6:00, 12:00 and 18:00 respectively, that is, under normal circumstances, the corresponding accommodation consumption data on the same day are the same, unless there is a temporary room change.
[0079] Step S104: defining the consumption nodes that do not include accommodation consumption data or food consumption data as missing nodes among the consumption nodes, and defining the remaining consumption nodes as valid nodes.
[0080] When a consumption node does not contain accommodation consumption data or food consumption data, it means that there is missing data in the consumption node, that is, the missing data of the consumption node needs to be supplemented. At this time, missing nodes and valid nodes are defined to distinguish different consumption nodes for subsequent analysis.
[0081] Step S105: Determine candidate personnel according to the duration of cultural travel and valid nodes, and calculate the consumption similarity based on the accommodation consumption data and food consumption data in the valid nodes of the candidate personnel and the accommodation consumption data and food consumption data in the valid nodes of the current personnel.
[0082] The alternative personnel are those who may supplement the data of the missing nodes of the current user after analyzing the cultural and travel duration and valid nodes. The specific determination method can refer to steps S200-S203; the consumption similarity is a numerical value reflecting the similarity of the consumption habits of two persons. The larger the value, the more similar the consumption habits of the two persons, that is, the consumption data of one person can be used as a data reference for the consumption data of another person. The specific determination method of consumption similarity can refer to steps S400-S405.
[0083] Step S106: define candidate persons whose consumption similarity is greater than the preset demand similarity as similar persons, and generate predicted consumption data based on the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar persons, and fill the predicted consumption data into the missing nodes.
[0084] The demand similarity is the minimum consumption similarity set by the staff to be met when determining that the consumption habits of two persons are relatively similar. By defining similar persons, it is possible to distinguish between candidate persons with different consumption similarities; the data of the consumption nodes corresponding to the missing nodes in the similar persons can be used to predict the missing data of the current user, that is, predict the consumption data, so that the predicted consumption data can be filled into the missing nodes to improve the integrity of the current user data; the method for determining the predicted consumption data can be to randomly select a user to fill the corresponding data into the missing node of the current user, or to determine it by calculating the average value based on the data of all users, or to determine it by the method of steps S500-step S504.
[0085] The steps to determine the candidate candidates based on the duration of the cultural tour and the effective nodes include:
[0086] Step S200: define persons whose cultural travel duration is consistent with that of the current user as preliminary candidates.
[0087] Define preliminary candidates to identify those with the same playing time. Only then will it be easy for the two to refer to each other's consumption habits, so as to facilitate the subsequent determination of candidate candidates.
[0088] Step S201: Count the valid nodes of the current user to determine the valid quantity, and calculate based on the valid quantity and a preset ratio parameter to determine the required quantity.
[0089] The effective number is the total number of effective nodes of the current user. The proportion parameter is a fixed value parameter set by the staff. Generally, the proportion parameter needs to be greater than 90%. The required number is the minimum number of corresponding consumption nodes that are both effective nodes required between two users when determining similar consumption habits. It is determined by multiplying the effective number by the proportion parameter and rounding up.
[0090] Step S202: define the consumption nodes corresponding to the valid nodes of the current user in the preliminary selection personnel as analysis nodes, and count the analysis nodes that are valid nodes to determine the analysis quantity.
[0091] Analysis nodes are defined to distinguish different consumption nodes, and the corresponding analysis quantity is the total number of determined analysis nodes.
[0092] Step S203: defining the preliminary candidates whose analysis quantity is not less than the effective quantity as candidate candidates.
[0093] When the number of analyses is not less than the effective number, it means that the number of nodes available for analysis between the two persons meets the requirements, and they can be defined as candidate persons.
[0094] After the candidates are determined, the cultural tourism consumption data management method based on big data also includes:
[0095] Step S300: Count the candidate personnel to determine the number of candidates.
[0096] The number of candidates is the number of candidates determined.
[0097] Step S301: Determine whether the candidate quantity is greater than a preset clear quantity.
[0098] The clear number is the maximum number of candidates allowed by the staff. The purpose of the judgment is to find out whether there are too many candidates to affect data analysis.
[0099] Step S3011: If the number of candidates is not greater than the clear number, the currently determined candidate personnel are maintained.
[0100] When the number of alternatives is not greater than the clear number, it means that there are not many alternatives, and normal maintenance is sufficient.
[0101] Step S3012: If the number of candidates is greater than the clear number, count the valid nodes in each candidate to determine the number of valid candidates.
[0102] When the number of alternatives is greater than the number of clear candidates, it means that there are too many alternatives. If all of them are analyzed, some invalid operations will occur, which will affect the efficiency of data analysis. Therefore, further analysis is needed; the number of valid alternatives is the total number of valid nodes among the alternatives.
[0103] Step S302: Sort the candidate personnel from large to small according to the number of valid candidates to output the personnel selection order, and maintain the candidate personnel at the front of the personnel selection order according to the number of clear candidates, and define and cancel the remaining candidate personnel.
[0104] The order of personnel selection is the order obtained by sorting the candidates from front to back according to the number of valid candidates from large to small. The candidates at the front are maintained to ensure that as much missing data as possible can be filled in later. The candidates with the same number of valid candidates are randomly sorted.
[0105] The step of calculating the consumption similarity based on the accommodation consumption data and the food consumption data in the valid node of the candidate personnel and the accommodation consumption data and the food consumption data in the valid node of the current personnel includes:
[0106] Step S400: define the analysis node that is a valid node among the candidate personnel and the valid node of the current user as binding nodes.
[0107] The method of defining each other as binding nodes is, for example, if the current user's valid node is 18:00 on October 2, then when the consumption node of the candidate at 18:00 on October 2 is the analysis node, the two nodes can be defined as binding nodes to each other, that is, the corresponding consumption nodes for analyzing the consumption habits of the two users are identified and distinguished.
[0108] Step S401: performing a difference calculation in a binding node based on the accommodation consumption data to determine the accommodation consumption difference, and performing a difference calculation based on the food consumption data to determine the food consumption difference.
[0109] The accommodation consumption difference is the numerical difference between the accommodation consumption data obtained at two mutually bound consumption nodes, and the difference is an absolute value; similarly, the food consumption difference is the numerical difference between the food consumption data obtained at two mutually bound consumption nodes, and the difference is also an absolute value.
[0110] Step S402: performing calculations based on the accommodation consumption difference, the food consumption difference and preset calculation parameters to determine the node similarity.
[0111] The calculation parameters are fixed parameter values set by the staff for calculation. The node similarity is a value reflecting the similarity of consumption between two nodes. The calculation formula is: Where δ is the node similarity, M Z is the difference in accommodation consumption, M Y is the dietary consumption difference, α and β are the calculation parameters.
[0112] Step S403: randomly selecting a node similarity from all node similarities as the central similarity, and defining the remaining node similarities as comparative similarities.
[0113] Define center similarity and comparative similarity to distinguish different node similarities for subsequent analysis.
[0114] Step S404: Calculate the similarity interval according to the center similarity and all the comparison similarities, and define the center similarity corresponding to the similarity interval with the smallest value as the standard similarity.
[0115] The similarity interval is the absolute value of the difference between the central similarity and the comparative similarity. This value can reflect the distance between the central similarity and the other comparative similarities. The smaller the value, the more representative the currently determined central similarity is of the other comparative similarities. At this time, it is defined as the standard similarity to distinguish different central similarities, which is convenient for subsequent analysis.
[0116] Step S405: Calculate based on the standard similarity and the preset proximity to construct a valid interval, and calculate the average value based on the similarities of all nodes in the valid interval to determine the consumption similarity.
[0117] The similarity is the maximum difference allowed between the standard similarity and the other similarities that are determined to be closer to the standard similarity, set by the staff. The valid interval means the interval in which the similarities of the nodes that are closer to the standard similarity need to be, and the interval is constructed with the values obtained by adding and subtracting the similarity from the standard similarity as endpoints. Only by calculating the mean of all the node similarities in the valid interval can some abnormal data be eliminated, thereby improving the accuracy of consumer similarity determination.
[0118] The steps of generating predicted consumption data according to the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar persons include:
[0119] Step S500: define the missing data type in the missing node of the current user as a missing type, and define associated data in the valid nodes corresponding to similar persons according to the missing type.
[0120] By defining the missing type, it is possible to know whether the data missing in the missing node of the current user is food consumption data or accommodation consumption data; by defining the associated data, the data that can provide reference significance can be identified to facilitate the subsequent determination of the predicted consumption data.
[0121] Step S501: randomly generate a simulated consumption data, and perform a difference calculation based on the simulated consumption data and the associated data to determine the simulated consumption difference.
[0122] The simulated consumption difference is the difference between the simulated consumption data and the associated data, and the difference is an absolute value.
[0123] Step S502: determining a correlation value corresponding to the consumption similarity according to a preset similarity matching relationship.
[0124] The correlation degree value is a numerical value that reflects the correlation between two users. The higher the consumption similarity, the more similar the consumption habits of the two users are, that is, the stronger the correlation is, and the greater the correlation degree value is. The similar matching relationship between the two is determined by the staff through multiple experiments in advance and entered into storage.
[0125] Step S503: Calculate based on all simulated consumption differences and corresponding correlation degree values to determine the deviation degree value.
[0126] The deviation degree value is the value obtained by adding the simulated consumption difference value multiplied by the corresponding correlation degree value.
[0127] Step S504: determining the smallest deviation value according to a preset sorting rule, and defining the simulated consumption data corresponding to the deviation value as predicted consumption data.
[0128] The sorting rule is a method set by the staff to sort the numerical values, such as the bubble method. The sorting rule can be used to determine the minimum deviation value of the numerical value, that is, the deviation between the simulated consumption data determined at this time and the related data when used as predicted consumption data is the smallest. At this time, it can be defined as predicted consumption data.
[0129] After the deviation degree value is determined, the cultural tourism consumption data management method based on big data also includes:
[0130] Step S600: Determine whether there are at least two simulated consumption data with the same and minimum deviation values.
[0131] The purpose of the judgment is to find out whether there are multiple simulated consumption data that meet the requirements, so as to determine the only predicted consumption data.
[0132] Step S6001: If there are not at least two simulated consumption data with the same and smallest deviation degree values, the simulated consumption data corresponding to the smallest deviation degree value is determined as the predicted consumption data.
[0133] When there are not at least two simulated consumption data with the same and smallest deviation values, it means that there is only one simulated consumption data that meets the requirements, and it can be determined as the predicted consumption data.
[0134] Step S6002: If there are at least two simulated consumption data with the same and smallest deviation degree values, the simulated consumption data corresponding to the smallest deviation degree value is determined as the candidate consumption data.
[0135] When there are at least two simulated consumption data with the same and smallest deviation degree values, it means that there are multiple simulated consumption data that meet the requirements. At this time, they are defined as alternative consumption data for identification to facilitate subsequent analysis.
[0136] Step S601: defining data corresponding to the missing type in other consumption nodes of the current user as relevant consumption data according to the missing type.
[0137] Relevant consumption data is defined to identify the data corresponding to the missing type at other time points of the current user, so that the predicted consumption data can be further determined based on the user's own consumption habits.
[0138] Step S602: performing calculations based on the relevant consumption data to determine the representative coefficient of each relevant consumption data, and defining the relevant consumption data corresponding to the largest representative coefficient as the reference consumption data.
[0139] The representative coefficient is a parameter value that reflects the possibility that a single relevant consumption data represents the rest of the effect consumption data. The larger the value, the more representative it is. The method for determining the representative coefficient is as follows: subtract one relevant consumption data from the rest of the relevant consumption data and calculate the absolute value, and determine it by summing up all the absolute values and then calculating the inverse; define control consumption data to identify the data that best reflects the user's short-term consumption habits, for subsequent analysis.
[0140] Step S603: performing difference calculation based on the candidate consumption data and the reference consumption data to determine the interval consumption data, and defining the candidate consumption data corresponding to the interval consumption data with the smallest value as the predicted consumption data.
[0141] The interval consumption data is the difference between the alternative consumption data and the control consumption data. The difference is an absolute value. When the interval consumption data is the smallest, it means that the current corresponding alternative consumption data is closest to the user's other consumption habits in the short term. At this time, it can be determined as the predicted consumption data.
[0142] After the predicted consumption data is determined, the cultural tourism consumption data management method based on big data also includes:
[0143] Step S700: randomly selecting one associated data from all associated data to exclude and combining the remaining associated data to determine an associated set, and defining the associated set that is not in the associated set as excluded data.
[0144] The association set is a set of remaining association data obtained after excluding only one association data. Exclusion data is defined to identify the excluded association data for the convenience of subsequent analysis.
[0145] Step S701: Analyze the association set to determine predicted consumption data corresponding to the excluded data, and define the predicted consumption data as virtual consumption data.
[0146] The virtual consumption data is the data of the consumption node corresponding to the excluded data that can be predicted when the consumption data is determined by predicting the associated data in the associated set.
[0147] Step S702: Calculate based on the virtual consumption data and the corresponding exclusion data to determine the data deviation ratio.
[0148] The data deviation ratio is the deviation ratio between the virtual consumption data and the corresponding exclusion data, which is determined by subtracting the corresponding exclusion data from the virtual consumption data and dividing it by the exclusion data.
[0149] Step S703: Calculate the reasonable deviation ratio based on all data deviation ratios, and update the currently determined predicted consumption data based on the reasonable deviation ratio.
[0150] By calculating the mean of all data deviation ratios, we can obtain the deviation between the two values in the predicted situation and the actual situation, that is, the reasonable deviation ratio. At this time, the predicted consumption data can be updated through the reasonable deviation ratio and the currently determined predicted consumption data, thereby improving the accuracy of the predicted consumption data.
[0151] Reference Figure 2 Based on the same inventive concept, an embodiment of the present invention provides a cultural tourism consumption data management system based on big data, including:
[0152] Acquisition module, used to obtain user consumption data;
[0153] A processing module, connected to the acquisition module and the judgment module, for storing and processing information;
[0154] A judgment module, connected with the acquisition module and the processing module, for judging the information;
[0155] The processing module performs feature recognition in the user consumption data to respectively determine the transportation consumption data, the accommodation consumption data and the food consumption data;
[0156] The processing module determines the cultural travel consumption time period based on the transportation consumption data, and determines the cultural travel duration based on the cultural travel consumption time period;
[0157] The processing module generates consumption nodes according to preset consumption rules during the cultural and tourism consumption period, and adds each accommodation consumption data and food consumption data to the consumption node according to the corresponding time point;
[0158] The processing module defines the consumption nodes that do not contain accommodation consumption data or food consumption data in each consumption node as missing nodes, and defines the remaining consumption nodes as valid nodes;
[0159] The processing module determines candidate personnel according to the duration of cultural travel and effective nodes, and calculates the accommodation consumption data and food consumption data in the effective nodes of the candidate personnel and the accommodation consumption data and food consumption data in the effective nodes of the current personnel to determine the consumption similarity;
[0160] The processing module defines the candidate personnel whose consumption similarity determined by the judging module is greater than the preset demand similarity as similar personnel, and generates predicted consumption data according to the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar personnel, and fills the predicted consumption data into the missing nodes;
[0161] A candidate determination module is used to determine the candidate;
[0162] The candidate correction module is used to correct the candidate when there are a large number of candidates;
[0163] A consumption similarity determination module, used to determine the consumption similarity between two users;
[0164] A forecast consumption data determination module is used to determine relatively accurate forecast consumption data;
[0165] A simulated consumption data screening module is used to screen multiple simulated consumption data that meet the requirements;
[0166] The predicted consumption data correction module corrects the predicted consumption data according to the deviations in the prediction.
[0167] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device and unit described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
Claims
1. A method for managing cultural tourism consumption data based on big data, characterized in that: include: Obtain user consumption data; Perform feature recognition in the user consumption data to determine transportation consumption data, accommodation consumption data, and food consumption data respectively; Determine the time period for cultural and tourism consumption based on transportation consumption data, and determine the duration of cultural and tourism consumption based on the time period for cultural and tourism consumption; Generate consumption nodes according to preset consumption rules during the cultural and tourism consumption period, and add each accommodation consumption data and food consumption data to the consumption nodes according to the corresponding time points; In each consumption node, a consumption node that does not contain accommodation consumption data or food consumption data is defined as a missing node, and the remaining consumption nodes are defined as valid nodes; Determine candidate personnel based on the duration of cultural travel and effective nodes, and calculate the consumption similarity based on the accommodation consumption data and food consumption data in the effective nodes of the candidate personnel and the accommodation consumption data and food consumption data in the effective nodes of the current personnel; Alternative persons whose consumption similarity is greater than the preset demand similarity are defined as similar persons, and predicted consumption data are generated based on the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar persons, and the predicted consumption data are filled into the missing nodes.
2. The method for managing cultural and tourism consumption data based on big data according to claim 1 is characterized in that: The steps to determine the candidate candidates based on the duration of the cultural tour and the effective nodes include: Define the people whose cultural travel time is the same as that of the current user as the preliminary candidates; Count the valid nodes of the current user to determine the valid quantity, and calculate the required quantity based on the valid quantity and the preset proportion parameters; The consumption nodes corresponding to the valid nodes of the current user are defined as analysis nodes among the preliminary candidates, and the analysis nodes that are valid nodes are counted to determine the analysis quantity; The candidates whose analysis number is not less than the effective number are defined as candidates.
3. The method for managing cultural and tourism consumption data based on big data according to claim 2 is characterized in that: After the candidates are determined, the cultural tourism consumption data management method based on big data also includes: Counting the candidates to determine the number of candidates; Determine whether the number of alternatives is greater than the preset clear number; If the number of candidates is not greater than the clear number, the currently determined candidates are maintained; If the number of candidates is greater than the clear number, count the valid nodes in each candidate to determine the number of valid candidates; According to the number of valid candidates, the candidate personnel are sorted from large to small to output the personnel selection order, and according to the clear number, the candidate personnel at the front of the personnel selection order are maintained, and the remaining candidate personnel are defined and cancelled.
4. The method for managing cultural and tourism consumption data based on big data according to claim 2 is characterized in that: The step of calculating the consumption similarity based on the accommodation consumption data and the food consumption data in the valid node of the candidate personnel and the accommodation consumption data and the food consumption data in the valid node of the current personnel includes: The analysis node that is a valid node among the candidate personnel and the valid node of the current user are defined as binding nodes to each other; In the binding node, a difference calculation is performed based on the accommodation consumption data to determine the accommodation consumption difference, and a difference calculation is performed based on the food consumption data to determine the food consumption difference; Calculate the node similarity based on the difference in accommodation consumption, the difference in food consumption and the preset calculation parameters; A node similarity is randomly selected from all node similarities as the central similarity, and the remaining node similarities are defined as comparative similarities; The similarity interval is determined by calculating the center similarity and all the comparison similarities, and the center similarity corresponding to the similarity interval with the smallest value is defined as the standard similarity; The calculation is performed based on the standard similarity and the preset proximity to construct a valid interval, and the consumption similarity is determined by performing an average calculation based on the similarities of all nodes in the valid interval.
5. The method for managing cultural and tourism consumption data based on big data according to claim 1 is characterized in that: The steps of generating predicted consumption data according to the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar persons include: Define the missing data type in the missing node of the current user as the missing type, and define the associated data in the valid nodes corresponding to similar persons according to the missing type; A simulated consumption data is randomly generated, and a difference calculation is performed based on the simulated consumption data and the associated data to determine the simulated consumption difference; Determine the correlation value corresponding to the consumption similarity according to the preset similarity matching relationship; Calculate the deviation degree value based on all simulated consumption differences and the corresponding correlation degree values; The deviation degree value with the smallest numerical value is determined according to a preset sorting rule, and the simulated consumption data corresponding to the deviation degree value is defined as the predicted consumption data.
6. The method for managing cultural and tourism consumption data based on big data according to claim 5 is characterized in that: After the deviation degree value is determined, the cultural tourism consumption data management method based on big data also includes: Determine whether there are at least two simulated consumption data with the same and minimum deviation degree values; If there are not at least two simulated consumption data with the same and smallest deviation degree values, the simulated consumption data corresponding to the smallest deviation degree value is determined as the predicted consumption data; If there are at least two simulated consumption data with the same and smallest deviation degree values, the simulated consumption data corresponding to the smallest deviation degree value is determined as the candidate consumption data; According to the missing type, the data corresponding to the missing type in the remaining consumption nodes of the current user is defined as the relevant consumption data; Calculate according to each relevant consumption data to determine the representative coefficient of each relevant consumption data, and define the relevant consumption data corresponding to the largest representative coefficient as the control consumption data; The difference calculation is performed based on the alternative consumption data and the control consumption data to determine the interval consumption data, and the alternative consumption data corresponding to the interval consumption data with the smallest value is defined as the predicted consumption data.
7. The method for managing cultural and tourism consumption data based on big data according to claim 6 is characterized in that: After the predicted consumption data is determined, the cultural tourism consumption data management method based on big data also includes: Randomly select one associated data from all associated data to exclude and combine them according to the remaining associated data to determine an associated set, and define the associated set that is not in the associated set as excluded data; Analyze the association set to determine the predicted consumption data corresponding to the excluded data, and define the predicted consumption data as virtual consumption data; Calculate the data deviation ratio based on the virtual consumption data and the corresponding excluded data; A calculation is performed based on all data deviation ratios to determine a reasonable deviation ratio, and the currently determined predicted consumption data is updated based on the reasonable deviation ratio.
8. A cultural tourism consumption data management system based on big data, characterized in that: include: Acquisition module, used to obtain user consumption data; A processing module, connected to the acquisition module and the judgment module, for storing and processing information; A judgment module, connected with the acquisition module and the processing module, for judging the information; The processing module performs feature recognition in the user consumption data to respectively determine the transportation consumption data, the accommodation consumption data and the food consumption data; The processing module determines the cultural travel consumption time period based on the transportation consumption data, and determines the cultural travel duration based on the cultural travel consumption time period; The processing module generates consumption nodes according to preset consumption rules during the cultural and tourism consumption period, and adds each accommodation consumption data and food consumption data to the consumption node according to the corresponding time point; The processing module defines the consumption nodes that do not contain accommodation consumption data or food consumption data in each consumption node as missing nodes, and defines the remaining consumption nodes as valid nodes; The processing module determines candidate personnel according to the duration of cultural travel and effective nodes, and calculates the accommodation consumption data and food consumption data in the effective nodes of the candidate personnel and the accommodation consumption data and food consumption data in the effective nodes of the current personnel to determine the consumption similarity; The processing module defines the candidate persons whose consumption similarity determined by the judgment module is greater than the preset demand similarity as similar persons, and generates predicted consumption data based on the accommodation consumption data and food consumption data of the valid nodes corresponding to the missing nodes among the similar persons, and fills the predicted consumption data into the missing nodes.