Customer Interaction Data Retrieval Method Based on Dynamic Data Indexing
Through segmented aggregation analysis and adaptive multi-scale detection algorithm, combined with dynamic data index structure, the problem that traditional indexing strategies cannot take into account long-term and short-term characteristics is solved, and efficient and accurate retrieval of customer interaction data is achieved.
Patent Information
- Application Number
- CN202411835389.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-12-13
AI Technical Summary
Traditional dynamic indexing strategies cannot take into account the long-term stable characteristics of customer interaction data and the short-term burst characteristics, resulting in insufficient capture of burst information, affecting the accuracy and real-time nature of data retrieval.
The stationary feature baseline of long-term data is extracted through segmented aggregation analysis, and short-term burst feature detection algorithm is combined with the adaptive multi-scale burst feature detection algorithm, and potential risks are evaluated through composite feature weight distribution and weak co-occurrence probability analysis, and dynamic data index structure is updated in real time.
It realizes effective separation and processing of different time features in customer interaction data, improves the accuracy and real-timeness of retrieval, and is suitable for scene requirements with high frequency changes and complex characteristics.
Smart Images

Figure CN119782383B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data retrieval, and more specifically, to a method for retrieving customer interaction data based on dynamic data indexing. Background Art
[0002] With the continuous growth of the scale and complexity of customer interaction data, dynamic data indexing technology has been widely applied to the organization and retrieval of real-time data. In practical applications, customer interaction data usually has significant temporal characteristics, including both long-term stable characteristics that change gradually over time and short-term burst characteristics that may suddenly occur due to specific events or behaviors; long-term stable characteristics refer to the statistical characteristics that remain stable within a long time range in customer interaction data, such as the stable values of average access times or durations; short-term burst characteristics are abnormal changes caused by specific events or behaviors, such as a sharp increase in access volume or a sharp change in stay time within a short period. The two are independent of each other in temporal characteristics but coexist simultaneously.
[0003] Traditional dynamic indexing strategies are mainly optimized for single characteristics and cannot take into account the coexistence relationship of multiple characteristic types in data. This limitation is particularly obvious in the processing of data involving complex customer behavior patterns; in the prior art, it is difficult for dynamic indexing to simultaneously process data with both long-term stable characteristics and short-term burst characteristics, which will lead to insufficient capture of burst information and affect the accuracy and real-time performance of data retrieval. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method for retrieving customer interaction data based on dynamic data indexing to solve the problems proposed in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for retrieving customer interaction data based on dynamic data indexing includes the following steps:
[0007] S1: Perform segmented aggregation analysis on historical customer interaction data within a long term, and extract stable characteristic values in the long-term data as a stable feature baseline;
[0008] S2: Apply an adaptive multi-scale burst feature detection algorithm to recent customer interaction data, and identify irregular short-term burst features by adjusting the time window and data scale;
[0009] S3: Compare the short-term burst features with the stable feature baseline, and eliminate feature items that do not meet the requirements according to a preset quality threshold to form a composite feature weight distribution;
[0010] S4: Monitor the dynamic stability of the feature groups with low correlation in the index. Based on the weak co-occurrence probability analysis between features, evaluate the potential risk of weak-related features triggering sudden behaviors in subsequent data, and record the potential incentives.
[0011] S5: Update the dynamic data index structure in real time based on the composite feature weight distribution and potential incentives, and retrieve the newly input customer interaction data based on the real-time updated dynamic data index structure.
[0012] In a preferred embodiment, perform segmented aggregation analysis on the historical customer interaction data in the long term, and extract the stable feature values in the long-term data as the stable feature baseline, specifically including:
[0013] S101: Obtain the historical customer interaction data in the long term, divide the historical customer interaction data into multiple consecutive time periods according to a predetermined time interval, and mark the start and end times of each time period.
[0014] S102: Extract data features from the historical customer interaction data in each time period according to the preset feature indicators. The data features include the number of customer interaction behaviors and the duration.
[0015] S103: Perform statistical processing on the data features in all time periods, analyze the distribution of the data features in each time period, and calculate the time change range of each data feature.
[0016] S104: According to the time change range results of the statistical analysis, screen out the data features with a small time change range and mark them as stable feature candidate values.
[0017] S105: Perform clustering analysis on the stable feature candidate values, extract the central value of the clustering result as the stable feature baseline, and record it as the stable feature value of the long-term data.
[0018] In a preferred embodiment, apply an adaptive multi-scale burst feature detection algorithm to the recent customer interaction data, and identify irregular short-term burst features by adjusting the time window and data scale, specifically including:
[0019] S201: Obtain the recent customer interaction data, and divide the recent customer interaction data into multiple consecutive time windows according to the preset time range.
[0020] S202: For the recent customer interaction data in each time window, calculate the recent feature values according to the multi-scale feature extraction rules. The recent feature values include the interaction frequency and the behavior persistence.
[0021] S203: Dynamically adjust the time window length and data scale parameters, calculate the multi-scale feature values under each time window respectively, and form a multi-scale feature matrix.
[0022] S204: Analyze the change patterns of the multi-scale feature matrix, screen and mark the features with relatively large mutation amplitudes as short-term burst features.
[0023] In a preferred embodiment, compare the short-term burst features with the stable feature baseline, and eliminate the feature items that do not meet the requirements according to the preset quality threshold to form a composite feature weight distribution, specifically including:
[0024] S301: Obtain the short-term burst feature set and the stable feature baseline set, and perform unified formatting processing;
[0025] S302: Pair the short-term burst features and the stable feature baseline item by item according to the feature type and time window to form a set of comparable feature pairs;
[0026] S303: For each group of feature pairs in the set of feature pairs, calculate the feature difference degree between the short-term burst feature and the stable feature baseline;
[0027] S304: Compare the feature difference degree with the preset quality threshold, eliminate the feature items whose difference degree exceeds the threshold range, and only retain the feature items that meet the requirements;
[0028] S305: Calculate the weights for the retained feature items according to the feature difference degree to generate a composite feature weight distribution.
[0029] In a preferred embodiment, perform dynamic stability monitoring on the feature groups with low correlation degrees in the index, analyze based on the weak co-occurrence probability between features, evaluate the potential risk of weak-related features triggering burst behaviors in subsequent data, and record potential incentives, specifically including:
[0030] S401: Extract the feature pairs with lower weights from the composite feature weight distribution to form a set of low-correlation feature groups;
[0031] S402: Calculate the co-occurrence probability between features based on the feature items in the set of low-correlation feature groups to obtain the co-occurrence matrix between each pair of features;
[0032] S403: Analyze the potential associations between low-correlation features according to the weak co-occurrence probability in the co-occurrence matrix, and identify potential incentives;
[0033] S404: Perform dynamic monitoring on the potential incentives in the low-correlation feature groups, track the trend of weak changes between features, and evaluate the potential risk of weak-related features triggering burst behaviors in subsequent data.
[0034] In a preferred embodiment, calculate the co-occurrence probability between features based on the feature items in the set of low-correlation feature groups to obtain the co-occurrence matrix between each pair of features, specifically as follows:
[0035] Count the co-occurrence times of the two features in the statistical feature pair within all historical time windows, and count the total occurrence times of each individual feature in the statistical feature pair;
[0036] Calculate the co-occurrence probability using the co-occurrence frequency and the total occurrence times. The formula is: where P q,co represents the co-occurrence probability of the feature pair, N1 represents the total occurrence times of the first feature in the feature pair, N2 represents the total occurrence times of the second feature in the feature pair, Min(N1, N2) represents the smaller value of the total occurrence times of the two features in the feature pair, and C q represents the co-occurrence times of the two features in the feature pair in the historical data;
[0037] The calculation result is represented in the form of a co-occurrence matrix: M = {P ab,co | a, b ∈ G}; where M represents the co-occurrence matrix of the low-correlation feature group, P ab,co represents the co-occurrence probability of the feature pair (a, b), a and b represent any two features in the low-correlation feature group, and G represents the set of the low-correlation feature group.
[0038] In a preferred embodiment, according to the weak co-occurrence probability in the co-occurrence matrix, analyze the potential associations between the low-correlation features and identify potential incentives. Specifically:
[0039] Set a co-occurrence probability threshold, extract the feature pairs that satisfy the co-occurrence probability of the feature pair being less than the co-occurrence probability threshold and mark them as low co-occurrence probability feature pairs;
[0040] For each group of low co-occurrence probability feature pairs, combine the eigenvalue fluctuations within the time window to determine whether there are potential incentives;
[0041] The specific potential incentive identification rule is: If the co-occurrence probability of the feature pair is less than the co-occurrence probability threshold, and the fluctuation patterns of the two features in the feature pair show a correlated trend, then mark it as a potential incentive.
[0042] In a preferred embodiment, based on the composite feature weight distribution and potential incentives, perform real-time updates on the dynamic data index structure, and retrieve the newly input customer interaction data based on the real-time updated dynamic data index structure. Specifically include:
[0043] S501: Extract the weight of each feature pair and the potential risk probability of the burst behavior from the composite feature weight distribution to generate an updated data set;
[0044] S502: Adjust the priority of the feature nodes according to the weight of the feature pair in the updated data set, and redefine the association strength between the feature pairs in combination with the potential risk probability of the burst behavior to optimize the index structure;
[0045] S503: Clean the feature nodes with low weights in the dynamic data index structure, and at the same time sort the feature nodes with high weights to improve the retrieval efficiency and the compactness of the structure;
[0046] S504: Use the dynamically updated dynamic data index structure to retrieve the newly input customer interaction data, match the high-weight feature pairs associated with the input data, and return the retrieval results.
[0047] Technical effects and advantages of the customer interaction data retrieval method based on dynamic data index of the present invention:
[0048] 1. By performing piecewise aggregation analysis to extract the stable feature baselines of long-term data, and combining with the adaptive multi-scale burst feature detection algorithm to accurately identify short-term burst features, the effective separation and comprehensive processing of different time features in customer interaction data are realized. It can identify, compare and fuse short-term burst features and long-term stable features respectively, overcoming the problem that traditional dynamic indexing strategies cannot take into account multiple feature patterns, and providing an efficient feature basis for the optimization of the subsequent index structure.
[0049] 2. By dynamically monitoring the stability of low-correlation feature groups and combining the weak co-occurrence probability analysis between features, the risk of burst behaviors that may be caused by weakly related features can be evaluated in advance. Further, through the real-time update and efficient retrieval of the dynamic data index structure, the rapid capture and processing of burst information and potential incentives in dynamic data are ensured. This index optimization strategy based on feature weights and risk probabilities significantly improves the accuracy and real-time performance of retrieval, effectively making up for the deficiencies of traditional methods in processing complex customer interaction data, and is applicable to the scenario requirements of high-frequency changes and complex features. Brief Description of the Drawings
[0050] Figure 1 It is a schematic diagram of the customer interaction data retrieval method based on dynamic data index of the present invention. Detailed Embodiments
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0052] Embodiment: Figure 1 The customer interaction data retrieval method based on dynamic data index of the present invention is given, which includes the following steps:
[0053] S1: Perform piecewise aggregation analysis on historical customer interaction data over a long term, and extract stable eigenvalue in the long-term data as the stationary feature baseline.
[0054] S2: Apply an adaptive multi-scale burst feature detection algorithm to recent customer interaction data, and identify irregular short-term burst features by adjusting the time window and data scale.
[0055] S3: Compare the short-term burst features with the stationary feature baseline, and eliminate the feature items that do not meet the requirements according to the preset quality threshold to form a composite feature weight distribution.
[0056] S4: Perform dynamic stability monitoring on the feature groups with low correlation in the index, and evaluate the potential risk of weak-related features triggering burst behavior in subsequent data based on the weak co-occurrence probability analysis between features, and record the potential incentives.
[0057] S5: Update the dynamic data index structure in real time based on the composite feature weight distribution and potential incentives, and retrieve the newly input customer interaction data based on the dynamically updated dynamic data index structure.
[0058] Perform piecewise aggregation analysis on historical customer interaction data over a long term, and extract stable eigenvalue in the long-term data as the stationary feature baseline, specifically including:
[0059] S101: Obtain historical customer interaction data over a long term, divide the historical customer interaction data into multiple consecutive time periods according to a predetermined time interval, and mark the start and end times of each time period:
[0060] For example, if the time interval is one day, the daily data corresponds to one time period; if the time interval is one week, the weekly data corresponds to one time period. The time period division rule should ensure that the time periods are consecutive and non-overlapping.
[0061] S102: Extract data features from the historical customer interaction data in each time period according to the preset feature indicators. The data features include the number of customer interaction behaviors and the duration:
[0062] After completing the time period division, extract features from the historical customer interaction data in each time period. Feature extraction needs to be completed based on the preset indicators, specifically including:
[0063] The number of customer interaction behaviors: Calculate the total number of all customer interaction behaviors in a time period as the basic indicator to measure the customer activity level.
[0064] The duration of customer interaction behaviors: Count the total duration of all customer interaction behaviors in a time period as the indicator to measure the depth of customer interaction participation.
[0065] By parsing the specific content of each interaction data, the eigenvalue of each time period is gradually calculated.
[0066] S103: Statistically process the data features within all time periods, analyze the distribution of data features in each time period, and calculate the time variation range of each data feature:
[0067] Statistically process the eigenvalues extracted from each time period, analyze the distribution of eigenvalues within all time periods, and calculate the time variation range of each feature. The specific analysis includes the following steps: statistically analyze the eigenvalue distribution within each time period, judge the central tendency and dispersion degree of eigenvalues; calculate the fluctuation range of each eigenvalue in different time periods to evaluate the stability of the feature in the time dimension; overall evaluate the eigenvalues within all time periods and record the change amplitude of each feature.
[0068] S104: According to the results of the time variation range obtained from the statistical analysis, screen out the data features with a smaller time variation range and mark them as candidate values for stable features:
[0069] According to the results of the statistical analysis, screen the data features. The screening rules need to be set in combination with the variation range of eigenvalues. For example: only screen out the eigenvalues with a smaller time variation range and mark them as candidate values for stable features; exclude the eigenvalues with a larger time variation range to ensure the stability of candidate features.
[0070] Among them, the specific screening conditions for the time variation range should be set according to business requirements and data characteristics, such as determining by setting a predefined fluctuation threshold.
[0071] S105: Perform clustering analysis on the candidate values for stable features, extract the central value of the clustering result as the stable feature baseline, and record it as the stable feature value of the long-term data:
[0072] Among them, the clustering analysis process includes the following steps:
[0073] Divide the candidate values for stable features into several categories, and calculate the clustering center according to the distribution of feature values in each category;
[0074] The clustering center is the stable feature baseline, which is used to represent the stable feature value of the data in the long-term range.
[0075] The specific method of clustering analysis can adopt an algorithm suitable for business requirements, such as the average clustering method based on eigenvalue distribution or other clustering tools with high computational capabilities. Finally, the stable feature baseline is recorded as the stable feature value of the long-term data.
[0076] Apply the adaptive multi-scale burst feature detection algorithm to the recent customer interaction data, and identify irregular short-term burst features by adjusting the time window and data scale, specifically including:
[0077] S201: Obtain recent customer interaction data, and divide the recent customer interaction data into multiple consecutive time windows according to a preset time range:
[0078] Among them, each time window has a clear start time and end time, and the length of the time window is determined by the actual analysis requirements.
[0079] S202: For the recent customer interaction data of each time window, calculate recent feature values according to the multi-scale feature extraction rules. The recent feature values include interaction frequency and behavior persistence:
[0080] The multi-scale feature extraction rules are as follows:
[0081] First, select different time scale parameters, such as a small time period (e.g., 10 minutes) and a large time period (e.g., 1 hour), and extract features on different time scales.
[0082] Then, calculate the average feature value and the change feature value within the time scale to describe the distribution law of customer behavior at short-term and long-term scales.
[0083] Finally, combine the data of different scales to form a comprehensive multi-scale feature representation.
[0084] Among them, the expression of the interaction frequency is: The expression of the behavior persistence is: Among them, F i,k represents the interaction frequency of the i-th time window under the k-th scale parameter, D i,k represents the behavior persistence of the i-th time window under the k-th scale parameter, n i,k represents the total number of interaction behaviors of the i-th time window under the k-th scale parameter, S k represents the k-th scale parameter, j represents the number of the j-th interaction behavior within a certain time window, e j represents the end time of the j-th interaction behavior, s j represents the start time of the j-th interaction behavior.
[0085] Among them, the 1 in the expression of the interaction frequency means that each interaction behavior contributes one count, which is used to count the frequency of interaction behaviors.
[0086] The comprehensive multi-scale feature representation includes two dimensions: interaction frequency and behavior persistence. Within each time window, a feature set is generated through multi-scale feature extraction, which is respectively represented as an interaction frequency feature set and a behavior persistence feature set; the interaction frequency feature set contains interaction frequency values at different scales, while the behavior persistence feature set contains behavior duration values at different scales; these two sets together constitute the comprehensive feature representation of each time window, which is used to describe the complete feature distribution of customer interaction behavior at multiple scales.
[0087] S203: Dynamically adjust the time window length and data scale parameters, and calculate the multi-scale feature values for each time window respectively to form a multi-scale feature matrix:
[0088] Adjustment of the time window length: The time window length is controlled by a parameter sequence, and data features with different time granularities are generated by setting different time intervals.
[0089] Adjustment of the data scale parameters: Generate the columns of the eigenvalue matrix at different scales by changing the data scale parameters.
[0090] Calculate the eigenvalue of the i-th time window under the k-th scale parameter, and the expression is as follows: M ik = f(W i , S k ); where, M ik represents the eigenvalue of the i-th time window under the k-th scale parameter, W i represents the i-th time window, and f(W i , S k ) represents the feature extraction function (determined jointly by the time window and the data scale).
[0091] S204: Analyze the change pattern of the multi-scale feature matrix, screen and mark the features with a large mutation amplitude as short-term burst features:
[0092] Calculate the change rate of the eigenvalues under adjacent time windows, and judge whether there is a mutation; screen the eigenvalues whose change rate exceeds their corresponding preset thresholds, and mark the qualified eigenvalues as short-term burst features.
[0093] Among them, the preset threshold corresponding to the change rate is a fixed value set according to specific business requirements and data characteristics, which is used to judge whether the change rate of the eigenvalue is abnormal. Each threshold corresponds to a specific time window and scale parameter, indicating the minimum amplitude at which the feature change is considered significant under this condition, and is used to screen the mutant eigenvalue.
[0094] Record the selected short-term burst features, and mark their corresponding time windows and related behavior data. For example, record the time window number, change rate and associated behavior features of the burst feature for subsequent use.
[0095] Compare the short-term burst features with the stable feature baseline, and eliminate the feature items that do not meet the requirements according to the preset quality threshold to form a composite feature weight distribution, specifically including:
[0096] S301: Obtain the short-term burst feature set and the stable feature baseline set, and perform unified formatting processing to ensure that the feature data can be directly compared:
[0097] Obtain the short-term burst feature set and the stable feature baseline set respectively. The feature set includes different types of data features (such as interaction frequency, behavior persistence, etc.). To ensure the comparability between feature sets, formatting processing of the feature sets is required. The formatting processing includes:
[0098] Normalize all feature values to the same numerical range, for example, standardize to [0, 1].
[0099] Classify and sort the feature set according to the feature type to ensure that features of the same type can be directly compared.
[0100] Align the feature sets in time according to the time window to ensure that there is a corresponding relationship between the short-term burst features and the stable feature baseline at the same time point.
[0101] S302: Pair the short-term burst features with the stable feature baseline item by item according to the feature type and time window to form a set of comparable feature pairs:
[0102] If the type of the short-term burst feature is the same as that of the stable feature baseline and both belong to the same time window, pair them; if the short-term burst feature meets the conditions with multiple stable feature baselines, select the baseline with the closest time distance for pairing.
[0103] S303: For each group of feature pairs in the feature pair set, calculate the feature difference degree between the short-term burst feature and the stable feature baseline:
[0104] Calculate the absolute value after taking the difference between the short-term burst feature and the stable feature baseline in the feature pair to obtain the feature difference degree.
[0105] Among them, the calculation of the feature difference degree needs to be carried out one by one for the feature pairs in all feature pair sets to generate a complete difference degree set.
[0106] S304: Compare the feature difference degree with the preset quality threshold, eliminate the feature items whose difference degree exceeds the threshold range, and only retain the feature items that meet the requirements:
[0107] If the feature difference degree is less than or equal to the preset quality threshold, retain the feature pair; if the feature difference degree is greater than the preset quality threshold, eliminate the feature pair.
[0108] Among them, the preset quality threshold is a fixed value set according to specific business requirements and characteristic attributes, used to determine whether the difference between the short-term burst characteristics and the stable characteristic baseline meets the requirements. Each characteristic type can correspond to a different quality threshold to measure the acceptable range of the characteristic difference and ensure the rationality and accuracy of the screening results.
[0109] S305: Calculate the weights for the retained feature items according to the feature difference degree to generate a composite feature weight distribution:
[0110] The weight of a feature pair is the reciprocal of the feature difference degree corresponding to the feature pair.
[0111] Integrate the weights of all feature pairs to form a composite feature weight distribution, and the composite feature weight distribution includes the weights of all feature pairs.
[0112] Conduct dynamic stability monitoring on the feature groups with low correlation degrees in the index. Based on the analysis of the weak co-occurrence probability between features, evaluate the potential risk of weak-related features triggering burst behaviors in subsequent data, and record the potential incentives, specifically including:
[0113] S401: Extract the feature pairs with lower weights from the composite feature weight distribution to form a set of low-correlation feature groups:
[0114] In the composite feature weight distribution, each feature pair has a corresponding weight value. The weight value is used to reflect the association strength between feature pairs, and a lower weight value represents a weaker association degree between the two features. First, extract the feature pairs with weight values lower than the preset threshold from the composite feature weight distribution to form a set of low-correlation feature groups.
[0115] The extraction rule is as follows:
[0116] Set the weight threshold and extract the set of feature pairs that meet the condition K q <T K of the feature pairs.
[0117] The extraction result is expressed as: G = {P q |K q <T K}; where G represents the set of low-correlation feature groups, P q represents the qth feature pair, K q represents the weight of the qth feature pair, and T K represents the weight threshold.
[0118] Through the above rules, ensure that only the feature pairs with lower weights are included in the subsequent analysis scope.
[0119] Among them, T KIt is a fixed value set according to specific business scenarios and feature association strengths, used to distinguish low - association - degree features from high - association - degree features. When the weight of a feature pair is lower than this threshold, it is considered to have a lower association degree and needs to be included in the low - association - degree feature group for further analysis and monitoring.
[0120] S402: Calculate the co - occurrence probability between features based on the feature items in the low - association - degree feature group set, and obtain the co - occurrence matrix between each pair of features:
[0121] In the low - association - degree feature group set, count the historical data of each group of feature pairs, calculate the co - occurrence probability of the feature pairs in the historical data. The co - occurrence probability is used to quantify the frequency of two features appearing simultaneously. The specific process is as follows:
[0122] Co - occurrence frequency statistics: Count the number of co - occurrences of the two features in the feature pair within all historical time windows.
[0123] Total occurrence times statistics: Count the total number of occurrences of each individual feature in the feature pair.
[0124] Calculate the co - occurrence probability: Calculate the co - occurrence probability using the co - occurrence frequency and the total number of occurrences. The formula is: where, P q,co represents the co - occurrence probability of the feature pair, used to quantify the frequency of the two features in the feature pair appearing simultaneously; N1 represents the total number of occurrences of the first feature in the feature pair; N2 represents the total number of occurrences of the second feature in the feature pair; Min(N1, N2) represents the smaller value of the total number of occurrences of the two features in the feature pair, used to standardize the co - occurrence probability; C q represents the number of co - occurrences of the two features in the feature pair in the historical data.
[0125] The calculation result is represented in the form of a co - occurrence matrix: M = {P ab,co |a, b ∈ G}; where, M represents the co - occurrence matrix of the low - association - degree feature group; P ab,co represents the co - occurrence probability of the feature pair (a, b), used to quantify the co - occurrence frequency of feature a and feature b in the historical data; a and b represent any two features in the low - association - degree feature group.
[0126] S403: Analyze the potential associations between low - association - degree features based on the weak co - occurrence probabilities in the co - occurrence matrix, and identify potential incentives:
[0127] Analyze each co - occurrence probability in the co - occurrence matrix, and combine with the historical distribution of the feature pairs to identify the potential associations between low - association - degree features. The steps of potential association analysis are as follows:
[0128] Set the co - occurrence probability threshold, extract the feature pairs that satisfy the co - occurrence probability of the feature pair being less than the co - occurrence probability threshold, and mark them as low - co - occurrence - probability feature pairs.
[0129] For each pair of low co-occurrence probability features, combined with the eigenvalue fluctuations within the time window, determine whether there are potential incentives.
[0130] The specific potential incentive identification rule is: if the co-occurrence probability of the feature pair is less than the co-occurrence probability threshold, and the fluctuation patterns of the two features in the feature pair show a correlated trend, then it is marked as a potential incentive.
[0131] Among them, the co-occurrence probability threshold is an important criterion for screening feature pairs. It is set according to the historical co-occurrence characteristics of the feature pairs to distinguish significant co-occurrence from weak co-occurrence. When the co-occurrence probability of a feature pair is lower than this threshold, it is considered that their co-occurrence association is weak and further analysis is required.
[0132] Use the following formula to calculate the correlated trend of the eigenvalues: where D q represents the correlated trend coefficient of the q-th feature pair, used to quantify the linear correlation degree between the two eigenvalues; X t and Y t respectively represent the eigenvalues of the two features in the feature pair at time t; and respectively represent the means of the two eigenvalues.
[0133] Among them, the eigenvalue is a numerical value used to quantify the performance of a feature at time t, usually including interaction frequency, behavior duration, etc.
[0134] According to the correlated trend coefficient, judge whether the fluctuation patterns of the two features in the feature pair show a correlated trend through its value range. When the correlated trend coefficient is greater than 0, the fluctuations of the two features are positively correlated, that is, the eigenvalues rise or fall synchronously with time; when the correlated trend coefficient is less than 0, the fluctuations of the two features are negatively correlated, that is, when one eigenvalue rises, the other eigenvalue falls; when the correlated trend coefficient is equal to 0, the two features have no correlation and the fluctuation patterns do not affect each other.
[0135] Set a positive correlation threshold and a negative correlation threshold. The positive correlation threshold and the negative correlation threshold are important indicators for judging the correlated trend of feature pairs. The positive correlation threshold represents the lowest coefficient value at which the positive correlation of the two features is significant, and the negative correlation threshold represents the highest coefficient value at which the negative correlation of the two features is significant. The two are jointly used to screen feature pairs with a significant correlated trend.
[0136] If the correlated trend coefficient is greater than or equal to the positive correlation threshold or the correlated trend coefficient is less than or equal to the negative correlation threshold, then it is considered that the fluctuation patterns of the two features in the feature pair show a correlated trend. The judgment result of the correlated trend can be used for further analysis of potential feature correlations or risk assessments.
[0137] S404: Dynamically monitor the potential incentives in the low-correlation feature group, track the trend of weak changes between features, and evaluate the potential risk that weakly correlated features may trigger sudden behaviors in subsequent data:
[0138] For the identified potential incentives, implement dynamic monitoring, track the change trend between features, and calculate the potential risk probability of triggering sudden behaviors in the future based on the feature fluctuations within the time window.
[0139] For each group of potential incentive feature pairs, continuously track the change of their feature values and record the fluctuation amplitude within different time windows.
[0140] Based on the fluctuation amplitude and time weight, calculate the potential risk probability of sudden behaviors. The potential risk probability of sudden behaviors is the product of the co-occurrence probability of the feature pair and the time weight.
[0141] Among them, the time weight is a dynamic parameter set according to the timeliness of the time window, used to reflect the impact of the latest data on feature analysis. The weight is usually calculated through a predefined decay function. The weight of the window closer to the current time is higher. Common methods include exponential decay or linear decay models.
[0142] The greater the potential risk probability of sudden behaviors, the greater the potential risk that weakly correlated features may trigger sudden behaviors in subsequent data, indicating that the co-occurrence relationship and time trend between weakly correlated features are more likely to trigger abnormal fluctuations. This shows that in subsequent data, the possibility of these feature pairs having sudden behaviors is higher, and they need to be key monitored and evaluated.
[0143] Based on the composite feature weight distribution and potential incentives, real-time update the dynamic data index structure. Retrieve the newly input customer interaction data based on the real-time updated dynamic data index structure, specifically including:
[0144] S501: Extract the weight of each feature pair and the potential risk probability of sudden behaviors from the composite feature weight distribution to generate an updated data set:
[0145] From the composite feature weight distribution, extract the weight of each feature pair one by one, as well as the potential risk probability of sudden behaviors obtained from the potential incentive analysis. The feature weight value represents the association strength of the feature pair, and the risk probability is used to quantify the risk that the feature pair may trigger sudden behaviors.
[0146] Combine the extracted feature weight values and risk probabilities according to the feature pairs to form an updated data set for adjusting the dynamic data index structure.
[0147] S502: Adjust the priority of the feature nodes according to the weight of the feature pairs in the updated data set, and redefine the association strength between feature pairs in combination with the potential risk probability of sudden behaviors to optimize the index structure:
[0148] According to the weight of each feature pair, the priority of the corresponding feature node in the index structure is dynamically adjusted. The feature node with a higher weight value has a higher priority in the index structure.
[0149] Adjustment rules: Increase the priority index level of high-weight feature nodes in the index structure to ensure that they participate in retrieval matching faster; lower the priority of low-weight feature nodes to reduce the resource usage of the index structure.
[0150] Based on the potential risk probability of the sudden behavior of each feature pair, the association strength between feature nodes is redefined. For feature pairs with higher potential risk probability of sudden behavior, their association strength in the index structure is strengthened so that they can be monitored and processed first.
[0151] S503: Clean up the low-weight feature nodes in the dynamic data index structure, and sort the high-weight feature nodes to improve retrieval efficiency and structure compactness:
[0152] Perform cleaning operations on feature nodes whose weight values are lower than the preset threshold in the dynamic data index structure: if the weight of a feature pair is less than the preset cleaning threshold, remove the feature node from the index structure; retain feature nodes whose feature pair weights are greater than or equal to the preset cleaning threshold, ensuring that only important features are retained in the index structure.
[0153] For the cleaned index structure, the remaining feature nodes are sorted in descending order using the weights of the feature pairs in the updated data set. The sorting rules are as follows: sort from high to low according to the weights of the feature pairs; feature nodes with high priority occupy high positions in the index structure to improve index retrieval efficiency.
[0154] After the cleaning and sorting operations, the resource positions in the index structure are reallocated to ensure that high-weight feature nodes occupy key positions while reducing the resource consumption of low-priority nodes to form a more compact index structure.
[0155] S504: Use the real-time updated dynamic data index structure to search for the newly input customer interaction data, match the high-weight feature pairs associated with the input data, and return the search results:
[0156] Preprocess the newly input customer interaction data and extract feature values according to the feature rules in the dynamic data index structure; for example, the newly input data may contain user operation records (such as click frequency, visit duration, etc.). By parsing these records, feature values such as "number of visits is 5 times" and "stay time is 120 seconds" can be extracted.
[0157] The extracted eigenvalue is matched one by one with the sorted high-weight feature nodes in the dynamic data index structure, and the matching rule is based on the correlation between the eigenvalue and the high-priority features in the index.
[0158] For example, a high-weight feature node in the index structure may record the feature of "the number of clicks is greater than 3 times and the stay time exceeds 100 seconds"; if the newly input data contains the eigenvalue of "the number of clicks is 5 times" and "the stay time is 120 seconds", it will match successfully with this node.
[0159] For all the feature nodes that match successfully, they are sorted in descending order according to their weight values, and the most relevant feature nodes are selected and returned as part of the retrieval result.
[0160] For example, if the newly input data matches two feature nodes (such as node A and node B) at the same time, where the weight of node A is 0.8 and the weight of node B is 0.6, then node A is returned to the retrieval result first; the final returned retrieval result can include the matching feature node numbers and their associated context information (such as feature descriptions, time ranges, etc.).
[0161] The returned retrieval result is provided to the subsequent analysis module to support user behavior pattern analysis, anomaly detection or other business logics. The format of the retrieval result may include information such as the matching feature node numbers, relevant descriptions, and matching eigenvalues.
[0162] The above formulas are all dimensionless and take their numerical calculations. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters and threshold selection in the formula are set by those skilled in the art according to the actual situation.
[0163] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0164] Those of ordinary skill in the art will realize that the modules and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0165] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0166] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.
[0167] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module. It may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0168] In addition, in each embodiment of this application, each functional module can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0169] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0170] As described above, this is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0171] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for retrieving customer interaction data based on dynamic data indexing, characterized in that, It includes the following steps: S1: Perform segmented aggregation analysis on historical customer interaction data in the long term, and extract stable eigenvalue in the long-term data as the stationary feature baseline; S2: Apply an adaptive multi-scale burst feature detection algorithm to recent customer interaction data, and identify irregular short-term burst features by adjusting the time window and data scale; S3: Compare the short-term burst features with the stationary feature baseline, and eliminate feature items that do not meet the requirements according to a preset quality threshold to form a composite feature weight distribution, specifically including: S301: Obtain the short-term burst feature set and the stationary feature baseline set, and perform unified formatting processing; S302: Pair the short-term burst features with the stationary feature baseline item by item according to the feature type and time window to form a set of comparable feature pairs; S303: Calculate the feature difference degree between the short-term burst feature and the stationary feature baseline for each group of feature pairs in the feature pair set; S304: Compare the feature difference degree with the preset quality threshold, eliminate the feature items whose difference degree exceeds the threshold range, and only retain the feature items that meet the requirements; S305: Calculate the weights for the retained feature items according to the feature difference degree to generate a composite feature weight distribution; S4: Perform dynamic stability monitoring on the feature groups with low correlation in the index, and evaluate the potential risk of weak-related features triggering burst behavior in subsequent data based on the weak co-occurrence probability analysis between features, and record potential incentives, specifically including: S401: Extract feature pairs with lower weights from the composite feature weight distribution to form a set of low-correlation feature groups; S402: Calculate the co-occurrence probability between features based on the feature items in the low-correlation feature group set to obtain a co-occurrence matrix for each pair of features; S403: Analyze the potential association between low-correlation features according to the weak co-occurrence probability in the co-occurrence matrix, and identify potential incentives; S404: Perform dynamic monitoring on the potential incentives in the low-correlation feature group, track the weak change trend between features, and evaluate the potential risk of weak-related features triggering burst behavior in subsequent data; S5: Real-time update the dynamic data index structure based on the composite feature weight distribution and potential incentives, and retrieve the newly input customer interaction data based on the real-time updated dynamic data index structure.
2. The method for retrieving customer interaction data based on dynamic data indexing according to claim 1, wherein Perform segmented aggregation analysis on historical customer interaction data in the long term, and extract stable eigenvalue in the long-term data as the stationary feature baseline, specifically including: S101: Obtain historical customer interaction data in the long term, divide the historical customer interaction data into multiple consecutive time periods according to a predetermined time interval, and mark the start and end times of each time period; S102: Extract data features from the historical customer interaction data in each time period according to preset feature indicators. The data features include the number of customer interaction behaviors and the duration; S103: Perform statistical processing on the data features in all time periods, analyze the distribution of data features in each time period, and calculate the time change range of each data feature; S104: According to the time change range results of the statistical analysis, screen out the data features with a smaller time change range and mark them as stationary feature candidate values; S105: Perform clustering analysis on the stable feature candidate values, extract the central value of the clustering result as the stable feature baseline, and record it as the stable feature value of the long-term data.
3. The method for retrieving customer interaction data based on dynamic data indexing according to claim 2, wherein Apply the adaptive multi-scale burst feature detection algorithm to the recent customer interaction data, and identify irregular short-term burst features by adjusting the time window and data scale, specifically including: S201: Obtain the recent customer interaction data, and divide the recent customer interaction data into multiple consecutive time windows according to the preset time range; S202: For the recent customer interaction data of each time window, calculate the recent feature values according to the multi-scale feature extraction rules. The recent feature values include interaction frequency and behavior persistence; S203: Dynamically adjust the time window length and data scale parameters, and calculate the multi-scale feature values under each time window respectively to form a multi-scale feature matrix; S204: Analyze the change pattern of the multi-scale feature matrix, and screen and mark the features with large mutation amplitudes as short-term burst features.
4. The method for retrieving customer interaction data based on dynamic data indexing according to claim 3, wherein Calculate the co-occurrence probability between features based on the feature items in the low-correlation feature group set to obtain the co-occurrence matrix between each pair of features, specifically: Count the co-occurrence times of the two features in the feature pair in all historical time windows, and count the total occurrence times of each individual feature in the feature pair; Calculate the co-occurrence probability using the co-occurrence frequency and the total number of occurrences. The formula is as follows: ; where represents the co-occurrence probability of the feature pair, represents the total number of occurrences of the first feature in the feature pair, represents the total number of occurrences of the second feature in the feature pair, represents the smaller value of the total number of occurrences of the two features in the feature pair, represents the number of co-occurrences of the two features in the feature pair in the historical data; The calculation results are represented in the form of a co-occurrence matrix: ; where represents the co-occurrence matrix of the low-correlation feature group, represents the co-occurrence probability of the feature pair , and represent any two features in the low-correlation feature group, represents the set of low-correlation feature groups.
5. The customer interaction data retrieval method based on dynamic data indexing according to claim 4, wherein, According to the weak co-occurrence probability in the co-occurrence matrix, analyze the potential association between low-correlation features and identify potential incentives, specifically: Set the co-occurrence probability threshold, extract the feature pairs that meet the condition that the co-occurrence probability of the feature pair is less than the co-occurrence probability threshold, and mark them as low co-occurrence probability feature pairs; For each group of low co-occurrence probability feature pairs, combine the eigenvalue fluctuations within the time window to judge whether there are potential incentives; The specific potential incentive identification rule is: if the co-occurrence probability of the feature pair is less than the co-occurrence probability threshold, and the fluctuation patterns of the two features in the feature pair show a relevant trend, then mark it as a potential incentive.
6. The method for retrieving customer interaction data based on dynamic data indexing according to claim 5, characterized in that Based on the composite feature weight distribution and potential incentives, perform real-time update on the dynamic data index structure, and retrieve the newly input customer interaction data based on the real-time updated dynamic data index structure, specifically including: S501: Extract the weight of each feature pair and the potential risk probability of burst behavior from the composite feature weight distribution to generate an updated data set; S502: Adjust the priority of the feature nodes according to the weight of the feature pairs in the updated data set, and redefine the association strength between the feature pairs in combination with the potential risk probability of burst behavior to optimize the index structure; S503: Clean up the feature nodes with low weights in the dynamic data index structure, and sort the feature nodes with high weights at the same time to improve the retrieval efficiency and the compactness of the structure; S504: Use the real-time updated dynamic data index structure to retrieve the newly input customer interaction data, match the high-weight feature pairs associated with the input data, and return the retrieval results.
Citation Information
Patent Citations
Real-time data query method and system based on multi-level index
CN118673043A
Method and system for hybrid information query
US20150058320A1