An AI-based personalized data intelligent evaluation system and method
Through the AI-based personalized data intelligent evaluation system, the time window is dynamically adjusted and the user labeling mechanism is introduced, which solves the problem that the fixed time window cannot adapt to changes in user behavior, realizes the efficiency and transparency of data evaluation, and improves the accuracy and user experience of personalized services.
Patent Information
- Application Number
- CN202510620638.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-14
AI Technical Summary
In the existing personalized data evaluation system, the fixed time window cannot adapt to the frequency of changes in different user behaviors, resulting in insufficient data in the cold start stage of new users or outdated data in active user data, and the evaluation process is opaque, and the user is disconnected from the system's cognition, which affects the accuracy of recommendation.
Using an AI-based personalized data intelligent evaluation system, fluctuation analysis is performed through the time window setting module, the default time window is dynamically output, and a user labeling interaction module is introduced for data labeling analysis. Combined with the dynamic window adjustment module, a converged dynamic time window is built to realize adaptive adjustment of the time window and user intervention.
The defects of the traditional fixed time window are optimized, and changes in user behavior are captured through dynamic time windows, reducing the problem of cold start of new users and outdated user data, improving data screening efficiency and accuracy, enhancing the interaction and transparency between users and the system, and optimizing the accuracy of personalized services.
Smart Images

Figure CN120181935B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data evaluation, and specifically relates to an AI-based personalized data intelligent evaluation system and method. Background Art
[0002] With the rapid development of artificial intelligence technology, AI-based personalized data intelligent evaluation systems have been widely used in e-commerce recommendations, financial risk control, education customization, health management and other fields. Such systems provide users with accurate personalized services by analyzing user behavior, preferences, scenarios and other data, significantly improving user experience and business efficiency.
[0003] However, in the process of personalized data evaluation, timeliness evaluation is a key link. In the existing technology, a fixed time window (such as 30 days, 90 days, etc.) is usually used to judge the effectiveness of personalized data. This "one-size-fits-all" approach has some defects, which are reflected in the following two aspects:
[0004] On the one hand, it cannot adapt to the individual differences in the frequency of changes in user behavior. For example, the amount of data in the cold start phase of new users is small, and the fixed window may only filter out a small amount of historical data, resulting in the model being unable to capture the initial features. For active users in high-frequency changing scenarios (such as short video preferences, fashion consumption, etc.), the fixed window fails to eliminate outdated data in a timely manner, resulting in delayed recommendations. On the other hand, the data evaluation process is not transparent, and users cannot intervene in the evaluation of data validity. For example, users believe that "historical collection data is still valid" (such as long-term interests), but the system automatically filters due to timeliness rules. Users will perceive that the recommendation is wrong, which in turn causes the system's evaluation logic to be out of touch with the user's cognition, resulting in a decrease in the accuracy of recommendations.
[0005] To this end, the present invention provides an AI-based personalized data intelligent evaluation system and method. Summary of the invention
[0006] In order to make up for the deficiencies of the prior art, at least one technical problem raised in the background technology is solved.
[0007] The technical solution adopted by the present invention to solve the technical problem is:
[0008] In the first aspect, a personalized data intelligent evaluation system based on AI specifically includes:
[0009] Time window setting module: collects user behavior data in all scenarios, performs fluctuation analysis on the collected behavior data, sets the initial time window based on the fluctuation analysis results, and dynamically outputs the default time window;
[0010] User annotation interaction module: Based on the real-time default time window, the system filters out the behavior data within the real-time default time window and the behavior data not within the real-time time window. The user conducts annotation analysis on the behavior data not within the real-time time window, and updates the fluctuation analysis based on the annotation analysis results to obtain the correction rules for the time window.
[0011] Dynamic window adjustment module: Combining the default time window and the correction rules of the time window, a fused dynamic time window is constructed to output the final valid data.
[0012] In the second aspect, an AI-based personalized data intelligent evaluation method, the specific steps are as follows:
[0013] Step S10: Collect the behavior data of the user in the full scenario, conduct fluctuation analysis on the collected behavior data, and dynamically output the default time window based on the fluctuation analysis results in combination with the set initial time window.
[0014] Step S20: Based on the real-time default time window, the system filters out the behavior data within the real-time default time window and the behavior data not within the real-time time window. The user conducts annotation analysis on the behavior data not within the real-time time window, and updates the fluctuation analysis based on the annotation analysis results to obtain the correction rules for the time window.
[0015] Step S30: Combining the default time window and the correction rules of the time window, a fused dynamic time window is constructed to output the final valid data.
[0016] The beneficial effects of the present invention are as follows:
[0017] 1. Through the dynamic time window setting and fluctuation analysis, the present invention optimizes the defects of the traditional fixed time window in personalized data evaluation. By constructing a dynamic weighted network, combining with the improved PageRank algorithm to identify key fluctuation nodes, and dynamically adjusting the time window range based on the node influence, it can capture the real-time changes of user behavior, and at the same time, weaken the long-term fluctuations through the time decay factor, strengthen the recent key data, reducing the problems of insufficient data in the cold start stage of new users and outdated data of active users. Moreover, by calculating the node influence and comprehensive fluctuation index in real time, the time window can be adaptively adjusted according to the user behavior pattern, achieving a balance between short-term fluctuation sensitivity and long-term trend stability, significantly improving the efficiency and accuracy of data screening.
[0018] 2. By introducing a user participation mechanism, the present invention collects users' annotations on the importance, relevance, and data quality of data outside the window through a visual annotation tool, and uses BERT semantic embedding and the analytic hierarchy process for annotation consistency analysis, filters low-confidence feedback, ensures the reliability of user feedback, transforms users' subjective judgments into objective correction rules, makes up for personalized scenarios that may be missed by the system's automatic analysis, improves the adaptability of the time window to complex scenarios, and enhances the interactivity and transparency between users and the system;
[0019] 3. By integrating the default time window and the correction rules generated by user annotations, the present invention constructs a dynamic time window. When there is a conflict between the expansion and contraction rules, it introduces a Transformer model to predict the user's intention and dynamically determines the priority of rule execution, realizing a human-computer collaborative data evaluation mechanism. It not only retains the efficiency of the system's automatic analysis but also incorporates users' personalized needs, ensuring that the time window can dynamically converge to a reasonable range according to users' behaviors. Through verification by the comprehensive fluctuation index after each adjustment, it further optimizes the window range, finally outputs high-value data that fits the user's true needs, improves the accuracy of personalized services and the user experience, and optimizes the problem of the disconnection between traditional evaluation logic and users' cognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention will be further described below with reference to the accompanying drawings.
[0021] Figure 1 is a block diagram of a personalized data intelligent evaluation system based on AI of the present invention;
[0022] Figure 2 is a flowchart of the steps of a personalized data intelligent evaluation method based on AI of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] In order to make the technical means, creative features, achieved purposes, and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.
[0024] Example 1
[0025] Please refer to Figure 1 as shown, a personalized data intelligent evaluation system based on AI described in an embodiment of the present invention includes:
[0026] A time window setting module 100: collects users' behavior data in the full scenario, performs fluctuation analysis on the collected behavior data, and dynamically outputs a default time window based on the fluctuation analysis results in combination with the set initial time window;
[0027] Among them, the real-time time window output by the time window setting module serves as the system judgment baseline, which is the time length for the system to default to judge the timeliness of data;
[0028] In this module, the execution process is as follows:
[0029] Access multi-source data interfaces, such as user operation logs, device sensors, and third-party platform APIs;
[0030] Collect real-time behavior data of users in different scenarios;
[0031] Specifically, the scenarios include but are not limited to: web browsing, APP interaction, and offline device use;
[0032] The behavior data includes but is not limited to: operation time, operation type, data object, and interaction duration;
[0033] Clean and preprocess the collected behavior data, filter out outliers and duplicate data, unify the data format and store it in the time series database;
[0034] First specifically, perform fluctuation analysis on the preprocessed behavior data based on the fluctuation propagation quantization model of complex network node influence, and output the initial time window. The process is as follows:
[0035] The fluctuation propagation quantization model of complex network node influence converts the time series of behavior data into a dynamic network, with each time point as a node, and the edge weight between nodes is defined by the fluctuation correlation intensity. The propagation effect of fluctuations is quantified through the node centrality index in graph theory to identify key fluctuation nodes and their impact on the overall data;
[0036] The time series formed by the change of user behavior data over time , calculate the fluctuation correlation degree of adjacent nodes , and the calculation formula is: , where is the minimum value to avoid the denominator of the fluctuation correlation degree being zero, is the eigenvalue of the behavior data;
[0037] Exemplarily, when and respectively represent the operation frequencies of the user at two adjacent moments, if is larger, and is relatively smaller, it indicates that the operation frequencies at these two moments fluctuate significantly, and the correlation degree will increase accordingly;
[0038] Integrate the fluctuation correlation degrees of all adjacent nodes to construct an adjacency matrix;
[0039] Introduce a time decay factor; the core role of the time decay factor is to quantify the decay effect of time on the propagation of fluctuations, in the form of an exponential decay function , where is the decay coefficient, is the time interval, enabling the dynamic weighted network to adaptively weaken long-term fluctuations and strengthen short-term fluctuations, thereby more accurately capturing the real-time change patterns of user behavior;
[0040] Construct a dynamic weighted network , where is the set of nodes, is the set of edges, representing the connection relationship between nodes, is the time-varying weight matrix, enabling the network to dynamically reflect the propagation characteristics of fluctuations at different times;
[0041] Adopt an improved PageRank algorithm to evaluate the node influence, which is used to measure the fluctuation influence of behavior data nodes;
[0042] Set the initial node influence as: , where is the mean of the time series, and the initial influence reflects the deviation degree of node data from the mean;
[0043] Iteratively calculate the node influence : The calculation formula is: ; where
[0044] is the fluctuation influence score of node i at the t-th iteration; d is the damping coefficient, with a default value of 0.85, and its specific value is set by those skilled in the art according to experience, representing the probability of continuous propagation of fluctuation influence in this embodiment; is the set of all edge nodes of node i, is the fluctuation correlation degree from node j to node i; is the sum of the weights of all out-edges of node j;
[0045] It should be noted that is the self-influence retention term. Even without external fluctuation conduction, the historical influence of the node itself will be retained to the next moment at a ratio of 1 - d, reflecting the inertia of fluctuation influence. For example, if a user performs high-frequency operations at a certain moment is relatively high, even if the operation frequency returns to the mean at subsequent moments, the influence at that moment will not disappear immediately, but will be partially retained after decaying according to the damping coefficient;
[0046] Exemplarily, when analyzing user consumption behavior, a large consumption on a certain day (node i) may reflect the user's temporary needs or the impact of promotional activities. Its influence will not completely disappear the next day, but will continuously affect the fluctuation evaluation of subsequent windows through the retention item, reducing the over-filtering due to short-term fluctuations;
[0047] is the influence propagation term of adjacent nodes, is the potential influence of the predecessor node j on the current node i, that is, the product of the fluctuation correlation degree and the historical influence of node j; is the total influence of node j distributed to subsequent nodes according to the weight ratio of each out-edge, avoiding the dilution of influence due to connecting multiple nodes. For example, if node j affects node i and k at the same time, and , then node i will receive more influence from j;
[0048] The node influence converts the isolated time-point data into a network model with causal relationships by simulating the inertial retention and chain propagation of fluctuations in reality, making the node influence contain both the "self-anomaly" and "global propagation importance" dual attributes, providing a quantitative basis for the intelligent adjustment of the time window: not only paying attention to the fluctuation amplitude of a single time point, but also mining the propagation path and key nodes of fluctuations in the time series, so as to achieve a deep understanding and efficient screening of user behavior data;
[0049] Through iterative calculation, the influence spreads layer by layer in the network; high-influence predecessor nodes (such as the time points of user cross-device switching operations) will spread the fluctuation influence to subsequent nodes through strong association edges (such as frequent device interaction records), and finally form an influence ranking reflecting the global propagation effect;
[0050] As the number of iterations increases, the influence scores of each node tend to stabilize, forming the final ranking of node influence. The stabilized not only reflects the abnormal degree of the node's own data (such as the deviation from the mean), but also includes its key role in the entire fluctuation propagation chain (such as serving as a "hub" to connect multiple high-fluctuation nodes);
[0051] Set the node extraction quantity ratio. According to the node extraction quantity ratio, extract the influence nodes within the node extraction quantity ratio in the final ranking of influence nodes as key fluctuation nodes;
[0052] Obtain the timestamps of the key fluctuation nodes; based on the timestamp of the latest key fluctuation node, trace back to the timestamp of the earliest key fluctuation node;
[0053] From the earliest key fluctuation node timestamp to the latest key fluctuation node timestamp, set a time step, and determine the initial time window; the start time of the initial time window is the difference between the earliest key fluctuation node timestamp and the time step, and the end time of the initial time window is the sum of the latest key fluctuation node timestamp and the time step;
[0054] It should be noted that if there are less than two key fluctuation nodes, those skilled in the art shall set the initial time window according to experience combined with the user's daily active period;
[0055] Second specifically, the process of dynamically outputting the default time window is as follows:
[0056] Obtain the user's behavior data in real time and calculate the influence of the current node in real time; compare the influence of the current node with the influence limit value. If the influence of the current node is greater than the influence limit value, it is determined as a real-time key fluctuation node. If the influence of the current node is less than or equal to the influence limit value, it is determined as a real-time non-key fluctuation node;
[0057] Among them, the influence limit value is α times the historical average node influence, where α is obtained by fitting historical data;
[0058] Based on the determination of a real-time key fluctuation node, trigger the condition for updating the time window;
[0059] Furthermore, the conditions for updating the time window include the extended window condition and the contraction window condition;
[0060] For the extended window condition: If the real-time key fluctuation node is within the current time window (i.e., the real-time time window of the initial time window), and the node influence value corresponding to the real-time key fluctuation node is within the ranking proportion threshold of the historical node influence sequence, add the time point of the real-time key fluctuation node to the current time window, and extend the interval boundary of the current time window to include the time points before and after the time point of the real-time key fluctuation node by each duration, to obtain the adjusted time window; where is the average time interval of historical key fluctuation nodes;
[0061] For the contraction window condition: When the proportion of key fluctuation nodes within the time window is lower than the proportion limit value, and the influence data of the current node fluctuates into a stable period, it indicates that there are too many node influence data in the window that contribute little to the fluctuation analysis, and the overall fluctuation intensity is low; among them, the determination process of the stable period is: the standard deviation of the node influence within Q consecutive time units is less than the preset standard deviation threshold;
[0062] At this time, starting from the latest key fluctuation node and shrinking forward to the timestamp of the nearest key fluctuation node can eliminate redundant low-value data and obtain an adjusted time window; in the stage where user operations tend to be stable, shrinking the window can avoid the interference of invalid data, enabling the analysis to focus on the key fluctuation node interval that truly has a fluctuating impact, improving the analysis efficiency and accuracy;
[0063] After each adjustment, perform a weighted average on the node influence within the adjusted time window to calculate the comprehensive fluctuation index BZ of the currently adjusted time window; the calculation formula is: ; where is the set of all nodes within the adjusted time window; is the node influence of node m, is the time interval between node m and the current latest time point; is the attenuation coefficient;
[0064] If the comprehensive fluctuation index is within the preset fluctuation range, then use the currently adjusted time window as the default time window; if the comprehensive fluctuation index is not within the preset fluctuation range, then continue to trigger the adjustment logic;
[0065] To sum up, the time window setting module constructs a dynamic weighted network through multi-source data collection and preprocessing to analyze the fluctuation propagation effect of user behavior data, and realizes the generation and adjustment of the initial time window;
[0066] Access the full-scenario behavior data of users, clean and convert it into a dynamic network. The nodes are time points, and the edge weights are defined by the fluctuation correlation degree of adjacent nodes. Introduce a time decay factor to construct a dynamic weighted network, evaluate the node influence based on the improved PageRank algorithm, taking into account the abnormality of the node itself and the importance of global propagation. Through iterative calculation, form a stable influence ranking, extract a preset proportion of high-influence nodes as key fluctuation nodes, determine the initial time window according to the timestamp range of the key fluctuation nodes and in combination with the historical average time step, and judge in real time whether the new node is a key fluctuation node, triggering the expansion or contraction condition. After each adjustment, calculate the weighted average of the node influence within the window. If it is within the preset interval, it is determined as the default time window; otherwise, continue to optimize;
[0067] It has the following effects: Effect 1: Capture the temporal correlation of user behavior through the fluctuation propagation model, replace the fixed time window, and optimize the problems of insufficient cold start data for new users and outdated data for active users. For example, in high-frequency change scenarios, timely eliminate low-influence old data, retain recent key fluctuation nodes, and optimize recommendation lag; in the cold start stage, supplement the initial window through empirical rules to reduce the situation of excessive valid data;
[0068] Effect 2: Based on the dual attributes of node influence (self-anomaly and global propagation), identify key fluctuation nodes that are highly valuable for user behavior assessment, eliminate low-contribution data, reduce ineffective calculations, and improve analysis efficiency;
[0069] Effect 3: Through real-time dynamic adjustment and verification of the comprehensive fluctuation index, the time window is adaptively adjusted according to the changes in user behavior patterns, balancing the short-term fluctuation sensitivity and long-term trend stability. For example, short-term high-impact behaviors such as large-scale consumption continue to affect the assessment through the attenuation retention mechanism;
[0070] Effect 4: The generated default time window serves as the system's judgment baseline. Subsequent user annotations for data outside the window can trigger correction rules based on this, forming a closed loop of "system intelligent analysis + user personalized intervention", optimizing the problem of the disconnection between traditional assessment logic and user cognition;
[0071] User annotation interaction module 200: Based on the real-time default time window, the system filters out the behavior data within the real-time default time window and the behavior data outside the real-time time window. Users perform annotation analysis on the behavior data outside the real-time time window, and update the fluctuation analysis in combination with the annotation analysis results to obtain the correction rules for the time window;
[0072] In this module, the execution process is as follows:
[0073] Based on the real-time default time window, identify whether the user's behavior data is within the real-time default time window;
[0074] Take the behavior data within the real-time default time window as the system's default valid data;
[0075] Take the behavior data outside the real-time default time window as the interaction to-be-evaluated data;
[0076] Display the interaction to-be-evaluated data for users and provide visualization annotation tools, including various annotation methods: importance annotation, relevance annotation, and data quality annotation;
[0077] Specifically, importance annotation: used for users to annotate the importance degree of the to-be-evaluated data, which can include: high importance degree, medium importance degree, and low importance degree. For example, users mark the usage behavior of a certain function as high importance degree;
[0078] Relevance annotation: used for users to annotate the relevance degree of the to-be-evaluated data, which can include: relevant and irrelevant. For example, users determine that a certain behavior data is relevant to consumption preferences;
[0079] Data quality annotation: Used for users to annotate the quality of the data to be evaluated, which may include: valid, missing, noise. For example, if a user determines that a certain operation results in incomplete records due to a system error, it is annotated as missing;
[0080] Perform annotation consistency analysis on the same user. Through consistency analysis, low-confidence annotations can be filtered out to improve the quality of the user annotation feedback data;
[0081] Furthermore, the process of performing consistency analysis from two dimensions of free text semantics and cross-time stability is as follows:
[0082] The text remarks added by the user during annotation (such as "This operation is related to account security") are used to calculate the cosine similarity as the semantic similarity through BERT semantic embedding;
[0083] Exemplarily, if two remarks are "Frequent abnormal logins" and "Abnormal login frequency" respectively, the semantic similarity is relatively high and they are regarded as consistent; if the remark is "Normal operation", the semantic similarity is relatively low and they are regarded as contradictory;
[0084] Calculate the difference rate of the annotation results of the same user for the same cluster of data at different times; the difference rate is the percentage of the ratio of the number of inconsistent annotations to the total number of annotations;
[0085] Exemplarily, if a user annotated "Irrelevant" for "Nighttime login" 30 days ago and currently annotates "Relevant", if the historical difference rate of this cluster > 40%, it is marked as "Temporally unstable annotation";
[0086] It should be noted that the same cluster of data refers to a data set with similar attributes formed by grouping user behavior data according to features through a data clustering method. Based on the feature similarity of user behavior data (such as operation type, data object, interaction duration, scenario context), the user behavior data is divided into multiple data clusters through a clustering algorithm (such as K-means, DBSCAN). The data within each cluster has highly similar behavior patterns or attributes, which is used to evaluate the temporal stability of the user's annotation of similar behaviors;
[0087] Example: The "Nighttime login" behavior of a user (regardless of the specific date on which it occurs) may form a cluster, and "Frequent product browsing" forms another cluster
[0088] Introduce the analytic hierarchy process to calculate the weight coefficients of the semantic similarity and the difference rate of the annotation results, and perform weighted fusion to obtain the confidence level of the user standard consistency;
[0089] It should be explained that the hierarchical analysis method can be: construct a 2×2 judgment matrix to compare the importance of semantic similarity and the difference rate of annotation results. For example, if the ratio of semantic similarity to annotation difference is 3:1, it means that semantic similarity is more important. The weight is calculated by the eigenvalue method to ensure that the consistency ratio is less than 0.1. The weight synthesis formula is: confidence = 0.6*semantic similarity + 0.4*(1-difference rate of annotation results);
[0090] If the confidence level is greater than or equal to the confidence level limit, it is determined that the user's annotations are consistent and the annotation quality is high;
[0091] If the confidence level is less than the confidence limit, it is determined that the user's annotation is biased, which triggers the user's secondary verification process and returns it to the user for secondary verification (such as a pop-up window prompting "There is a contradiction between your two annotations, please confirm");
[0092] The process of obtaining the correction rule of the time window is:
[0093] For importance labels and relevance labels driving extended window rules: filter the behavior data marked as highly important and relevant by users in the data to be evaluated as key labeled data. Although these data are outside the real-time default time window, they are of high value for evaluating user behavior.
[0094] When key annotation data is detected, the extended window rule is triggered. Specifically, the time point containing the key annotation data is included in the current time window, and the time point is used as the benchmark to expand the forward and backward time lengths (the average time interval of historical key fluctuation nodes), thereby expanding the scope of the time window to cover more data that is valuable for user behavior evaluation;
[0095] For data quality label-driven shrinking window rules: filter the behavior data marked as missing or noisy by the user in the data to be evaluated as low-quality data; count the number of low-quality data, and calculate the proportion of the number of low-quality data to the total amount of behavior data in the default time window;
[0096] At the same time, determine whether the current node influence data has entered a stable period, that is, if the standard deviation of the node influence within Q consecutive time units is less than the preset standard deviation threshold, it is a stable period;
[0097] If the proportion of low-quality data is greater than the proportion of low-quality data and is in a stable period, the shrinking window rule is triggered, starting from the latest key fluctuation node, shrinking forward to the timestamp of the most recent key fluctuation node, eliminating redundant low-value data in the time window, and reducing the time window range; otherwise, the shrinking window rule is not triggered;
[0098] In summary, the user annotation interaction module filters out the data to be evaluated outside the window for user annotation through the default time window, filters out low-confidence feedback through multi-dimensional annotation consistency analysis, and obtains the time window correction rule;
[0099] Take the data within the default time window as the effective data of the system, and the data outside the window as the data to be evaluated. Provide three types of annotation methods: importance, relevance, and data quality. Support users to supplement annotation basis through text remarks. From the two dimensions of semantic similarity and time stability, the annotation confidence is obtained through the analytic hierarchy process weighted fusion. If the confidence is lower than the threshold, trigger secondary verification to ensure the annotation quality;
[0100] It has the following effects: Effect 1: Through annotation interaction, transform the user's subjective judgment of data value into an objective correction rule, making up for personalized scenarios that may be missed by the system's automatic fluctuation analysis. For example, users manually mark historical low-frequency but important operations, reducing evaluation bias caused by system default window filtering;
[0101] Effect 2: Through the double verification of semantic similarity and time stability, combined with the dynamic weighting of the analytic hierarchy process, effectively filter out contradictory or unstable annotations (such as opposite annotations of the same user for the same data in a short period), improving the reliability of feedback data and avoiding misleading system adjustments by low-quality annotations;
[0102] Effect 3: Retain the efficiency of the system's automatic analysis and introduce user domain knowledge (such as business scenario understanding). For example, in the financial risk control scenario, users mark "abnormal login" data as highly important, and force the expansion of the window to include related time points, improving the comprehensiveness of risk identification;
[0103] Effect 4: Through the contraction rule driven by data quality annotation, automatically clean up noise data (such as invalid operation records generated by system errors), and combine the judgment of the stable period to avoid excessive contraction, finding a balance between user participation and system automation, and improving the adaptability of the time window to complex scenarios;
[0104] The dynamic window adjustment module 300: Combine the default time window and the correction rule of the time window to construct a fused dynamic time window and output the final effective data;
[0105] In this module, the execution process is as follows:
[0106] Obtain the default time window and the correction rule;
[0107] Identify the obtained correction rule. If the expansion window rule and the contraction window rule overlap, determine the execution priority. The process is as follows:
[0108] Introduce the Transformer model to predict the true intention of the current operation by analyzing the user's historical annotation behavior data, so as to dynamically determine the rule priority;
[0109] Train an intention prediction model through the user's historical behavior data, and associate the user's behavior pattern with the data expansion requirement or data reduction requirement;
[0110] Optionally, the training data is the user's annotation records in the past 30 days. The output features include: annotation timestamp, data cluster type (one-hot encoding), node influence value, historical window adjustment record; the output is a binary classification result (0 for contraction requirement, 1 for expansion requirement), and it is trained using the cross-entropy loss function. The accuracy of the validation set needs to be ≥85%;
[0111] Furthermore, the annotation timestamp: reflects the time correlation of user intervention (recent annotations better reflect the current intention), the data cluster type: based on user behavior clustering (such as "night login" "high-frequency browsing" clusters), which conforms to the setting of the "same cluster data" annotation consistency analysis in the text, the node influence value: directly correlates with the fluctuation analysis result, reflecting the importance of the data itself, and the historical adjustment record: captures the user's long-term preference (such as frequently expanding the window may indicate that the user tends to retain more data);
[0112] When there is a conflict between the expansion and contraction rules, input the current behavior data into the model to predict whether the user is more inclined to expand the data range or shrink the data range;
[0113] If the prediction result is an expansion requirement, the expansion window rule is preferentially executed; if it is a contraction requirement, the contraction window rule is preferentially executed;
[0114] Exemplarily, when the user recently frequently annotates the data outside the window as highly important, the model predicts that the user may be exploring new behavior patterns, and at this time, the window is preferentially expanded; if the user continuously annotates low-quality data, the window is preferentially contracted;
[0115] By introducing the Transformer model to predict the user's intention, it not only continues the efficiency of automatic analysis but also incorporates the flexibility of user personalized intervention, which is a key implementation link of the "human-machine collaboration" data evaluation mechanism;
[0116] Execute the selected rule and adjust the time window, specifically:
[0117] Expanded window execution: If the expanded window rule is executed first, the time points containing critical annotation data are included in the current time window. Based on this time point and the average time interval of historical critical fluctuation nodes, the corresponding duration is extended forward and backward. Meanwhile, the node set within the time window is updated, the influence of each node is recalculated, and the influence ranking of the nodes is re-evaluated based on the improved PageRank algorithm to determine the new critical fluctuation nodes;
[0118] Contracted window execution: If the contracted window rule is executed first, starting from the latest critical fluctuation node, it contracts forward to the timestamp of the nearest critical fluctuation node, eliminating redundant low-value data. The influence of the nodes within the contracted time window is weighted and averaged, and the comprehensive fluctuation index BZ of the current time window is recalculated. If the comprehensive fluctuation index is not within the preset fluctuation range, the adjustment logic is continuously triggered until the comprehensive fluctuation index is within the preset range;
[0119] Based on the adjusted dynamic time window, the behavior data within this time window is filtered out as the final valid data; meanwhile, these data are cleaned and preprocessed again, including removing residual outliers and duplicate data, ensuring the uniformity of the data format, and improving the data quality;
[0120] In summary, the dynamic window adjustment module works as follows: It obtains the default time window and the correction rules generated based on user annotations. When there is a conflict between the expansion and contraction rules due to overlapping time intervals, it predicts the true intention of the user's current operation by introducing a Transformer model, and dynamically determines the priority of rule execution based on the prediction result. If the prediction is an expansion requirement, the time points of critical annotation data are first included in the window, and it is extended forward and backward according to the average time interval of historical critical fluctuation nodes, updating the node set and re-evaluating the influence ranking. If it is a contraction requirement, starting from the latest critical node, the low-value data interval is eliminated, and the comprehensive fluctuation index is recalculated by weighted-averaging the node influence until it falls within the preset range. Based on the adjusted dynamic window, the behavior data is filtered, and after cleaning and preprocessing, high-quality final valid data is output;
[0121] It has the following effects: Effect 1: By correcting the time window through user intervention, a closed-loop of system intelligent analysis and user personalized intervention is achieved;
[0122] Effect 2: After each adjustment, the comprehensive fluctuation index is recalculated. If it does not fall within the preset range, it continues to be optimized to ensure that the time window dynamically converges to a reasonable range with the user's behavior pattern, balancing the short-term fluctuation sensitivity and the long-term trend stability;
[0123] Effect 3: It continues the high efficiency of the system's automatic analysis, and through user annotation, it realizes personalized calibration, optimizing the problems of insufficient adaptation to individual differences and opaque evaluation process of traditional fixed windows, and finally outputs high-value data that meets the real needs of users, improving the accuracy of personalized services and the user experience.
[0124] Embodiment 2
[0125] This application also provides an AI-based personalized data intelligent evaluation method, which is applicable to this kind of AI-based personalized data intelligent evaluation system. Please refer to Figure 2 As shown in the figure, it is a schematic flowchart of an AI-based personalized data intelligent evaluation method provided by this application, including the following steps:
[0126] Step S10: Collect the user's behavior data in the full scenario, perform fluctuation analysis on the collected behavior data, and based on the fluctuation analysis results, combine with the set initial time window to dynamically output the default time window;
[0127] The specific process of this step is as follows: First, access the user's full-scenario behavior data, clean it and convert it into a dynamic network. The nodes are time points, and the edge weights are defined by the fluctuation correlation degree of adjacent nodes. Introduce a time decay factor to construct a dynamic weighted network, and then evaluate the node influence based on the improved PageRank algorithm, taking into account the abnormality of the node itself and the importance of global propagation. Through iterative calculation, a stable influence ranking is formed, and a preset proportion of high-influence nodes are extracted as key fluctuation nodes. According to the time stamp range of the key fluctuation nodes, combined with the historical average time step, determine the initial time window, and real-time judge whether the new node is a key fluctuation node, trigger the expansion or contraction condition, and calculate the weighted average of the node influence within the window after each adjustment. If it is within the preset interval, it is determined as the default time window, otherwise continue to optimize;
[0128] Step S20: Based on the real-time default time window, the system filters out the behavior data within the real-time default time window and the behavior data not within the real-time time window. The user performs annotation analysis on the behavior data not within the real-time time window, and updates the fluctuation analysis in combination with the annotation analysis results to obtain the correction rule of the time window;
[0129] The specific process of this step is as follows: First, take the data within the default time window as the effective data of the system, and the data outside the window as the data to be evaluated. Provide three types of annotation methods: importance, relevance, and data quality, support the user to supplement the annotation basis through text remarks, and obtain the annotation confidence through weighted fusion from the two dimensions of semantic similarity and time stability. If the confidence is lower than the threshold, trigger secondary verification to ensure the annotation quality;
[0130] Step S30: Combine the default time window with the correction rules of the time window to construct a fused dynamic time window and output the final valid data;
[0131] The specific process of this step is as follows: Obtain the default time window and the correction rules generated based on user annotations. When conflicts occur due to overlapping time intervals in the expansion and contraction rules, introduce a Transformer model to predict the true intention of the user's current operation, and dynamically determine the rule execution priority according to the prediction result. If the prediction is an expansion requirement, first include the key annotation data time points in the window, expand forward and backward according to the average time interval of historical key fluctuation nodes, update the node set and re-evaluate the influence ranking. If it is a contraction requirement, start from the latest key node to eliminate low-value data intervals, and recalculate the comprehensive fluctuation index by weighted average of node influence until it falls within the preset interval. Filter the behavior data based on the adjusted dynamic window, and output high-quality final valid data after cleaning and preprocessing.
[0132] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI-based personalized data intelligent evaluation system, characterized in that: Specifically include: Time window setting module: Collect the behavior data of the user in the full scenario, conduct fluctuation analysis on the collected behavior data, and based on the results of the fluctuation analysis, combine with the set initial time window to dynamically output the default time window; The process of the above-mentioned fluctuation analysis is as follows: Convert the time series of behavior data into a dynamic weighted network, where the nodes are time points, and the edge weights are based on the fluctuation correlation degree of adjacent nodes. Introduce a time decay factor to construct the adjacency matrix; Use the improved PageRank algorithm to evaluate the node influence. Take the deviation degree of the node eigenvalue from the mean as the initial influence, and combine with the initial influence to iteratively calculate the node influence; Extract the key fluctuation nodes with a preset ratio according to the node influence ranking; The process of the above-mentioned dynamically outputting the default time window is as follows: Determine the initial time window according to the time stamp range of the key fluctuation nodes and the historical average time step; Obtain the user's behavior data in real time, and calculate the influence of the current node in real time; If it is greater than the influence limit value, it is determined as a real-time key fluctuation node; Trigger the condition for updating the time window to adjust the time window. After each adjustment, perform a weighted average on the node influence within the adjusted time window, and calculate the comprehensive fluctuation index of the currently adjusted time window; If the comprehensive fluctuation index is within the preset fluctuation range, then use the currently adjusted time window as the default time window; User annotation interaction module: Based on the real-time default time window, the system screens out the behavior data within the real-time default time window and the behavior data not within the real-time time window. The user conducts annotation analysis on the behavior data not within the real-time time window, and updates the fluctuation analysis in combination with the annotation analysis results to obtain the correction rule of the time window; The correction rule of the above-mentioned time window is as follows: For the importance label and relevance label-driven extended window rule: Screen out the behavior data marked by the user as highly important and relevant in the data to be evaluated as the key annotation data; Trigger the extended window rule; Incorporate the time points containing the key annotation data into the current time window, and take this time point as the benchmark to extend the duration forward and backward; For the data quality label-driven shrinking window rule: Screen out the behavior data marked by the user as missing or noisy in the data to be evaluated as low-quality data; Count the number of low-quality data, and calculate the proportion of the number of low-quality data in the total amount of behavior data within the default time window; Judge whether the current node influence data enters a stable period; If the proportion of low-quality data is greater than the proportion of low-quality data and is in a stable period, trigger the shrinking window rule, starting from the latest key fluctuation node, and shrink forward to the time stamp of the nearest key fluctuation node to narrow the time window range; Dynamic window adjustment module: Combine the default time window and the correction rule of the time window to construct a fused dynamic time window and output the final valid data.
2. The personalized data intelligent evaluation system based on AI according to claim 1, wherein: The above-mentioned updating the time window includes the extended window condition and the shrinking window condition; For the extended window condition: If the real-time critical fluctuation node is within the real-time time window of the initial time window, and the node influence value corresponding to the real-time critical fluctuation node is within the ranking proportion threshold of the historical node influence sequence, add the time point of the real-time critical fluctuation node to the current time window, and extend the interval boundary of the current time window to include the average time interval duration of each historical critical fluctuation node before and after the time point of the real-time critical fluctuation node, to obtain an adjusted time window; For the shrinking window condition: When the proportion of critical fluctuation nodes within the time window is lower than the proportion limit value, and the fluctuation of the current node influence data enters a stable period; The stable period is when the standard deviation of the node influence within consecutive Q time units is less than the preset standard deviation threshold; Starting from the latest critical fluctuation node, shrink forward to the time stamp of the nearest critical fluctuation node to obtain an adjusted time window.
3. An AI-based personalized data intelligent evaluation system according to claim 1, characterized in that: The initial time window is: Obtain the time stamp of the critical fluctuation node; Based on the time stamp of the latest critical fluctuation node, trace back to the time stamp of the earliest critical fluctuation node; From the time stamp of the earliest critical fluctuation node to the time stamp of the latest critical fluctuation node, and set a time step to determine the initial time window; The start time of the initial time window is the difference between the time stamp of the earliest critical fluctuation node and the time step, and the end time of the initial time window is the sum of the time stamp of the latest critical fluctuation node and the time step.
4. An AI-based personalized data intelligent evaluation system according to claim 1, characterized in that: The annotation analysis process is: Based on the real-time default time window, identify whether the user's behavior data is within the real-time default time window; Take the behavior data that is not within the real-time default time window as the interaction data to be evaluated; Display the interaction data to be evaluated to the user and provide it to the visualization annotation tool, which includes importance annotation, relevance annotation, and data quality annotation; And perform consistency analysis on the user's annotation results.
5. An AI-based personalized data intelligent evaluation system according to claim 4, characterized in that: The consistency analysis process is: For the text remarks added by the user during annotation, calculate the cosine similarity as the semantic similarity through BERT semantic embedding; Calculate the difference rate of the annotation results of the same user for the same cluster data at different times; The difference rate is the percentage of the ratio of the number of inconsistent annotations to the total number of annotations; Introduce the analytic hierarchy process to calculate the weight coefficients of the semantic similarity and the annotation result difference rate, and perform weighted fusion to obtain the confidence level of the user standard consistency; If the confidence level is greater than or equal to the limit value of the confidence level, it is determined that the user annotation is consistent; If the confidence level is less than the limit value of the confidence level, it is determined that there is a deviation in the user annotation, and the user secondary verification process is triggered and returned to the user for secondary verification.
6. An AI-based personalized data intelligent evaluation system according to claim 1, characterized in that: The process of outputting the final valid data is: Identify the obtained correction rules. When the extended window rule and the shrinking window rule overlap; Train an intent prediction model through the Transformer model algorithm, input the current behavior data into the model, and predict whether the user is more inclined to expand the data range or shrink the data range. If the prediction result is an expansion requirement, the extended window rule is preferentially executed; If it is a shrinking requirement, the shrinking window rule is preferentially executed; Execute the selected rules and adjust the time window. Based on the adjusted dynamic time window, filter out the behavior data within this time window as the final valid data.
7. An AI-based personalized data intelligent evaluation method, applied to an AI-based personalized data intelligent evaluation system according to any one of claims 1-6, characterized in that: The specific steps are as follows: Step S10: Collect the behavior data of the user in the full scenario, perform fluctuation analysis on the collected behavior data, and dynamically output the default time window based on the fluctuation analysis results in combination with the set initial time window; Step S20: Based on the real-time default time window, the system filters out the behavior data within the real-time default time window and the behavior data not within the real-time time window. The user performs annotation analysis on the behavior data not within the real-time time window, and updates the fluctuation analysis in combination with the annotation analysis results to obtain the correction rules for the time window; Step S30: Combine the default time window and the correction rules for the time window to construct a fused dynamic time window and output the final valid data.
Citation Information
Patent Citations
Effective dynamic network node influence measurement method
CN107958032A