Product optimization system based on user behavior data mining
By mining user behavior data, features are extracted, critical paths are identified, abnormal paths are analyzed, and risks are quantified. This solves the problem of unstable risk assessment in existing technologies, realizes the comprehensiveness and applicability of risk scoring, and provides more accurate product optimization suggestions.
Patent Information
- Application Number
- CN202511060916.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-11
AI Technical Summary
Existing product optimization technologies rely on fixed rules or single weights for risk assessment, lacking a normalization mechanism. This leads to inconsistent scoring results that are susceptible to extreme data and fail to accurately reflect actual risks.
By mining user behavior data, we extract user behavior features, identify critical paths, analyze abnormal path segments, perform structural mapping and aggregation, quantify risk scores, and use a multi-dimensional indicator comprehensive scoring model with adjustable weights to improve the stability and applicability of the scores.
It achieves comprehensiveness and comparability of risk scoring results, improves the stability and flexibility of the model, adapts to different product needs and risk preferences, and provides more accurate guidance for product optimization.
Smart Images

Figure CN120929758A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of product optimization technology, specifically a product optimization system based on user behavior data mining. Background Technology
[0002] With the development of intelligent and digital operations in products, user behavior data generated during product use is widely used in tasks such as product optimization, user profile building, and strategy feedback and adjustment. Existing product optimization technologies typically rely on user behavior logs (such as click records, usage frequency, function calls, etc.) for feature extraction and behavior modeling, and use classification models, clustering algorithms, sequence analysis, and other methods to identify user preferences, abnormal behaviors, or product defect usage scenarios, thereby assisting in the optimization of product structure, functional modules, or interface processes.
[0003] Existing risk assessment methods have shortcomings: most existing risk assessment methods use fixed rules or threshold judgments, isolated event assessments, etc. They typically use fixed parameters or single weights, such as single failure rate scoring methods, and lack normalization mechanisms. They are easily affected by extreme data, resulting in large fluctuations and inconsistent scoring results that are not comparable and cannot accurately reflect actual risks. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a product optimization system based on user behavior data mining to solve the problems mentioned in the background section.
[0005] To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a product optimization system based on user behavior data mining, comprising the following steps: S1. Extract user behavior record sequences to obtain user behavior features; S2. Identify critical paths based on user behavior characteristics to obtain critical usage paths; S3. Perform path deviation comparison analysis based on the key usage path to obtain abnormal path segments; S4. Perform structural mapping and aggregation based on abnormal path fragments to obtain local behavior clusters; S5. Quantify the risk based on local behavioral clusters to obtain the risk score results; S6. Compare versions based on risk scoring results to obtain optimization results.
[0006] To further optimize this technical solution, the path deviation comparison analysis in step S3 includes: Based on the obtained key usage paths, abnormal path segments in the user's actual behavior path are identified, and each abnormal path segment is assigned a semantic label, thereby elevating low-level path deviation comparison analysis to high-level user intent understanding, making the results of path deviation comparison analysis more interpretable and valuable for application.
[0007] To further optimize this technical solution, the structure mapping aggregation in step S4 includes: The obtained abnormal path fragments are subjected to structural mapping and aggregation. Abnormal path fragments with similarity are merged together to aggregate representative local behavior clusters, thereby discovering structural problem points and providing targeted input for subsequent scoring and product optimization.
[0008] To further optimize this technical solution, the risk quantification in step S5 includes: Based on the obtained local behavioral clusters, a risk scoring model is used to quantify the risk of these clusters and obtain risk score results to quantify potential negative impacts.
[0009] To further optimize this technical solution, the risk scoring model includes:
[0010] in: Cluster Comprehensive risk score; Cluster The number of users; Total number of users across all clusters; Cluster Behavioral risk score; The weighting coefficient representing the percentage of users in the cluster; : Weighting coefficients for behavioral risk scores.
[0011] To further optimize this technical solution, the behavioral risk scoring includes:
[0012] in: Cluster Functional failure rate; Cluster Path deviation; Cluster The complexity of behavior; Weighting coefficient for functional failure rate; : Weighting coefficient for path deviation; Weighting coefficients for behavioral complexity; The cluster's behavioral risk score is obtained by comprehensively calculating the functional failure rate, path deviation, and behavioral complexity.
[0013] To further optimize this technical solution, the functional failure rate includes:
[0014] in: :Function The failure rate; Cluster The collection of all functions involved; Cluster The number of functions; The functional failure rate is obtained by calculating the failure rate of all functions involved in the cluster and taking the average.
[0015] To further optimize this technical solution, the path deviation includes:
[0016] in: : Abnormal path Standard Critical Path Edit distance; Standard Critical Path Length; Cluster A collection of abnormal paths; Cluster The number of abnormal paths; The path deviation is obtained by summing and averaging the normalized edit distances between abnormal paths and standard critical paths in the cluster.
[0017] To further optimize this technical solution, the behavioral complexity includes:
[0018] in: : Sigmoid function; : Function sensitivity adjustment factor; The behavioral complexity is obtained by calculating the ratio of the length of the abnormal path to the length of the standard critical path, normalizing it, and averaging it.
[0019] This technical solution has been further optimized, including the following functional modules: The module includes a feature extraction module, a path recognition module, an anomaly analysis module, a structure aggregation module, a risk scoring module, and an optimization evaluation module.
[0020] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, they implement the steps of a product optimization system based on user behavior data mining as described in the first aspect of the present invention.
[0021] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, they implement the steps of a product optimization system based on user behavior data mining as described in the first aspect of the present invention.
[0022] Compared with existing technologies, the present invention provides a product optimization system based on user behavior data mining, which has the following beneficial effects: This product optimization system based on user behavior data mining introduces a comparison mechanism with the standard critical path through a risk scoring model. It uses normalized multi-dimensional indicators for comprehensive scoring, which improves the comprehensiveness and comparability of risk scoring and enhances the stability of the model. Furthermore, it uses adjustable weights to adapt to different product needs and risk preferences, thereby improving the flexibility and applicability of risk scoring. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating a product optimization system based on user behavior data mining proposed in this invention. Figure 2 This is a flowchart illustrating a risk scoring model for a product optimization system based on user behavior data mining, as proposed in this invention. Figure 3 This is a schematic diagram of a product optimization system based on user behavior data mining proposed in this invention. Detailed Implementation
[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0026] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0027] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0028] Example 1
[0029] Reference Figures 1-2 This is the first embodiment of the present invention, which provides a product optimization system based on user behavior data mining, including the following steps: S1. Extract user behavior record sequences to obtain user behavior features.
[0030] In this embodiment, the extraction of the user behavior record sequence includes: During product optimization, users' actual operational behaviors are the most direct source of data reflecting product interaction efficiency, feature usage preferences, and potential problems. User behavior record sequences, as a temporal expression of this data, contain rich information. Therefore, it is necessary to extract these behaviors in a structured manner to obtain user behavior characteristics.
[0031] The purpose of this step is to extract representative behavioral features from the collected historical interaction data between users and products, i.e., user behavior record sequences, and construct a set of user behavior features, thereby improving the processability and modeling efficiency of behavioral data and providing a data input basis for subsequent steps.
[0032] The implementation methods for this step include: First, user operation logs from clients, servers, or embedded terminals are collected and reconstructed into a chronological sequence of user behavior events. Each event sequence includes multi-dimensional attributes such as behavior type (e.g., click, swipe, pause), timestamp, operation object, and device type. To improve data usability, each user behavior event sequence is divided into behavior segments within a fixed behavior window (a small value is recommended, or it can be set according to the product's average number of operation steps, such as 3-5 behavior operations, with a swipe step size of 1, i.e., swiping one behavior at a time, which is beneficial for constructing a temporal structure). Feature extraction is performed on each behavior segment, including: Behavior type frequency statistics: Calculate the frequency of various behavioral operations (such as clicking, accessing, and collecting) to reflect the user's operational tendencies over a period of time and measure the user's dependence on or anxiety about the function or resource. For example, frequent clicking on product details may reflect a strong purchase intention. Behavioral path length and average dwell time calculation: Path length refers to the number of operations performed by a user in one behavior, that is, the hierarchical span from the start page to the end page. Average dwell time is calculated by the time difference between consecutive page browsing, which represents the average dwell time of the user at each step, thereby characterizing the depth of the user's single visit and the average attention time to the content, and evaluating the depth of the user's visit and the degree of attention to the content within a behavioral segment. For example, a long path and a high dwell time often indicate a clear goal orientation, while the opposite may be a shallow and casual browsing. Operation sequence pattern extraction: By recording the order in which users perform operations, a directional user behavior path is constructed, such as "login → browse products → favorite products → add to cart". This is used to characterize the structure and logic of the behavior path, which helps in subsequent analysis of common usage paths and abnormal jump behaviors. Behavior termination interval assessment: Calculate the time interval between the user's last action and the current time. A short time interval indicates high activity, while a long time interval may indicate that the user has moved elsewhere or lost interest in the product. This is used to infer the user's activity status and churn risk. During the feature extraction process, a unified feature template is used to encode multidimensional features into vector form, thereby constructing a set of user behavior features.
[0033] S2. Identify critical paths based on user behavior characteristics to obtain critical usage paths.
[0034] In this embodiment, the critical path identification includes: In step S1, user behavior features are extracted from user behavior records to form a set of user behavior features, which includes relevant data on user behavior paths. In order to discover the core behavior paths of users when using the product, this step identifies key paths based on the user behavior paths in the user behavior features. By generating candidate usage paths and using a path concentration discrimination model, the importance of the path is measured, representative key usage paths are identified, and the core process of users when using the product is determined. This allows for a more accurate focus on highly representative behavioral structures as the path basis for subsequent analysis.
[0035] Furthermore, the path concentration discrimination model includes:
[0036] in: : No. The concentration score of the candidate paths; The number of user behavior paths indicates the number of user behavior paths in the user behavior records. : No. Candidate path and the first The edit distance of a user behavior path represents the minimum number of edit operations required to transform one sequence into another. It is used to measure the similarity between two paths. The more similar the two sequences are, the fewer edit operations are required, and the smaller the edit distance is. Conversely, the greater the difference, the larger the edit distance is. Edit operations typically include insertion (adding a behavior operation to the path, such as adding a checkout operation), deletion (removing a behavior operation from the path, such as deleting an add-to-cart operation), and replacement (replacing a behavior operation in the path with another operation, such as replacing a favorite operation with an add-to-cart operation). : No. Candidate paths; : No. A user behavior path, that is, the user behavior path of each behavior segment obtained in step S1, for example, login → browse products → favorite products → add to cart; Distance normalization factor: Used to edit the effect of distance on centrality scores. The smaller the value, the smaller the difference in paths, and the higher the score. The larger the value, the smoother the function, and the scores of all paths tend to be averaged.
[0037] Furthermore, the distance normalization factor includes:
[0038] in: The number of candidate paths; The distance normalization factor is obtained by calculating the average edit distance between all candidate paths and user behavior paths.
[0039] This model describes how to identify critical paths by using centralized scoring to identify candidate paths based on user behavior paths in user behavior characteristics.
[0040] Traditional path recognition methods often rely on behavior frequency or complete path comparison, which can easily miss low-frequency but representative paths. This results in low sensitivity to structural differences and insufficient generalization ability during path recognition, ignoring the diversity of paths. In contrast, this model extracts behavior fragments and constructs candidate paths through a sliding window, which is not constrained by frequency thresholds. This improves the retention rate of rare but critical paths, enhances the ability to capture complex user behaviors, and improves the stability of critical path recognition. It adopts a centralized scoring mechanism to improve the representativeness and refinement of the identified paths, and adjusts the centralized score through a distance normalization factor to improve the flexibility of path recognition.
[0041] The steps for using this model include: Candidate path generation: Based on the complete user behavior path (not just behavior fragments) extracted from each user behavior event sequence obtained in step S1, the window length is set (based on the average number of operation steps of the product, thus identifying a complete user behavior path as a path fragment; for example, a common product purchase path includes browsing, clicking, adding to cart, checkout, and payment, totaling 5 steps, so the window length can be set to 4-6). Then, a sliding window slicing method is used (the sliding step length is the same as the window length, extracting the complete path structure that can express the semantics of the product flow, and distinguishing it from the behavior fragment division in step S1 to avoid semantic overlap between behavior fragments and candidate paths), dividing the user behavior path into multiple path fragments. Finally, duplicate fragments are removed to obtain candidate paths. ; Path concentration determination: based on the obtained candidate paths Combined with the user behavior path obtained in step S1 Calculate the distance normalization factor Edit distance between paths Then, combined with the distance normalization factor Edit distance between paths Calculate the centrality score of the candidate paths. ; Critical path selection: Based on the concentration score of the calculated candidate paths. Selecting clustered scores Paths with a value greater than or equal to a threshold are designated as critical use paths, forming a set of critical use paths. The threshold can be set using the quantile method to make it adapt to the data distribution. For example, if the 75th percentile of the centrality score of all paths is used as the threshold, then paths with a centrality score in the top 25% range (i.e., higher scores) will be retained as critical use paths.
[0042] S3. Perform path deviation comparison analysis based on the key usage path to obtain abnormal path segments.
[0043] In this embodiment, the path deviation comparison analysis includes: In step S2, the critical usage path was obtained, and the core process of users using the product was determined. However, in actual use, different users may deviate from the critical path due to issues such as cognitive habits, task understanding, and insufficient interface guidance.
[0044] The purpose of this step is to identify anomalous path segments in the user's actual behavior path that deviate significantly from the key usage path of their group, based on the obtained key usage path. Furthermore, these deviations are classified and explained, and each anomalous path segment is assigned a semantic label. This elevates the low-level path deviation comparative analysis to a high-level understanding of user intent, making the results of the path deviation comparative analysis more interpretable and valuable for application.
[0045] The implementation methods for this step include: Key path matching of user's actual behavior path to the group: Extract the behavior path of the user to be analyzed within the specified task. Based on the key usage path obtained in step S2, assign the key usage path to each user's actual behavior path. The assignment method can be based on the minimum edit distance to select the optimal key usage path, that is, select the key usage path with the smallest edit distance as the key path of the group to which the user's actual behavior path belongs. Two types of path temporal alignment: The Dynamic Time Warping (DTW) algorithm is used to calculate the temporal alignment between the user's actual behavior path and its key usage path, and output the offset distance matrix to quantify the degree of local offset between the user's actual behavior path and the key usage path. Abnormal path segment identification: In the offset distance matrix, identify continuous sub-path segments whose offset distance exceeds a preset threshold. The threshold can be set based on the 75th percentile of the historical path deviation score (or set according to experience). Continuous sub-path segments whose offset distance is in the top 25% range (i.e., larger offset distance) are abnormal path segments that deviate from the critical usage path. Multidimensional Deviation Indicator Extraction: For each abnormal path segment, the following indicators are statistically calculated using the data obtained from feature extraction to characterize behavioral deviation features, including: function call order reversal rate (reflecting sequence consistency deviation; comparing the abnormal path segment with the corresponding key usage path to count the number of function nodes with misaligned order and calculate their proportion in the total number of operations), page jump interruption number (reflecting path completion deviation; traversing the abnormal path segment to detect whether there are cases of premature exit or failure to jump to the next key node after page dwell, and accumulating the count), function click return rate (reflecting operation trial deviation; counting the number of all "click-return" pairs in the abnormal path segment and dividing by the total number of clicks), and abnormal dwell time distribution coefficient (reflecting attention distribution deviation; obtaining the difference between the dwell time distribution of each page in the abnormal path segment and the corresponding key usage path by calculating the mean square error, etc.). Path Deviation Tag Generation: Each abnormal path segment is assigned a semantic path deviation tag to improve the semantic interpretability of the abnormal path segment for understanding user intent. The tag is generated by combining the following factors, including the degree of deviation from the target (whether the behavioral target has changed), the degree of difference in behavioral content (whether the function used has changed significantly), and the relative offset of the behavioral position (the starting point of the offset behavior in the critical path). Typical tags include: target conversion interruption type (the path is interrupted and the target has not been achieved, indicating that the user may have encountered obstacles), repeated function exploration type (repeatedly clicking on the same type of function, indicating that the user may have uncertainty or confusion about a certain function), and jump loop trap type (the path has multiple jump loops, indicating that the path logic may be closed or the navigation design is unclear). Output of abnormal path segments: The final output is a set of abnormal path segments, the deviation index corresponding to each segment, and the path deviation label of each segment.
[0046] S4. Perform structural mapping and aggregation based on abnormal path fragments to obtain local behavior clusters.
[0047] In this embodiment, the structure mapping aggregation includes: In step S3, multiple abnormal path segments were identified, reflecting abnormal user behavior during use. However, these abnormal paths are complex in origin, numerous, and scattered in distribution. Direct analysis not only fails to identify common problems but also cannot be effectively used for subsequent design optimization. Therefore, it is necessary to perform structural mapping and aggregation on these abnormal path segments to facilitate the identification of problem sources.
[0048] The purpose of this step is to perform structural mapping and aggregation based on the obtained abnormal path fragments, grouping paths with similar operational structures or behavioral patterns together. This identifies "structural operational obstacles" and "high-frequency abnormal structures" commonly encountered by users in the product's functional flow, aggregating representative local behavioral clusters to discover structural problem points, such as redundant jumps and complex page backs. This provides targeted input for subsequent scoring and product optimization, weakens path position dependence, emphasizes the essence of path structure, and avoids misjudgments caused by path disturbances when analyzing anomalies using traditional methods based on path frequency or time sorting.
[0049] The implementation methods for this step include: Structural Behavioral Feature Vector Extraction: Based on the obtained abnormal path fragments, relevant behavioral features are obtained from the feature extraction in step S1, including behavior type frequency (reflecting the user's operational tendencies over a period of time), behavior path length (the number of operations performed by the user in one behavior, i.e., the hierarchical span from the start page to the end page), average dwell time (used to measure page attention), and behavior termination interval (used to infer the user's activity status and churn risk). All feature vectors are numerically normalized (e.g., minimum-maximum normalization) to ensure that each dimension is comparable under the same scale. Finally, all features are concatenated into a unified feature vector. Structural behavior vector aggregation modeling: The feature vectors of all abnormal path segments are combined into a sample set for clustering modeling. Unsupervised clustering algorithms based on distance or density (such as DBSCAN, HDBSCAN, K-means++, etc.) are used to aggregate the structure of this set to generate multiple local behavior clusters. Each cluster contains abnormal path segments with similar structural behaviors. Cluster analysis and semantic tagging: First, frequency analysis is performed on common behavioral combinations (such as certain jump paths, repeated returns, etc.) in each cluster type to count the frequency of patterns in each cluster. Then, based on behavioral characteristics, operational intent analysis is performed to determine the behavioral tendency of the cluster and construct semantic explanation tags for each type of abnormal behavior. For example, a high frequency of page returns and a short average path length indicate "frequent return-type disorientation behavior," which suggests that the product may have a confusing page structure or navigation logic, making it difficult for users to clearly identify the operation path. Repeatedly clicking the same function button without any subsequent action indicates "single-operation type invalid path behavior," which suggests that the product may have issues such as button function logic not working or unclear operation feedback. Starting from the homepage and repeatedly jumping to the recharge page indicates "goal-oriented payment intent behavior," which suggests that the product may have issues such as a complex payment process or insufficient compatibility.
[0050] S5. Quantify the risk based on local behavioral clusters to obtain the risk score results.
[0051] In this embodiment, the risk quantification includes: In step S4, abnormal path fragments are aggregated to obtain local behavior clusters. Each cluster reflects the concentrated deviation behavior of a certain type of user in certain functions or operations. The purpose of this step is to use a risk scoring model to quantify the risk of these clusters based on the obtained local behavior clusters and obtain risk scoring results, so as to quantify the potential negative impact of the behavior represented by each cluster on product stability, user experience, etc.
[0052] Furthermore, the risk scoring model includes:
[0053] in: Cluster The comprehensive risk score reflects the overall impact of the cluster on product risk; the higher the value, the higher the risk. Cluster The number of users, the number of users included in the cluster, is greater than 0; The total number of users across all clusters is greater than 0. Cluster The behavioral risk score reflects the magnitude of the risk that user behavior in the cluster poses to the product; the higher the value, the higher the risk. The weighting coefficient for the proportion of users in the cluster ranges from 0 to 1, and the sum of the weighting coefficients is 1. It is set according to the degree of concern about the severity and spread of the risk. For example, if there is a large-scale and frequent occurrence of low-risk operations, the impact of the spread on the overall risk score will be increased. and These values can be set to 0.7 and 0.3 respectively to prevent minor errors from accumulating into major problems. If a small number of users experience an extremely high failure rate, the severity rating's impact on the overall risk score can be increased. and These can be set to 0.2 and 0.8 respectively, to place greater emphasis on the severity of the anomaly, or to balance the impact of both. and Both can be set to 0.5; : The weighting coefficient for behavioral risk scores, ranging from 0 to 1.
[0054] Furthermore, the behavioral risk score includes:
[0055] in: Cluster The function failure rate, ranging from 0 to 1, is the probability of failure when a user uses a function. It reflects the explicit problems in the product, such as system crashes and function abnormalities. Cluster The path deviation, ranging from 0 to 1, reflects the degree of deviation between the cluster path and the standard critical path. The standard critical path is the standard business process defined by the product design specifications or functional process documents, which reflects the expected functional execution path of the product and is the embodiment of the ideal design process. It is different from the critical usage path that reflects the mainstream user behavior. Cluster The behavioral complexity ranges from 0 to 1, reflecting the complexity of the user's operation steps; The weighting coefficient for the functional failure rate ranges from 0 to 1, and the sum of the weighting coefficients is 1. It is set according to the product's usage requirements. If the product has poor fault tolerance and failure is intolerable, then the weighting coefficient is increased. If the product emphasizes the standardization of the path, then improve... If complex paths cause user churn or cognitive load issues, then improve... For example, for trading platforms where outages have a significant impact, , and These values can be set to 0.6, 0.2, and 0.2 respectively to mitigate the impact of function failures, particularly for work order approvals that emphasize behavioral path control. , and These values can be set to 0.3, 0.5, and 0.2 respectively to mitigate the impact of user behavior path deviations. For end-user-oriented products, the settings should not be overly complex. , and These can be set to 0.2, 0.2, and 0.6 respectively to increase the impact of behavioral complexity. For products without obvious points of interest, , and These values can be set to 0.4, 0.3, and 0.3 respectively, to balance the various dimensions. The weighting coefficient for path deviation, ranging from 0 to 1; : Weighting coefficient for behavioral complexity, ranging from 0 to 1; The cluster's behavioral risk score is obtained by comprehensively calculating the functional failure rate, path deviation, and behavioral complexity.
[0056] Furthermore, the failure rate of the function includes:
[0057] in: :Function The failure rate is the ratio of the number of times a user's operation failed to execute to the total number of times the function was executed in the cluster. Cluster The collection of all functions involved, that is, all functions involved in user operation behavior; Cluster The number of functions is used to calculate the average failure rate; The functional failure rate is obtained by calculating the failure rate of all functions involved in the cluster and taking the average.
[0058] Furthermore, the path deviation includes:
[0059] in: : Abnormal path Standard Critical Path The edit distance reflects the difference between the abnormal path and the standard critical path; Standard Critical Path The length of the edit distance, i.e., the maximum value of the edit distance, is used for normalization; Cluster The set of abnormal paths, that is, the set of all abnormal path segments contained in the cluster; Cluster The number of abnormal paths; The path deviation is obtained by summing and averaging the normalized edit distances between abnormal paths and standard critical paths in the cluster.
[0060] Furthermore, the behavioral complexity includes:
[0061] in: The Sigmoid function is used to normalize values between 0 and 1 to prevent situations where the length of an abnormal path exceeds the standard critical path, resulting in a complexity greater than 1. ; : Function sensitivity adjustment factor, usually 3 to 5; The behavioral complexity is obtained by calculating the ratio of the length of the abnormal path to the length of the standard critical path, normalizing it, and averaging it.
[0062] This model describes how to calculate the risk score for each local behavior cluster based on the characteristics of the user's actual behavior path in the local behavior cluster.
[0063] Traditional risk quantification methods often employ fixed rules or threshold judgments, isolated event assessments, and other approaches. They typically use fixed parameters or single weights, such as single failure rate scoring methods, which ignore behavioral structures and interaction complexity. They are highly susceptible to extreme data and exhibit large fluctuations. In contrast, this model uses normalized multi-dimensional indicators for comprehensive scoring, improving the comprehensiveness and comparability of risk scoring, enhancing model stability, and using adjustable weights to adapt to different product needs and risk preferences, thereby increasing the flexibility and applicability of risk scoring.
[0064] The steps for using the above model include: Data Acquisition: Obtain local behavioral clusters and various data information of the clusters from step S4, including the collection of abnormal paths of each cluster. The collection of all functions involved in each cluster Failure rate of each function execution in the cluster Number of users in the cluster , wait; Parameter calculation: Based on the acquired data, calculate the parameters in the model, including the edit distance between the abnormal path and the standard critical path. Functional failure rate Path deviation Behavioral complexity Behavioral risk score ; Risk scoring: A comprehensive risk score is obtained by calculating the calculated parameters and other data. This serves as data support for subsequent optimization decisions.
[0065] S6. Compare versions based on risk scoring results to obtain optimization results.
[0066] In this embodiment, the version comparison includes: In step S5, the risk scores of each cluster under the current version of the product are calculated, and the risk score results are obtained to support the optimization of the product. However, in the user behavior path analysis of complex products, relying solely on the risk score results of a single version is insufficient to determine whether structural adjustments or optimizations have truly improved the user experience or reduced the risk of anomalies. Therefore, a method is needed to compare and analyze the risk score results between different versions to measure the optimization effect.
[0067] The purpose of this step is to compare the risk scores between different product versions based on the obtained risk scores, identify whether product optimization has truly played a positive role in reducing risk, simplifying operations, and providing clear guidance, determine the merits of different versions and verify the optimization effect, thereby quantifying the differences before and after product optimization, identifying whether product optimization has led to an improvement in risk scores, and which product adjustments are more effective, providing a basis for continuous product optimization.
[0068] The implementation methods for this step include: Version identification and risk score result acquisition: For two product versions that need to be compared (such as version V1 and version V2), execute steps S1 to S5 respectively to obtain the risk score results of each local behavior cluster in each version; Cluster alignment and matching: Based on semantic tag similarity (e.g., both are "single-operation type" or "behavior-oriented type"), each cluster in the two product versions is matched one-to-one to ensure the comparability of the comparison results; Rating Difference Calculation and Normalization: For each pair of matched rating objects, calculate the rating difference between the new version and the old version of the product (obtained by subtracting the risk score of the old version from the risk score of the new version), and use the tanh function for normalization (while preserving the directionality of the difference, suppressing the interference of extreme rating changes on the overall score), so that the rating difference ranges between (-1, 1), which is used to quantify the increase or decrease in risk score between product versions; Cluster optimization trend judgment: The optimization effect is judged based on the positive and negative signs and magnitude of the score difference. If the score difference is less than 0, it means that the new version performs better in the cluster, the optimization is successful, and the risk is reduced. If the score difference is greater than 0, it means that the optimization of the new version has not achieved the expected results or has a negative effect, and the risk is increased. Comprehensive optimization result judgment: Based on product and business needs, assign weights to each cluster, calculate the weighted sum of the score differences between the old and new versions of each cluster, evaluate the overall effectiveness of product optimization, and use it to guide product decisions and version rollback.
[0069] Example 2: Reference Figure 3 This is the second embodiment of the present invention, which provides a product optimization system based on user behavior data mining, including the following functional modules: Feature extraction module: Extracts representative behavioral features from the collected historical interaction data between users and products, i.e., user behavior record sequences, and constructs a set of user behavior features; Path identification module: Based on user behavior characteristics, it identifies critical paths, measures the importance of paths, identifies representative key usage paths, and determines the core processes when users use the product, which serve as the basis for subsequent analysis. Anomaly Analysis Module: By identifying anomalous path segments through key usage paths, each anomalous path segment is assigned a semantic label, thereby elevating low-level path deviation comparison analysis to high-level user intent understanding, making the results of path deviation comparison analysis more interpretable and applicable. The structure aggregation module performs structure mapping and aggregation on abnormal path segments through unsupervised clustering. Paths with similar operational structures or behavioral patterns are grouped together to form multiple local behavioral clusters, which are used to identify typical but abnormal behavioral patterns for targeted analysis. Risk scoring module: Analyzes each local behavioral cluster, calculates the comprehensive risk score for each cluster using a risk scoring model, and generates a risk scoring result; Optimization Assessment Module: Compare and analyze the risk scoring results of the above different versions of the structure to determine whether the structural adjustments or functional optimizations have achieved positive results.
[0070] Example 3: This embodiment also provides a computer device applicable to a product optimization system based on user behavior data mining, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize a product optimization system based on user behavior data mining as proposed in the above embodiment.
[0071] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements a product optimization system based on user behavior data mining as proposed in the above embodiments.
[0072] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0073] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0074] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0075] More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0076] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0077] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A product optimization system based on user behavior data mining, characterized in that, Includes the following steps: S1. Extract user behavior record sequences to obtain user behavior features; S2. Identify critical paths based on user behavior characteristics to obtain critical usage paths; S3. Perform path deviation comparison analysis based on the key usage path to obtain abnormal path segments; S4. Perform structural mapping and aggregation based on abnormal path fragments to obtain local behavior clusters; S5. Quantify the risk based on local behavioral clusters to obtain the risk score results; S6. Compare versions based on risk scoring results to obtain optimization results.
2. The product optimization system based on user behavior data mining according to claim 1, characterized in that, The path deviation comparison analysis in step S3 includes: Based on the obtained key usage paths, abnormal path segments in the user's actual behavior path are identified, and each abnormal path segment is assigned a semantic label, thereby elevating low-level path deviation comparison analysis to high-level user intent understanding, making the results of path deviation comparison analysis more interpretable and valuable for application.
3. The product optimization system based on user behavior data mining according to claim 1, characterized in that, The structure mapping aggregation in step S4 includes: The obtained abnormal path fragments are subjected to structural mapping and aggregation. Abnormal path fragments with similarity are merged together to aggregate representative local behavior clusters, thereby discovering structural problem points and providing targeted input for subsequent scoring and product optimization.
4. The product optimization system based on user behavior data mining according to claim 1, characterized in that, The risk quantification in step S5 includes: Based on the obtained local behavioral clusters, a risk scoring model is used to quantify the risk of these clusters and obtain risk score results to quantify potential negative impacts.
5. A product optimization system based on user behavior data mining according to claim 4, characterized in that, The risk scoring model includes: , in: Cluster Comprehensive risk score; Cluster The number of users; Total number of users across all clusters; Cluster Behavioral risk score; The weighting coefficient representing the percentage of users in the cluster; : Weighting coefficients for behavioral risk scores.
6. A product optimization system based on user behavior data mining according to claim 5, characterized in that, The behavioral risk score includes: , in: Cluster Functional failure rate; Cluster Path deviation; Cluster The complexity of behavior; Weighting coefficient for functional failure rate; : Weighting coefficient for path deviation; Weighting coefficients for behavioral complexity; The cluster's behavioral risk score is obtained by comprehensively calculating the functional failure rate, path deviation, and behavioral complexity.
7. A product optimization system based on user behavior data mining according to claim 6, characterized in that, The failure rate of the function includes: , in: :Function The failure rate; Cluster The collection of all functions involved; Cluster The number of functions; The functional failure rate is obtained by calculating the failure rate of all functions involved in the cluster and taking the average.
8. A product optimization system based on user behavior data mining according to claim 6, characterized in that, The path deviation includes: , in: : Abnormal path Standard Critical Path Edit distance; Standard Critical Path Length; Cluster A collection of abnormal paths; Cluster The number of abnormal paths; The path deviation is obtained by summing and averaging the normalized edit distances between abnormal paths and standard critical paths in the cluster.
9. A product optimization system based on user behavior data mining according to claim 6, characterized in that, The behavioral complexity includes: , in: : Sigmoid function; : Function sensitivity adjustment factor; The behavioral complexity is obtained by calculating the ratio of the length of the abnormal path to the length of the standard critical path, normalizing it, and averaging it.
10. A product optimization system based on user behavior data mining according to claim 1, characterized in that, Includes the following functional modules: The module includes a feature extraction module, a path recognition module, an anomaly analysis module, a structure aggregation module, a risk scoring module, and an optimization evaluation module.