Software function iteration method and device

By constructing a hybrid model for anomaly detection and missing value filling, combining generative adversarial networks and clustering algorithms for data cleaning, building a multi-model pool for multimodal fusion feature extraction, and determining the association rules of software function usage based on frequent item set mining algorithms, the shortcomings of traditional software function statistical methods in multi-dimensional analysis are solved, and the efficiency and accuracy of software function iteration are improved.

CN120669957APending Publication Date: 2025-09-19BEIJING C H L ROBOTICS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510512752.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional software function statistics methods lack multi-dimensional analysis and cannot accurately capture users' usage habits and needs in different time periods and scenarios, resulting in software developers lacking sufficient basis for optimizing user experience and improving software performance.

Method used

Construct a hybrid model for anomaly detection and missing value filling, combine generative adversarial networks and clustering algorithms for data cleaning, build a multi-model pool for multimodal fusion feature extraction, and determine the association rules of software function usage based on frequent item set mining algorithms.

Benefits of technology

It improves the efficiency and accuracy of software function iteration, enables a better understanding of user needs and behavior patterns, and makes more reasonable function iteration decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669957A_ABST
    Figure CN120669957A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a software function iteration method and device, and the method comprises the steps: building a hybrid model through a preset isolated forest model and a preset time sequence model, carrying out the anomaly detection of initial user behavior features, removing abnormal values, building a generator and discriminator frame, and carrying out the recognition of the abnormal values. Carrying out missing value filling on the user behavior characteristics according to the generative adversarial network, complementing missing values, cleaning the processed user behavior characteristics based on a preset clustering algorithm, constructing a multi-model pool, carrying out multi-modal fusion on the preprocessed user behavior characteristics, and carrying out multi-modal fusion on the preprocessed user behavior characteristics; combining the multi-modal fusion feature with a preset service label to determine a corresponding user behavior data set, performing frequent item mining on the user behavior data set based on a frequent item set mining algorithm, determining a corresponding software function use association rule, and performing software function iteration according to the software function use association rule to obtain a software function set; the efficiency and accuracy of software functions can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically to a software function iteration method and device. Background Art

[0002] In the field of software development, improving user experience has always been a key focus for developers. To optimize user experience and enhance software performance, in-depth statistics and analysis of software feature usage are often required. However, traditional statistical methods have numerous shortcomings, limiting the effectiveness and accuracy of optimization decisions.

[0003] Traditional statistical methods typically provide only a simple summary of click counts, a single-dimensional statistical information that falls far short of meeting the needs of modern software development. The lack of comprehensive analysis across multiple dimensions, such as usage duration, usage scenarios, and user groups, leaves developers with insufficient basis for making decisions about feature iterations. Specifically, traditional methods are unable to accurately capture user habits and needs across different time periods and scenarios, nor can they meticulously differentiate the preferences and behavior patterns of different user groups.

[0004] Therefore, under current technological conditions, software developers face numerous challenges in optimizing user experience and improving software performance. A more comprehensive, accurate, and efficient statistical and analytical method is urgently needed to better understand user needs and behavior patterns, thereby making more reasonable decisions about feature iteration. Summary of the Invention

[0005] In response to the problems in the prior art, the present application provides a software function iteration method and device, which can improve the efficiency and accuracy of software functions.

[0006] In order to solve at least one of the above problems, the present application provides the following technical solutions:

[0007] In a first aspect, the present application provides a software function iteration method, comprising:

[0008] Acquire historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data to determine a corresponding initial user behavior feature, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior feature based on the hybrid model, determine a corresponding outlier, remove the outlier, and determine a corresponding first user behavior feature;

[0009] Constructing a generator and discriminator framework, performing conditional constraint training on the generator framework based on user behavior consistency constraints, determining a corresponding generative adversarial network, filling missing values ​​for the first user behavior feature using the generative adversarial network, determining a corresponding second user behavior feature, cleaning the second user behavior feature based on a preset clustering algorithm, and determining a corresponding preprocessed user behavior feature;

[0010] Construct a multi-model pool, input the preprocessed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

[0011] Furthermore, the step of constructing a hybrid model based on a preset isolation forest model and a preset time series model, performing an anomaly detection operation on the initial user behavior feature based on the hybrid model, determining corresponding outliers and removing the outliers, and determining the corresponding first user behavior feature includes:

[0012] Performing a global anomaly location operation on the initial user behavior features according to a preset isolation forest model to determine the corresponding global sparse anomaly;

[0013] A local anomaly location operation is performed on the initial user behavior features according to a preset time series model to determine the corresponding local time series anomaly, wherein the preset time series model adopts an attention mechanism to locate key abnormal time points in the user behavior sequence.

[0014] The global sparse anomaly and the local time series anomaly are weightedly fused, the scores obtained after the weighted fusion are screened according to a preset anomaly threshold, the corresponding outliers are determined and eliminated, and the corresponding first user behavior feature is determined.

[0015] Furthermore, the construction of the generator and discriminator framework includes:

[0016] Constructing a corresponding generator framework according to the residual fully connected network, wherein the input of the generator framework includes a missing data matrix, a mask matrix and a random noise vector, and the output of the generator framework includes a complete data matrix;

[0017] A corresponding discriminator framework is constructed according to the spectral normalization fully connected layer, where the input of the discriminator framework is the complete data matrix, and the output of the discriminator framework is a scalar probability.

[0018] Furthermore, the generator framework is subjected to conditional constraint training based on the user behavior consistency constraint to determine the corresponding generative adversarial network, including:

[0019] Constructing a user behavior consistency constraint according to preset business rules, constructing a loss function according to the user behavior consistency constraint, a preset reconstruction loss, and a preset adversarial loss, and training the generator framework;

[0020] The discriminator framework is trained according to a preset binary cross entropy loss, and the trained discriminator framework and the trained generator framework are jointly trained to determine a corresponding generative adversarial network.

[0021] Furthermore, the cleaning of the second user behavior feature based on a preset clustering algorithm to determine the corresponding pre-processed user behavior feature includes:

[0022] Using the K-Means++ algorithm to determine the optimal number of clusters according to a preset silhouette coefficient, and determining the corresponding user group according to the optimal number of clusters;

[0023] Based on the percentile of the behavior distribution of each user group, a corresponding differentiation threshold is determined, and the second user behavior feature is cleaned according to the differentiation threshold to determine the corresponding pre-processed user behavior feature.

[0024] Furthermore, the multi-model pool is constructed, the pre-processed user behavior features are input into the multi-model pool, weights are assigned to the models in the multi-model pool according to a meta-learning algorithm, and the multimodal fusion features output by the multi-model pool are determined, including:

[0025] According to the different characteristics of user behavior data, complementary models are selected to construct a multi-model pool, wherein the multi-model pool includes at least one of a decision tree model, a long short-term memory network, a graph neural network, and a Bayesian network;

[0026] Based on the meta-learning algorithm, the weights of each model in the multi-model pool are dynamically allocated to output multimodal fusion features.

[0027] Furthermore, frequent item mining is performed on the user behavior dataset based on a frequent item set mining algorithm to determine corresponding software function usage association rules, including:

[0028] Adopting the frequent item set mining algorithm, the pseudo rules are filtered according to the preset support, confidence and lift parameters to determine the corresponding strong correlation association rules;

[0029] The contribution of the target variables in the strongly correlated association rules is interpreted according to a preset SHAP analysis algorithm, and the corresponding software function usage association rule is determined according to the target variable with the highest contribution.

[0030] In a second aspect, the present application provides a software function iteration device, comprising:

[0031] A user behavior data outlier processing module is configured to obtain historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data, determine a corresponding initial user behavior feature, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior feature based on the hybrid model, determine a corresponding outlier, remove the outlier, and determine a corresponding first user behavior feature;

[0032] A user behavior data filling and cleaning module is used to build a generator and discriminator framework, perform conditional constraint training on the generator framework based on user behavior consistency constraints, determine a corresponding generative adversarial network, fill missing values ​​for the first user behavior feature based on the generative adversarial network, determine a corresponding second user behavior feature, clean the second user behavior feature based on a preset clustering algorithm, and determine a corresponding preprocessed user behavior feature;

[0033] The software function usage association rule determination module is used to build a multi-model pool, input the preprocessed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

[0034] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the software function iteration method when executing the program.

[0035] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the software function iteration method when executed by a processor.

[0036] In a fifth aspect, the present application provides a computer program product, comprising a computer program / instruction, which implements the steps of the software function iteration method when executed by a processor.

[0037] It can be seen from the above technical solution that the present application provides a software function iteration method and device, which constructs a hybrid model by presetting an isolation forest model and a preset time series model, performs anomaly detection on the initial user behavior characteristics, eliminates the outliers, and then constructs a generator and discriminator framework, fills the missing values ​​of the user behavior characteristics according to the generative adversarial network, makes up the missing values, cleans the processed user behavior characteristics based on the preset clustering algorithm, constructs a multi-model pool to perform multimodal fusion on the preprocessed user behavior characteristics, combines the multimodal fusion features with the preset business labels to determine the corresponding user behavior data set, performs frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determines the corresponding software function usage association rules, and performs software function iteration according to the software function usage association rules, thereby improving the efficiency and accuracy of the software function. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0039] Figure 1 This is one of the flowcharts of the software function iteration method in the embodiment of the present application;

[0040] Figure 2 This is a second flowchart of the software function iteration method in an embodiment of the present application;

[0041] Figure 3 This is the third flow chart of the software function iteration method in the embodiment of the present application;

[0042] Figure 4 This is a fourth flowchart of the software function iteration method in an embodiment of the present application;

[0043] Figure 5 This is the fifth flowchart of the software function iteration method in the embodiment of the present application;

[0044] Figure 6 This is the sixth flowchart of the software function iteration method in the embodiment of the present application;

[0045] Figure 7 This is the seventh flowchart of the software function iteration method in the embodiment of the present application;

[0046] Figure 8 This is a structural diagram of a software function iteration device in an embodiment of the present application;

[0047] Figure 9 Schematic diagram of the structure of the electronic device in the embodiment of the present application.

[0048] Reference numerals:

[0049] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver program storage unit 9144, antenna 9111, speaker 9131, microphone 9132. DETAILED DESCRIPTION

[0050] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0051] The acquisition, storage, use, and processing of data in this application's technical solution comply with relevant national laws and regulations.

[0052] Considering the many challenges faced by software developers in optimizing user experience and improving software performance, the present application provides a method and device for software function iteration, which constructs a hybrid model by presetting an isolation forest model and a preset time series model, performs anomaly detection on the initial user behavior characteristics, removes the outliers, and then constructs a generator and discriminator framework. The missing values ​​of the user behavior characteristics are filled according to the generative adversarial network, the missing values ​​are supplemented, and the processed user behavior characteristics are cleaned based on the preset clustering algorithm. A multi-model pool is constructed to perform multimodal fusion on the preprocessed user behavior characteristics, and the multimodal fusion features are combined with the preset business labels to determine the corresponding user behavior data set. The user behavior data set is frequently mined based on the frequent item set mining algorithm to determine the corresponding software function usage association rules. The software function is iterated according to the software function usage association rules, thereby improving the efficiency and accuracy of the software function.

[0053] In order to improve the efficiency and accuracy of software functions, this application provides an embodiment of a software function iteration method, see Figure 1 , the software function iteration method specifically includes the following contents:

[0054] Step S101: Acquire historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data to determine a corresponding initial user behavior feature, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior feature based on the hybrid model, determine corresponding outliers, remove the outliers, and determine a corresponding first user behavior feature;

[0055] Optionally, in this embodiment, first, background introduction, in the software function design stage, improving user experience has always been the focus of developers. Usually, the optimization design of software functions is based on manual statistics of commonly used software functions and manual analysis and optimization. However, this single-dimensional statistical information is far from meeting the needs of modern software development. The lack of comprehensive analysis of multiple dimensions such as usage time, usage scenarios, and user groups has led to a lack of sufficient basis for developers to make function iteration decisions. With the development of technical means, this embodiment proposes a more comprehensive, accurate, and efficient statistical and analysis method for generating software function association rules to improve user experience and software usage efficiency. In the process of data statistics and analysis, it is necessary to ensure that the analyzed data has clarity, correctness, and availability in order to ensure the efficiency of statistical analysis, so data preprocessing is necessary.

[0056] Optionally, in this embodiment, this step is the first step in data preprocessing, which is the process of eliminating outliers in historical user behavior data based on artificial intelligence technology, thereby improving the accuracy of the data and laying the foundation for subsequent statistical analysis based on the data.

[0057] Specifically, the original user function click logs are collected from the target software, including fields such as timestamp, user ID, and function module. In software user behavior data analysis, the original data often contains noise, including outliers. Traditional rule cleaning (such as threshold filtering and mean filling) requires manual definition of rules, which is difficult to adapt to complex scenarios and inefficient. This embodiment solves the shortcomings of the existing technology as follows:

[0058] First, this step extracts multi-dimensional features of user behavior, including time series features, statistical features, contextual features, etc. For example, time series features include click frequency (times / minute) and session interval (such as the interval between two function uses); statistical features include weekly active days and the standard deviation of daily usage time; and contextual features include function usage paths (such as whether "A→B→C" is common). It is understandable that combining time series, statistical, and contextual features can enhance sensitivity to "hidden anomalies" (such as normal click volume but abnormal path) and improve the accuracy of data anomaly identification.

[0059] Next, we trained a hybrid model of Isolation Forest and Time Series LSTM to identify complex anomaly patterns within the multi-dimensional features of user behavior, avoiding the limitations of traditional single-detection methods. Isolation Forest processes global sparse anomalies, while LSTM captures localized time series anomalies.

[0060] Specifically, Isolation Forest training randomly samples a subset of data, constructs a binary tree to segment outliers, and outputs a global anomaly score (0-1), where the higher the score, the more abnormal it is.

[0061] Specifically, LSTM training inputs the user behavior time series sequence, combines the two-layer LSTM and attention mechanism to identify key time points, and outputs the time series reconstruction error. The larger the error, the higher the probability of local anomaly.

[0062] After the weights are determined through grid search optimization, the outputs of the isolation forest and LSTM are weightedly fused to obtain the final anomaly value. The anomaly value is judged based on the anomaly threshold. Values ​​above the threshold are considered anomalies, and the user behavior features corresponding to the anomaly values ​​are removed.

[0063] Let's take an example to illustrate: suppose that the original user behavior log is received.

[0064] The first step is to extract features from the original user behavior logs.

[0065] Extract its temporal features: click frequency (such as the number of operations per minute); session interval (the interval between two adjacent function uses); time window statistics (such as the number of times a function was used in the past hour).

[0066] Statistical characteristics: user activity (number of active days per week); behavioral dispersion (such as standard deviation of usage time).

[0067] Contextual features: Function usage paths (e.g., whether "Export → Save → Close" is compliant); cross-module associations (e.g., whether the "3D Modeling" module is usually followed by "Rendering"). For example, in modeling software, the number of "Preview" operations before the user's "Save" operation is extracted as a key feature, and abnormal paths (e.g., direct save) will be flagged.

[0068] In the second step, a hybrid architecture of isolation forest and LSTM is constructed to perform weighted voting based on the outputs of the two.

[0069] Assume that the isolation forest score (weight 0.6) + LSTM error (weight 0.4) is used to obtain an anomaly value. This anomaly value is compared with the preset threshold. Assuming the threshold is 0.8, a total score greater than 0.8 is considered an anomaly.

[0070] An embodiment of the present application is as follows: Isolation Forest discovers that a user clicks the "Undo" button 500 times in a single day (global anomaly); LSTM detects that the user has been in an "Undo→Redo" cycle for 10 consecutive minutes (local anomaly); after fusion, it is determined to be malicious script behavior.

[0071] Step S102: Constructing a generator and discriminator framework, performing conditional constraint training on the generator framework based on user behavior consistency constraints, determining a corresponding generative adversarial network, filling missing values ​​for the first user behavior feature based on the generative adversarial network, determining a corresponding second user behavior feature, cleaning the second user behavior feature based on a preset clustering algorithm, and determining a corresponding preprocessed user behavior feature;

[0072] Optionally, in this embodiment, this step is the second and third steps in data preprocessing, which uses artificial intelligence technology to supplement missing values ​​and clean user behavior data, improve data accuracy and availability, and lay the foundation for subsequent statistical analysis based on the data.

[0073] Optionally, in this embodiment, in the analysis of user behavior data, missing values ​​(such as unrecorded usage time, interrupted session logs) will seriously affect the accuracy of decision-making. Therefore, this application proposes the following solutions:

[0074] The second step is to supplement the missing values ​​of the data. Specifically, business logic is built and the missing values ​​are supplemented based on the generative adversarial network after business constraints.

[0075] First, we design the corresponding generator and discriminator architectures. The generator inputs data with missing values ​​and outputs complete data, filling in the missing parts. The discriminator distinguishes generated data from real data, driving the generator to improve the filling quality.

[0076] Specifically, the generator network is constructed: the input of the generator is the missing data matrix, the mask matrix and the random noise vector, and the output of the generator is the complete data matrix.

[0077] The generator network architecture uses a fully connected network at the core layer for tabular data processing, employing residual connections (ResNet) to prevent vanishing gradients, and processes numerical and categorical features in two separate channels. Preferably, the core layer for time series data is modified to use a Transformer network, employing an attention mechanism to capture long-term dependencies and performing sinusoidal position encoding on timestamps.

[0078] The generator's loss function design specifically incorporates business rule constraints. For example, "preview must occur before saving" can be implemented through a rule-based loss function. Embedding business rules ensures that the generated data complies with logic. The generator is trained based on business rule constraints, reconstruction loss, and adversarial loss. The reconstruction loss only calculates the error for the non-missing parts, while the adversarial loss optimizes the generator through feedback from the discriminator.

[0079] Specifically, the discriminator network is constructed as follows: the input of the discriminator is the complete data matrix, and the input is the scalar probability. Among them, the scalar probability p∈[0,1] (1=real, 0=generated).

[0080] The discriminator network architecture adds spectral normalization to the core layer of tabular data processing to stabilize training, while using gradient penalty (WGAN-GP) to avoid mode collapse.

[0081] The loss function of the discriminator is designed to be a binary cross entropy loss.

[0082] After constructing the generator and discriminator, we alternately train the generator and discriminator to obtain a trained generative adversarial network for filling in missing data. It can be understood that the rule-based constraint loss ensures the rationality of the filling.

[0083] Optionally, in this embodiment, after data anomaly removal and missing value processing, the obtained high-quality data is cleaned and adapted.

[0084] The third step is to clean the data. Specifically, the task of this stage is to set dynamic thresholds for users based on clustering algorithms to formulate group-specific anomaly analysis rules.

[0085] First, based on the K-Means++ clustering algorithm, we improved the initial center point selection to avoid the local optimality problem of traditional K-Means. We used the silhouette coefficient to determine the optimal number of clusters, and based on this number of clusters, we first identified all user groups. Next, for each group, we calculated the 99th percentile of key indicators and dynamically adjusted the threshold to avoid misjudgments caused by fixed values.

[0086] It can be understood that this step of dividing users into groups based on the clustering algorithm and setting dynamic thresholds is to work in conjunction with the anomaly detection in step S101, and the normal threshold of the specific group output by this step is used to correct the results of the anomaly detection module and optimize the threshold rules.

[0087] In step S101-step S102, this embodiment constructs a complete progressive collaborative data preprocessing mode. First, the anomaly detection based on the hybrid model identifies and eliminates noise data, and outputs clean data with high confidence. Secondly, the generative adversarial network based on the rule constraints repairs the incomplete data and outputs a coherent and complete data set. Finally, the dynamic threshold setting based on the clustering algorithm adapts to the analysis needs of different user groups, and the final dynamic anomaly threshold is fed back to the anomaly detection module to optimize the threshold rules, which can maximize the retention of real data and prevent misjudgment.

[0088] Specifically, the application scenarios of this step include:

[0089] 1. Novice users may stay for a long time due to their unfamiliarity with the software, while expert users are efficient and use the software for a shorter period of time. If a fixed threshold (such as "single operation > 1 hour = abnormality") is used for anomaly detection, it will lead to a large number of misjudgments. At this time, the threshold is adjusted dynamically, and novice users and expert users are labeled as groups based on the clustering algorithm. The threshold for the novice group is set to 45 minutes, and the threshold for the expert group is set to 20 minutes. Based on this method, the final data cleaning is performed, and the threshold rules are determined and fed back to the anomaly detection module in step S101 for outlier screening to minimize misjudgments.

[0090] 2. The normal usage time of certain functions (such as 3D rendering) is significantly longer than that of basic functions (such as clicking the Save button) and requires differentiated judgment.

[0091] 3. This embodiment can adapt to changes in data distribution (such as users growing from novices to experts) by updating the clustering model and thresholds every week, providing an analytical basis for subsequent data analysis.

[0092] To illustrate with an example, assume that on an online learning platform, when students watch videos, a fixed threshold cannot distinguish between "attentive viewing" (long stay) and "idle brushing" (abnormally long stay).

[0093] This embodiment uses clustering to identify "active learning type" (frequent pause / replay) and "passive viewing type" (continuous playback), and sets the threshold value as active type single viewing > 2 hours = abnormal, passive type > 1 hour = abnormal. The obtained abnormal values ​​are eliminated or corrected. At the same time, data marked as abnormal can be manually reviewed.

[0094] It is worth noting that in this embodiment, the logical relationship of data processing presents a progressive coordination and cannot be reversed, that is, first identify abnormal data, then correct missing values, and finally cluster and group to set dynamic thresholds.

[0095] Specifically, if missing values ​​are directly filled in for data containing outliers (e.g., GAN), the generator may learn noise patterns (e.g., filling in "abnormal click duration" as normal values). Dynamic thresholds, on the other hand, require the calculation of user clustering features (e.g., average daily usage time per person) based on complete data. If missing values ​​are not corrected, the clustering results may be biased. Finally, outlier correction based on dynamic thresholds is part of the closed-loop linkage in the data preprocessing stage of this embodiment.

[0096] Taking the embodiment in the above-mentioned application scenario 1 as an example, after dividing the two groups, arbitrary rules can be customized based on the groups' needs for using the software. The final effect may be to strengthen short-term and high-frequency click detection for the "novice group" (to prevent misoperation); and relax long-term usage judgment for the "expert group" (to avoid misjudgment).

[0097] Step S103: Construct a multi-model pool, input the preprocessed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

[0098] Optionally, in this embodiment, after the above steps S101 and S102, user behavior characteristics that have undergone data preprocessing have been obtained. This characteristic can be used for data analysis, and software function usage association rules are generated based on the analysis results to perform functional iteration on the software.

[0099] Specifically, based on the different characteristics of user behavior data (such as temporal sequence, sparsity, and category imbalance), complementary models are selected to build a model pool for feature analysis and output multi-dimensional fusion features. The models in the model pool include but are not limited to:

[0100] Decision Tree (CART / Random Forest): used to process discrete features (such as function click classification);

[0101] LSTM (Long Short-Term Memory Network): used to capture the temporal dependencies of user operation sequences;

[0102] Graph Neural Networks (GNNs): Analyze the topological relationships between user-function interactions (e.g., “Users of function A often use function B together”);

[0103] Bayesian networks: Infer probabilistic associations between feature usage and user attributes (e.g., “the probability that the designer community prefers module X is 78%”).

[0104] The aforementioned models are candidate models. A meta-learning algorithm adjusts the output weights of each model based on real-time data distribution, screening the prediction results of one or more models with the highest weight ratios for fusion output. This fusion output approach can capture complex nonlinear relationships in user behavior (such as the dynamic correlation between feature usage frequency and user retention).

[0105] Next, the output of the model pool is combined with business tags (such as "high-value users" and "churn risk") to mine frequent itemsets using the Apriori algorithm. The goal of this step is to extract strong association rules from user behavior data, such as "feature usage patterns → business outcomes" (e.g., "high frequency of feature A → improved user retention").

[0106] Specifically, based on the frequent itemset mining algorithm, the minimum support and confidence are set and the lift is used to filter out pseudo rules (lift>1 indicates positive correlation). The rules of the data set with business labels are filtered to retain the rules that are strongly related to the business goals (such as associating "function usage" with "retention rate"), and redundant rules are eliminated (such as only one of "function A→function B" and "function B→function A" is retained).

[0107] Transform rules into executable recommendations based on business semantics, for example:

[0108] Rule: IF "Tolerance Analysis" usage time > 3 minutes AND number of errors ≥ 2 → THEN trigger real-time help (confidence level 85%);

[0109] Suggestion: Embed a “Smart Guide Button” in this function.

[0110] Next, to make the association rules output by the model easier to understand and quantify the impact of features on business indicators, we used SHAP value analysis to improve the interpretability of the model.

[0111] Specifically, the interpretability processing quantifies the contribution of each feature to the target variable (such as retention rate) by calculating the SHAP value.

[0112] For example, there are currently 100,000 user session logs, including usage sequence features of 20 functions such as "Tolerance Analysis" and "Sketch Tool"; business tag: user subscription renewal rate.

[0113] FP-Growth discovery rule: IF "Tolerance Analysis" is used ≥ 3 times / week AND single session duration > 5 minutes → THEN renewal probability increases by 25% (lift = 2.1)

[0114] SHAP analysis shows that "Tolerance Analysis" contributes most to renewals (SHAP = 0.22), so a report is generated: "Expert users make extensive use of the tolerance analysis function. We recommend adding an advanced tutorial entry."

[0115] In this way, we deeply explore the relationship between software functions and user needs, aiming to improve the accuracy and efficiency of user experience optimization decisions.

[0116] This example demonstrates how this embodiment preprocesses the user's software operation behavior data and performs statistical analysis based on the preprocessed data to obtain software function usage association rules that are most relevant to user needs, and optimizes the software based on the rules.

[0117] From the above description, it can be seen that the software function iteration method provided in the embodiment of the present application can construct a hybrid model through a preset isolation forest model and a preset time series model, perform anomaly detection on the initial user behavior characteristics, eliminate outliers, and then construct a generator and discriminator framework. The missing values ​​of the user behavior characteristics are filled according to the generative adversarial network, the missing values ​​are supplemented, and the processed user behavior characteristics are cleaned based on the preset clustering algorithm. A multi-model pool is constructed to perform multimodal fusion on the pre-processed user behavior characteristics, and the multimodal fusion features are combined with preset business labels to determine the corresponding user behavior data set. Frequent item mining is performed on the user behavior data set based on the frequent item set mining algorithm to determine the corresponding software function usage association rules, and software function iteration is performed according to the software function usage association rules, thereby improving the efficiency and accuracy of the software function.

[0118] In one embodiment of the software function iteration method of the present application, see Figure 2 , and can also include the following:

[0119] Step S201: performing a global anomaly location operation on the initial user behavior features according to a preset isolation forest model to determine a corresponding global sparse anomaly;

[0120] Step S202: performing a local anomaly location operation on the initial user behavior features according to a preset time series model to determine the corresponding local time series anomaly, wherein the preset time series model adopts an attention mechanism to locate key abnormal time points in the user behavior sequence;

[0121] Step S203: performing weighted fusion on the global sparse anomaly and the local time series anomaly, screening the scores obtained after the weighted fusion according to a preset anomaly threshold, determining corresponding outliers and removing the outliers, and determining the corresponding first user behavior feature.

[0122] Optionally, in this embodiment, this step is a process of eliminating outliers in data preprocessing.

[0123] Specifically, the isolation forest algorithm detects global outliers in user behavior data. The core implementation steps are as follows:

[0124] Data sampling: Randomly extract several subsets (e.g., 256 records per subset) from the user behavior dataset to ensure that the sampling covers different functional modules and user groups.

[0125] Binary tree construction: recursively perform the following operations on each subset:

[0126] Randomly select a feature (such as "click duration" or "session interval");

[0127] Randomly generate a split threshold (e.g. click duration > 300 seconds) and divide the data into left and right subtrees;

[0128] Repeat the splitting until all data points are isolated or the tree depth limit is reached (e.g. depth = 8)

[0129] Anomaly Scoring: Calculate the path length of each data point (i.e., the number of splits from the root node to that point). The shorter the path, the higher the probability of anomaly. A global anomaly score (0-1) is output through normalization. A score > 0.7 is considered an anomaly.

[0130] Specifically, the long short-term memory network (LSTM) is combined with the attention mechanism to detect local abnormal patterns in user behavior time series data. The specific implementation steps are as follows:

[0131] Data sampling: input user behavior time series (such as the number of function clicks per minute), group them by user ID and standardize them; construct a sliding window (such as window length = 30 minutes, step length = 5 minutes) to generate training samples.

[0132] Model construction: Network structure: 2-layer LSTM (number of hidden layer units = 64) + attention mechanism layer (calculate the weight of each time point), the output layer reconstructs the input sequence by a fully connected layer, and the loss function uses the mean square error (MSE).

[0133] Training phase: input normal behavior sequence and minimize reconstruction error

[0134] Anomaly detection: Input the sequence to be tested, calculate the reconstruction error, and determine if the error exceeds the threshold (such as MSE>3σ) as an anomaly.

[0135] Specifically, after determining the weights through grid search optimization, the outputs of the isolation forest and LSTM are weightedly fused to obtain the final anomaly value. The anomaly value is judged based on the anomaly threshold. Values ​​above the threshold are considered anomalies, and the user behavior features corresponding to the anomaly values ​​are removed.

[0136] Let's take an example to illustrate: suppose that the original user behavior log is received.

[0137] The first step is to extract features from the original user behavior logs.

[0138] Extract its temporal features: click frequency (such as the number of operations per minute); session interval (the interval between two adjacent function uses); time window statistics (such as the number of times a function was used in the past hour).

[0139] Statistical characteristics: user activity (number of active days per week); behavioral dispersion (such as standard deviation of usage time).

[0140] Contextual features: Function usage paths (e.g., whether "Export → Save → Close" is compliant); cross-module associations (e.g., whether the "3D Modeling" module is usually followed by "Rendering"). For example, in modeling software, the number of "Preview" operations before the user's "Save" operation is extracted as a key feature, and abnormal paths (e.g., direct save) will be flagged.

[0141] In the second step, a hybrid architecture of isolation forest and LSTM is constructed to perform weighted voting based on the outputs of the two.

[0142] Assume that the isolation forest score (weight 0.6) + LSTM error (weight 0.4) is used to obtain an anomaly value. This anomaly value is compared with the preset threshold. Assuming the threshold is 0.8, a total score greater than 0.8 is considered an anomaly.

[0143] An embodiment of the present application is as follows: Isolation Forest discovers that a user clicks the "Undo" button 500 times in a single day (global anomaly); LSTM detects that the user has been in an "Undo→Redo" cycle for 10 consecutive minutes (local anomaly); after fusion, it is determined to be malicious script behavior.

[0144] Through step S203 , this embodiment successfully identifies data outliers, laying the foundation for subsequent data gap filling.

[0145] In one embodiment of the software function iteration method of the present application, see Figure 3 , and can also include the following:

[0146] Step S301: constructing a corresponding generator framework based on the residual fully connected network, wherein the input of the generator framework includes a missing data matrix, a mask matrix and a random noise vector, and the output of the generator framework includes a complete data matrix;

[0147] Step S302: constructing a corresponding discriminator framework according to the spectral normalization fully connected layer, wherein the input of the discriminator framework is the complete data matrix, and the output of the discriminator framework is a scalar probability.

[0148] Optionally, this step builds the generator architecture and the discriminator architecture.

[0149] Specifically, the generator network is constructed: the input of the generator is the missing data matrix, the mask matrix and the random noise vector, and the output of the generator is the complete data matrix.

[0150] The generator network architecture uses a fully connected network at the core layer for tabular data processing, employing residual connections (ResNet) to prevent vanishing gradients, and processes numerical and categorical features in two separate channels. Preferably, the core layer for time series data is modified to use a Transformer network, employing an attention mechanism to capture long-term dependencies and performing sinusoidal position encoding on timestamps.

[0151] The generator's loss function design specifically incorporates business rule constraints. For example, "preview must occur before saving" can be implemented through a rule-based loss function. Embedding business rules ensures that the generated data complies with logic. The generator is trained based on business rule constraints, reconstruction loss, and adversarial loss. The reconstruction loss only calculates the error for the non-missing parts, while the adversarial loss optimizes the generator through feedback from the discriminator.

[0152] Specifically, the discriminator network is constructed as follows: the input of the discriminator is the complete data matrix, and the input is the scalar probability. Among them, the scalar probability p∈[0,1] (1=real, 0=generated).

[0153] The discriminator network architecture adds spectral normalization to the core layer of tabular data processing to stabilize training, while using gradient penalty (WGAN-GP) to avoid mode collapse.

[0154] The loss function of the discriminator is designed to be a binary cross entropy loss.

[0155] After constructing the generator and discriminator, we alternately train the generator and discriminator to obtain a trained generative adversarial network for filling in missing data. It can be understood that the rule-based constraint loss ensures the rationality of the filling.

[0156] Through step S302, this embodiment successfully trains a generative adversarial network, which is based on a business logic constraint generator so that the generated data meets the business logic requirements, laying the foundation for subsequent data missing value filling.

[0157] In one embodiment of the software function iteration method of the present application, see Figure 4 , and can also include the following:

[0158] Step S401: constructing a user behavior consistency constraint according to a preset business rule, constructing a loss function according to the user behavior consistency constraint, a preset reconstruction loss, and a preset adversarial loss, and training the generator framework;

[0159] Step S402: The discriminator framework is trained according to a preset binary cross entropy loss, and the trained discriminator framework and the trained generator framework are jointly trained to determine a corresponding generative adversarial network.

[0160] Optional, this step is to generate adversarial network loss function design,

[0161] Specifically, the generator's loss function design incorporates business rule constraints. For example, "preview must occur before saving," which can be implemented through a rule-based loss function. Embedding business rules ensures that the generated data complies with logic. The generator is trained based on business rule constraints, reconstruction loss, and adversarial loss. The reconstruction loss only calculates the error for the non-missing data, while the adversarial loss optimizes the generator through feedback from the discriminator.

[0162] Specifically, the loss function of the discriminator is designed to be a binary cross entropy loss.

[0163] After constructing the generator and discriminator, we alternately train the generator and discriminator to obtain a trained generative adversarial network for filling in missing data. It can be understood that the rule-based constraint loss ensures the rationality of the filling.

[0164] Through step S402, this embodiment successfully trains a generative adversarial network, which is based on a business logic constraint generator so that the generated data meets the business logic requirements, laying the foundation for subsequent data missing value filling.

[0165] In one embodiment of the software function iteration method of the present application, see Figure 5 , and can also include the following:

[0166] Step S501: using the K-Means++ algorithm to determine the optimal number of clusters according to a preset silhouette coefficient, and determining the corresponding user group according to the optimal number of clusters;

[0167] Step S502: Based on the percentile of each of the user group behavior distributions, a corresponding differentiation threshold is determined, and the second user behavior feature is cleaned according to the differentiation threshold to determine a corresponding pre-processed user behavior feature.

[0168] This step is optional and involves cleaning and adapting the high-quality data obtained after data anomaly removal and missing value processing. The task of this stage is to set dynamic thresholds for users based on the clustering algorithm to develop group-specific anomaly analysis rules.

[0169] First, based on the K-Means++ clustering algorithm, we improve the initial center point selection to avoid the local optimality problem of traditional K-Means. We use the silhouette coefficient to determine the optimal number of clusters, and then first obtain all user groups based on the number of clusters. For example, we divide users into groups such as "novice", "normal", and "expert", and extract key indicators for each user type (such as the duration of a single operation).

[0170] Next, for each group, the 99% quantile of the key indicator is calculated (that is, 99% of the data is lower than this value), and the threshold is dynamically adjusted to avoid misjudgment caused by fixed values.

[0171] For example, for novice users, the P99 session duration is 45 minutes; for expert users, the P99 session duration is 15 minutes. User behavior data is reviewed in real time. If it exceeds the P99 threshold for the group, it is marked as an anomaly. The P99 threshold is recalculated weekly to adapt to changes in data distribution (such as user skill improvement).

[0172] It can be understood that this step of dividing users into groups based on the clustering algorithm and setting dynamic thresholds is to work in conjunction with the anomaly detection in step S101, and the normal threshold of the specific group output by this step is used to correct the results of the anomaly detection module and optimize the threshold rules.

[0173] Through step S502, this embodiment successfully sets an abnormal threshold based on the user group according to the clustering algorithm, laying the foundation for subsequent linkage with the abnormality detection module to form a data processing closed loop.

[0174] In one embodiment of the software function iteration method of the present application, see Figure 6 , and can also include the following:

[0175] Step S601: Select complementary models to construct a multi-model pool based on different characteristics of user behavior data, wherein the multi-model pool includes at least one of a decision tree model, a long short-term memory network, a graph neural network, and a Bayesian network;

[0176] Step S602: Dynamically assign the weights of each model in the multi-model pool based on a meta-learning algorithm, and output multimodal fusion features.

[0177] Optionally, in this embodiment, this step is a process of fusion features based on multi-model input.

[0178] Specifically, based on the different characteristics of user behavior data (such as temporal sequence, sparsity, and category imbalance), complementary models are selected to build a model pool for feature analysis and output multi-dimensional fusion features. The models in the model pool include but are not limited to:

[0179] Decision Tree (CART / Random Forest): used to process discrete features (such as function click classification);

[0180] LSTM (Long Short-Term Memory Network): used to capture the temporal dependencies of user operation sequences;

[0181] Graph Neural Networks (GNNs): Analyze the topological relationships between user-function interactions (e.g., “Users of function A often use function B together”);

[0182] Bayesian networks: Infer probabilistic associations between feature usage and user attributes (e.g., “the probability that the designer community prefers module X is 78%”).

[0183] The aforementioned models are candidate models. A meta-learning algorithm adjusts the output weights of each model based on real-time data distribution, screening the prediction results of one or more models with the highest weight ratios for fusion output. This fusion output approach can capture complex nonlinear relationships in user behavior (such as the dynamic correlation between feature usage frequency and user retention).

[0184] Through step S602 , this embodiment successfully obtains fusion features, which lays a foundation for subsequent mining of association rules in the fusion features and formulating corresponding software iteration strategies.

[0185] In one embodiment of the software function iteration method of the present application, see Figure 7 , and can also include the following:

[0186] Step S701: using a frequent itemset mining algorithm, filtering pseudo rules according to preset support, confidence, and lift parameters to determine corresponding strongly correlated association rules;

[0187] Step S702: performing contribution interpretation on the target variables in the strongly correlated association rules according to a preset SHAP analysis algorithm, and determining the corresponding software function usage association rule according to the target variable with the highest contribution.

[0188] Optionally, this embodiment is a process of identifying strong association rules based on a frequent mining algorithm.

[0189] Specifically, based on the frequent itemset mining algorithm, the minimum support and confidence are set and the lift is used to filter out pseudo rules (lift>1 indicates positive correlation). The rules of the data set with business labels are filtered to retain the rules that are strongly related to the business goals (such as associating "function usage" with "retention rate"), and redundant rules are eliminated (such as only one of "function A→function B" and "function B→function A" is retained).

[0190] Transform rules into executable recommendations based on business semantics, for example:

[0191] Rule: IF "Tolerance Analysis" usage time > 3 minutes AND number of errors ≥ 2 → THEN trigger real-time help (confidence level 85%);

[0192] Suggestion: Embed a “Smart Guide Button” in this function.

[0193] Next, to make the association rules output by the model easier to understand and quantify the impact of features on business indicators, we used SHAP value analysis to improve the interpretability of the model.

[0194] Specifically, the interpretability processing quantifies the contribution of each feature to the target variable (such as retention rate) by calculating the SHAP value.

[0195] For example, there are currently 100,000 user session logs, including usage sequence features of 20 functions such as "Tolerance Analysis" and "Sketch Tool"; business tag: user subscription renewal rate.

[0196] FP-Growth discovery rule: IF "Tolerance Analysis" is used ≥ 3 times / week AND single session duration > 5 minutes → THEN renewal probability increases by 25% (lift = 2.1)

[0197] SHAP analysis shows that "Tolerance Analysis" contributes most to renewals (SHAP = 0.22), so a report is generated: "Expert users make extensive use of the tolerance analysis function. We recommend adding an advanced tutorial entry."

[0198] In this way, we deeply explore the relationship between software functions and user needs, aiming to improve the accuracy and efficiency of user experience optimization decisions.

[0199] Through step S702, this embodiment successfully mines the relationship between software functions and user needs, improves the software functions based on user needs, implements software function iteration, and improves user experience and software usage efficiency.

[0200] In order to improve the efficiency and accuracy of software functions, the present application provides an embodiment of a software function iteration device for implementing all or part of the content of the software function iteration method, see Figure 8 , the software function iteration device specifically includes the following contents:

[0201] A user behavior data outlier processing module 10 is configured to obtain historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data, determine corresponding initial user behavior features, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior features based on the hybrid model, determine corresponding outliers, remove the outliers, and determine corresponding first user behavior features;

[0202] A user behavior data filling and cleaning module 20 is used to construct a generator and discriminator framework, perform conditional constraint training on the generator framework based on user behavior consistency constraints, determine a corresponding generative adversarial network, fill missing values ​​for the first user behavior feature based on the generative adversarial network, determine a corresponding second user behavior feature, clean the second user behavior feature based on a preset clustering algorithm, and determine a corresponding preprocessed user behavior feature;

[0203] The software function usage association rule determination module 30 is used to build a multi-model pool, input the pre-processed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

[0204] From the above description, it can be seen that the software function iteration device provided in the embodiment of the present application can construct a hybrid model through a preset isolation forest model and a preset time series model, perform anomaly detection on the initial user behavior characteristics, eliminate outliers, and then construct a generator and discriminator framework. The missing values ​​of the user behavior characteristics are filled according to the generative adversarial network, the missing values ​​are supplemented, and the processed user behavior characteristics are cleaned based on the preset clustering algorithm. A multi-model pool is constructed to perform multimodal fusion on the pre-processed user behavior characteristics, and the multimodal fusion features are combined with preset business labels to determine the corresponding user behavior data set. The user behavior data set is frequently mined based on the frequent item set mining algorithm to determine the corresponding software function usage association rules, and the software function is iterated according to the software function usage association rules, thereby improving the efficiency and accuracy of the software function.

[0205] From a hardware perspective, in order to improve the efficiency and accuracy of software functions, the present application provides an embodiment of an electronic device for implementing all or part of the software function iteration method. The electronic device specifically includes the following:

[0206] A processor, a memory, a communications interface, and a bus; wherein the processor, the memory, and the communications interface communicate with each other via the bus; the communications interface is used to implement information transmission between the software function iteration method and related devices such as the core business system, the user terminal, and related databases; the logic controller can be a desktop computer, a tablet computer, a mobile terminal, etc., but this embodiment is not limited thereto. In this embodiment, the logic controller can be implemented with reference to the embodiment of the software function iteration method in the embodiment, as well as the embodiment of the software function iteration method, the contents of which are incorporated herein, and repeated parts are not repeated.

[0207] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.

[0208] In practical applications, portions of the software function iteration method may be executed on the electronic device side as described above, or all operations may be performed on the client device. The specific selection may be based on the processing capabilities of the client device and the limitations of the user's usage scenario. This application does not impose any restrictions on this. If all operations are performed on the client device, the client device may also include a processor.

[0209] The client device may include a communication module (i.e., a communication unit) that can establish a communication connection with a remote server to implement data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a server structure of a distributed device.

[0210] Figure 9 Schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Figure 9 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that the Figure 9 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0211] In one embodiment, the software function iteration method function may be integrated into the central processing unit 9100. The central processing unit 9100 may be configured to perform the following control:

[0212] Step S101: Acquire historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data to determine a corresponding initial user behavior feature, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior feature based on the hybrid model, determine corresponding outliers, remove the outliers, and determine a corresponding first user behavior feature;

[0213] Step S102: Constructing a generator and discriminator framework, performing conditional constraint training on the generator framework based on user behavior consistency constraints, determining a corresponding generative adversarial network, filling missing values ​​for the first user behavior feature based on the generative adversarial network, determining a corresponding second user behavior feature, cleaning the second user behavior feature based on a preset clustering algorithm, and determining a corresponding preprocessed user behavior feature;

[0214] Step S103: Construct a multi-model pool, input the preprocessed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

[0215] From the above description, it can be seen that the electronic device provided in the embodiment of the present application constructs a hybrid model by using a preset isolation forest model and a preset time series model, performs anomaly detection on the initial user behavior characteristics, eliminates the outliers, and then constructs a generator and discriminator framework. The missing values ​​of the user behavior characteristics are filled according to the generative adversarial network, the missing values ​​are supplemented, and the processed user behavior characteristics are cleaned based on the preset clustering algorithm. A multi-model pool is constructed to perform multimodal fusion on the pre-processed user behavior characteristics, and the multimodal fusion features are combined with the preset business labels to determine the corresponding user behavior data set. The user behavior data set is frequently mined based on the frequent item set mining algorithm to determine the corresponding software function usage association rules, and the software function is iterated according to the software function usage association rules, thereby improving the efficiency and accuracy of the software function.

[0216] In another embodiment, the software function iteration method can be configured separately from the central processing unit 9100. For example, the software function iteration method can be configured as a chip connected to the central processing unit 9100, and the functions of the software function iteration method are realized through the control of the central processing unit.

[0217] like Figure 9 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Figure 9 In addition, the electronic device 9600 may also include all components shown in Figure 9 For components not shown, reference may be made to the prior art.

[0218] like Figure 9 As shown, the central processing unit 9100 is sometimes also referred to as a controller or operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.

[0219] Memory 9140 can be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It can store the aforementioned failure-related information and also store programs that execute the relevant information. The CPU 9100 can execute the programs stored in memory 9140 to implement information storage or processing.

[0220] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 may be, for example, a keypad or touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display objects such as images and text. The display may be, for example, an LCD display, but is not limited thereto.

[0221] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), or a SIM card. Alternatively, it may be a memory that retains information even when power is off, can be selectively erased, and is provided with more data. Examples of such memory are sometimes referred to as EPROMs. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 by the central processing unit 9100.

[0222] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various driver programs for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0223] The communication module 9110 is a transmitter / receiver that transmits and receives signals via the antenna 9111. The communication module 9110 is coupled to the central processing unit 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.

[0224] Based on different communication technologies, multiple communication modules 9110 can be provided in the same electronic device, such as a cellular network module, a Bluetooth module, and / or a wireless local area network module. The communication module 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby implementing common telecommunication functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Furthermore, the audio processor 9130 is also coupled to the central processing unit 9100, enabling local recording via the microphone 9132 and playback of stored audio via the speaker 9131.

[0225] The embodiments of the present application also provide a computer-readable storage medium capable of implementing all steps of the software function iteration method in the above-mentioned embodiment, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the computer program implements all steps of the software function iteration method in the above-mentioned embodiment, where the execution subject is a server or a client. For example, when the processor executes the computer program, the following steps are implemented:

[0226] Step S101: Acquire historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data to determine a corresponding initial user behavior feature, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior feature based on the hybrid model, determine corresponding outliers, remove the outliers, and determine a corresponding first user behavior feature;

[0227] Step S102: Constructing a generator and discriminator framework, performing conditional constraint training on the generator framework based on user behavior consistency constraints, determining a corresponding generative adversarial network, filling missing values ​​for the first user behavior feature based on the generative adversarial network, determining a corresponding second user behavior feature, cleaning the second user behavior feature based on a preset clustering algorithm, and determining a corresponding preprocessed user behavior feature;

[0228] Step S103: Construct a multi-model pool, input the preprocessed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

[0229] From the above description, it can be seen that the computer-readable storage medium provided in the embodiment of the present application constructs a hybrid model by presetting an isolation forest model and a preset time series model, performs anomaly detection on the initial user behavior characteristics, eliminates the outliers, and then constructs a generator and discriminator framework, fills the missing values ​​of the user behavior characteristics according to the generative adversarial network, makes up the missing values, cleans the processed user behavior characteristics based on the preset clustering algorithm, constructs a multi-model pool to perform multimodal fusion on the preprocessed user behavior characteristics, combines the multimodal fusion features with the preset business labels to determine the corresponding user behavior data set, performs frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determines the corresponding software function usage association rules, and iterates the software function according to the software function usage association rules, thereby improving the efficiency and accuracy of the software function.

[0230] The embodiments of the present application also provide a computer program product capable of implementing all steps of the software function iteration method in the above-mentioned embodiment, where the execution subject is a server or a client. When the computer program / instruction is executed by a processor, the computer program / instruction implements the steps of the software function iteration method. For example, the computer program / instruction implements the following steps:

[0231] Step S101: Acquire historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data to determine a corresponding initial user behavior feature, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior feature based on the hybrid model, determine corresponding outliers, remove the outliers, and determine a corresponding first user behavior feature;

[0232] Step S102: Constructing a generator and discriminator framework, performing conditional constraint training on the generator framework based on user behavior consistency constraints, determining a corresponding generative adversarial network, filling missing values ​​for the first user behavior feature based on the generative adversarial network, determining a corresponding second user behavior feature, cleaning the second user behavior feature based on a preset clustering algorithm, and determining a corresponding preprocessed user behavior feature;

[0233] Step S103: Construct a multi-model pool, input the preprocessed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

[0234] From the above description, it can be seen that the computer program product provided in the embodiment of the present application constructs a hybrid model by presetting an isolation forest model and a preset time series model, performs anomaly detection on the initial user behavior characteristics, eliminates outliers, and then constructs a generator and discriminator framework. The missing values ​​of the user behavior characteristics are filled according to the generative adversarial network, the missing values ​​are supplemented, and the processed user behavior characteristics are cleaned based on the preset clustering algorithm. A multi-model pool is constructed to perform multimodal fusion on the preprocessed user behavior characteristics, and the multimodal fusion features are combined with preset business labels to determine the corresponding user behavior data set. The user behavior data set is frequently mined based on the frequent item set mining algorithm to determine the corresponding software function usage association rules, and the software function is iterated according to the software function usage association rules, thereby improving the efficiency and accuracy of the software function.

[0235] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0236] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (apparatus), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0237] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0238] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0239] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. A software function iteration method, characterized in that: The method comprises: Acquire historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data to determine a corresponding initial user behavior feature, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior feature based on the hybrid model, determine a corresponding outlier, remove the outlier, and determine a corresponding first user behavior feature; Constructing a generator and discriminator framework, performing conditional constraint training on the generator framework based on user behavior consistency constraints, determining a corresponding generative adversarial network, filling missing values ​​for the first user behavior feature using the generative adversarial network, determining a corresponding second user behavior feature, cleaning the second user behavior feature based on a preset clustering algorithm, and determining a corresponding preprocessed user behavior feature; Construct a multi-model pool, input the preprocessed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

2. The software function iteration method according to claim 1, characterized in that: The step of constructing a hybrid model based on a preset isolation forest model and a preset time series model, performing an anomaly detection operation on the initial user behavior feature based on the hybrid model, determining corresponding outliers and removing the outliers, and determining the corresponding first user behavior feature includes: Performing a global anomaly location operation on the initial user behavior features according to a preset isolation forest model to determine the corresponding global sparse anomaly; Performing a local anomaly location operation on the initial user behavior features according to a preset time series model to determine the corresponding local time series anomaly, wherein the preset time series model adopts an attention mechanism to locate key abnormal time points in the user behavior sequence; The global sparse anomaly and the local time series anomaly are weightedly fused, the scores obtained after the weighted fusion are screened according to a preset anomaly threshold, the corresponding outliers are determined and eliminated, and the corresponding first user behavior feature is determined.

3. The software function iteration method according to claim 1, characterized in that: The generator and discriminator frameworks are constructed, including: Constructing a corresponding generator framework according to the residual fully connected network, wherein the input of the generator framework includes a missing data matrix, a mask matrix and a random noise vector, and the output of the generator framework includes a complete data matrix; A corresponding discriminator framework is constructed according to the spectral normalization fully connected layer, where the input of the discriminator framework is the complete data matrix, and the output of the discriminator framework is a scalar probability.

4. The software function iteration method according to claim 1, characterized in that: The step of performing conditional constraint training on the generator framework based on the user behavior consistency constraint to determine the corresponding generative adversarial network includes: Constructing a user behavior consistency constraint according to preset business rules, constructing a loss function according to the user behavior consistency constraint, a preset reconstruction loss, and a preset adversarial loss, and training the generator framework; The discriminator framework is trained according to a preset binary cross entropy loss, and the trained discriminator framework and the trained generator framework are jointly trained to determine a corresponding generative adversarial network.

5. The software function iteration method according to claim 1, characterized in that: The cleaning of the second user behavior feature based on a preset clustering algorithm to determine a corresponding preprocessed user behavior feature includes: Using the K-Means++ algorithm to determine the optimal number of clusters according to a preset silhouette coefficient, and determining the corresponding user group according to the optimal number of clusters; Based on the percentile of the behavior distribution of each user group, a corresponding differentiation threshold is determined, and the second user behavior feature is cleaned according to the differentiation threshold to determine the corresponding pre-processed user behavior feature.

6. The software function iteration method according to claim 1, characterized in that: The step of constructing a multi-model pool, inputting the pre-processed user behavior features into the multi-model pool, performing weight assignment on each model in the multi-model pool according to a meta-learning algorithm, and determining a multimodal fusion feature output by the multi-model pool includes: According to the different characteristics of user behavior data, complementary models are selected to construct a multi-model pool, wherein the multi-model pool includes at least one of a decision tree model, a long short-term memory network, a graph neural network, and a Bayesian network; Based on the meta-learning algorithm, the weights of each model in the multi-model pool are dynamically allocated to output multimodal fusion features.

7. The software function iteration method according to claim 1, characterized in that: Performing frequent item mining on the user behavior dataset based on a frequent item set mining algorithm to determine corresponding software function usage association rules, including: Adopting the frequent item set mining algorithm, the pseudo rules are filtered according to the preset support, confidence and lift parameters to determine the corresponding strong correlation association rules; The contribution of the target variables in the strongly correlated association rules is interpreted according to a preset SHAP analysis algorithm, and the corresponding software function usage association rule is determined according to the target variable with the highest contribution.

8. A software function iteration device, characterized in that: The device comprises: A user behavior data outlier processing module is configured to obtain historical user behavior data, perform a multi-dimensional feature extraction operation on the historical user behavior data, determine a corresponding initial user behavior feature, construct a hybrid model based on a preset isolation forest model and a preset time series model, perform an anomaly detection operation on the initial user behavior feature based on the hybrid model, determine a corresponding outlier, remove the outlier, and determine a corresponding first user behavior feature; A user behavior data filling and cleaning module is used to build a generator and discriminator framework, perform conditional constraint training on the generator framework based on user behavior consistency constraints, determine a corresponding generative adversarial network, fill missing values ​​for the first user behavior feature based on the generative adversarial network, determine a corresponding second user behavior feature, clean the second user behavior feature based on a preset clustering algorithm, and determine a corresponding preprocessed user behavior feature; The software function usage association rule determination module is used to build a multi-model pool, input the preprocessed user behavior features into the multi-model pool, assign weights to the various models in the multi-model pool according to the meta-learning algorithm, determine the multimodal fusion features output by the multi-model pool, combine the multimodal fusion features with preset business tags to determine the corresponding user behavior data set, perform frequent item mining on the user behavior data set based on the frequent item set mining algorithm, determine the corresponding software function usage association rules, and perform software function iteration according to the software function usage association rules.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the software function iteration method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the software function iteration method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Data mining-based SaaS software product demand generation method and system

    CN121092124A