AB test effect evaluation method based on dynamic confidence interval

By using a dynamic confidence interval A/B testing method, combined with rolling window aggregation, Bayes' theorem, and autoregressive analysis, the problems of misjudgment and decision delay caused by static confidence interval settings in A/B testing are solved, achieving a more accurate and faster evaluation of experimental results.

CN121070804AActive Publication Date: 2025-12-05MOJI FENGYUN BEIJING SOFTWARE TECH DEV CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511603892.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2025-12-05
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

In existing A/B testing methods, the confidence interval is set statically and fails to be adaptively adjusted according to the dynamic changes of the time series. This may lead to misjudgments in the early or late stages of the test, and it cannot accurately reflect the real-time changes in the experimental results. Furthermore, it does not effectively integrate historical experimental data for Bayesian prior inference, which means that new experiments need to wait for a sufficiently long testing period to obtain reliable results, increasing the delay in business decision-making.

Method used

By collecting exposure event data streams and conversion event data streams from A/B testing experiments, rolling window aggregation and time series analysis are performed to generate dynamic confidence intervals. Combined with Bayes' theorem and autoregressive analysis, the posterior distribution is dynamically updated to construct a dual-optimized dynamic confidence interval, and decisions and parameter recommendations are made based on this.

Benefits of technology

It achieves prior guidance for effect estimation, improves the stability and reliability of small-sample stage assessment, enhances the accuracy and robustness of confidence intervals, and reduces decision delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121070804A_ABST
    Figure CN121070804A_ABST
Patent Text Reader

Abstract

The invention discloses an AB test effect evaluation method based on a dynamic confidence interval, and relates to the technical field of big data analysis, and the method comprises the steps: carrying out the autoregression analysis of a historical time sequence, extracting an autoregression coefficient, obtaining a dynamic variance dilatation factor according to a time sequence analysis theory, and obtaining a dynamic variance dilatation coefficient; correcting the posterior distribution variance by using the dynamic variance expansion factor to generate a corrected posterior distribution variance; constructing a dual optimization dynamic confidence interval through the mean value of the posterior distribution and the corrected posterior distribution variance, comparing the dual optimization dynamic confidence interval with a preset sequential test boundary, making a decision, and generating an evaluation decision; and based on the evaluation decision, using the mean value of the posterior distribution and the corrected posterior distribution variance as recommendation parameters, setting initialization parameter recommendation of the AB test in the next stage, and integrating to generate an AB test effect evaluation report. According to the invention, the accuracy and robustness of the confidence interval are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data analytics, and in particular to a method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals. Background Technology

[0002] A / B testing, as a core experimental method for iterating and optimizing internet products, has evolved from simple controlled experiments to a complex evaluation system based on big data analysis and statistical inference. With the expansion of internet business scale and the explosive growth of user behavior data, A / B testing plays an increasingly important role in product decision-making. Traditional A / B testing methods mainly rely on hypothesis testing and fixed confidence intervals, evaluating effectiveness by comparing the conversion rate differences between the experimental and control groups. In recent years, with the introduction of Bayesian statistics and machine learning techniques, A / B testing methods have improved in efficiency and accuracy; methods such as Bayesian A / B testing and sequential analysis are beginning to be applied in real-world scenarios.

[0003] Existing A / B testing techniques have two main shortcomings: First, the confidence intervals are set statically and fail to be adaptively adjusted according to the dynamic changes in the time series, which may lead to misjudgments in the early or late stages of testing and fail to accurately reflect the real-time changes in experimental results. Second, the lack of effective integration of historical experimental data for Bayesian prior inference means that new experiments need to wait for a sufficiently long testing period to obtain reliable results, which reduces experimental efficiency and increases the delay in business decision-making. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals to address the problems of high risk of misjudgment and decision delay.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] This invention provides a method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals, comprising:

[0008] The system collects exposure event data streams and conversion event data streams from A / B tests, aggregates them through a rolling window using a time window, obtains the cumulative exposure counts and conversion counts of the A / B tests within the current time window, and stores them in chronological order to generate the current window's A / B test data and historical time series.

[0009] Based on the current window AB test data, a set of similar final effect size estimates is retrieved. Through statistical inference methods, a prior distribution of the test effect is generated and updated according to Bayes' theorem. The prior distribution and the current window AB test data are then merged to generate the mean and variance of the posterior distribution.

[0010] Autoregressive analysis is performed on historical time series to extract autoregressive coefficients. Based on time series analysis theory, dynamic variance inflation factor is obtained. The dynamic variance inflation factor is used to correct the posterior distribution variance to generate the corrected posterior distribution variance.

[0011] A dual-optimized dynamic confidence interval is constructed using the mean of the posterior distribution and the corrected variance of the posterior distribution. The dual-optimized dynamic confidence interval is compared with the preset sequential test boundary and a decision is made to generate an evaluation decision.

[0012] Based on the evaluation decision, the mean of the posterior distribution and the corrected variance of the posterior distribution are used as recommended parameters to set the initialization parameters for the next stage of the A / B test, and an A / B test effect evaluation report is generated.

[0013] As a preferred embodiment of the method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in this invention, the acquisition of exposure event data stream and conversion event data stream of the A / B testing experiment refers to acquiring the original event stream of the A / B testing experiment, classifying the event types, filtering and recording events with abnormal timestamps and missing key fields, and generating exposure event data stream and conversion event data stream.

[0014] As a preferred embodiment of the method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals described in this invention, the steps of performing rolling window aggregation through a time window to obtain the cumulative exposure and conversion counts of the A / B test within the current time window, and storing them in chronological order to generate the current window A / B test data and historical time series are as follows.

[0015] The continuous exposure event data stream and conversion event data stream are divided into discrete time intervals of fixed length by time windows, and grouped according to the test version number to generate a windowed event set;

[0016] Perform multi-dimensional aggregation operations on the exposure event data stream and conversion event data stream in the windowed event set to obtain the cumulative exposure count and conversion count, and perform correlation matching to generate the current time window metric statistical record;

[0017] All time window indicator statistics are combined into a continuous historical time series, and then linked and integrated with the current time window indicator statistics to generate the current window A / B test data.

[0018] As a preferred embodiment of the method for evaluating the effect of A / B testing based on dynamic confidence intervals according to the present invention, the steps of retrieving a set of similar final effect size estimates based on the current window of A / B testing data and generating a prior distribution of the experimental effect through statistical inference methods are as follows:

[0019] Based on the current window AB test data, AB test records that meet the similarity threshold are selected from the historical test knowledge base, and the final effect size estimate is extracted to generate a set of historical test effect size estimates.

[0020] The distribution of the historical experimental effect size estimates was fitted using the Bayesian hierarchical inference method to obtain the prior distribution mean and variance, and then integrated into a prior distribution.

[0021] As a preferred embodiment of the method for evaluating the effectiveness of AB testing based on dynamic confidence intervals described in this invention, the steps of updating according to Bayes' theorem, fusing the prior distribution and the current window of AB testing data, and generating the mean and variance of the posterior distribution are as follows:

[0022] Based on the cumulative exposure and conversion counts in the current time window's statistical records, calculate the point estimate and naive variance estimate of the effect size for the current window, and generate a likelihood function representation.

[0023] Based on Bayes' theorem, the prior distribution and the likelihood function are probabilistically fused, and the specific parameter representation of the posterior distribution is obtained through closed-form analytical derivation. The mean and variance of the posterior distribution are then obtained.

[0024] As a preferred embodiment of the method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals as described in this invention, the steps of performing autoregressive analysis on historical time series, extracting autoregressive coefficients, and obtaining the dynamic variance inflation factor according to time series analysis theory are as follows.

[0025] Linear interpolation is used to fill in missing time points in historical time series data, outlier data points are corrected using the three-standard-deviation principle, and stationarity is tested to generate standard historical time series data.

[0026] We use an autoregressive analysis framework to model and analyze standard historical time series data, identify the optimal order of autoregression, and obtain the estimated values ​​of autoregressive coefficients through the least squares estimation method to generate autoregressive analysis results.

[0027] Based on the autoregressive analysis results, the variance inflation factor value for the current time window is obtained through the standard formula of time series theory, and its rationality is verified to generate a reasonable dynamic variance inflation factor value.

[0028] As a preferred embodiment of the method for evaluating the experimental effect of AB test based on dynamic confidence intervals as described in this invention, the generation of the corrected posterior distribution variance refers to combining the reasonable dynamic variance inflation factor value with the posterior distribution variance to correct the posterior distribution variance, and generating the corrected posterior distribution variance through nonnegativity verification and standardized formatting.

[0029] As a preferred embodiment of the method for evaluating the experimental effect of AB test based on dynamic confidence intervals as described in this invention, the construction of the dual-optimized dynamic confidence interval refers to obtaining the upper and lower bounds of the dual-optimized dynamic confidence interval based on the mean of the posterior distribution and the corrected variance of the posterior distribution through the standard normal distribution quantile method, and constructing the dual-optimized dynamic confidence interval.

[0030] As a preferred embodiment of the method for evaluating the effectiveness of AB testing based on dynamic confidence intervals described in this invention, the steps for generating the evaluation decision are as follows:

[0031] The dual-optimized dynamic confidence interval object is compared with the preset sequential test boundary value to dynamically generate preliminary decision conclusions;

[0032] Based on the preliminary decision conclusions, compliance verification and risk factor assessment are conducted to generate an assessment decision.

[0033] As a preferred embodiment of the method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals described in this invention, the steps of setting initialization parameters for the next stage of the A / B testing based on evaluation decisions, using the mean of the posterior distribution and the corrected variance of the posterior distribution as recommended parameters, and integrating these parameters to generate an A / B testing effectiveness evaluation report, are as follows:

[0034] Based on the evaluation decision, dynamic judgment is made, and the posterior distribution mean parameter and the corrected posterior distribution variance parameter are extracted as core recommendation parameters to generate recommendation parameters;

[0035] The recommended parameter set is mapped and converted into the prior distribution center value and minimum detectable effect value of the next stage AB test, and the scaling factor of the posterior distribution variance is adjusted to optimize the sensitivity, generating the initialization parameter recommendations;

[0036] By integrating the initialization parameter recommendations, evaluation decisions, dual-optimization dynamic confidence interval data, and recommended parameter sets, an A / B test effect evaluation report is generated.

[0037] The beneficial effects of this invention are as follows: By constructing a Bayesian prior distribution through the fusion of historical experimental data and dynamically updating it with real-time data, prior guidance for effect estimation is achieved, improving the stability and reliability of small-sample stage assessments. By extracting the dynamic variance inflation factor through autoregressive analysis and performing time-series correction on the posterior variance, the underestimation of variance caused by data autocorrelation is effectively compensated, enhancing the accuracy and robustness of confidence intervals. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a flowchart of a method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals.

[0040] Figure 2 This is a flowchart of the AB test data acquisition and preprocessing process.

[0041] Figure 3 A flowchart for generating Bayesian prior and posterior distributions.

[0042] Figure 4 A flowchart for dynamic correction and evaluation decisions. Detailed Implementation

[0043] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0044] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0045] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0046] Reference Figures 1-4 This is one embodiment of the present invention, which provides a method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals, including the following steps:

[0047] S1: Collect exposure event data stream and conversion event data stream of AB test, perform rolling window aggregation through time window, obtain the cumulative exposure number and conversion number of AB test within the current time window, and store them in time order to generate AB test data of the current window and historical time series;

[0048] S1.1: Collect the raw event stream of the A / B test, classify the event types, filter and record events with abnormal timestamps and missing key fields, and generate exposure event data stream and conversion event data stream;

[0049] Furthermore, the raw event stream of the A / B test is collected, and the events are classified in real time according to the event type identifier of the raw event stream. All events with the event type identifier of exposure type are assigned to the exposure event set, and all events with the event type identifier of conversion type are assigned to the conversion event set. Anomaly handling is performed on the classified events. The value of the event timestamp field is checked to see if it is within a reasonable time range (e.g., earlier than January 1, 2020 or later than the current server time). Events with timestamp values ​​outside the reasonable range are marked as abnormal events and filtered out. At the same time, it is checked whether the event is missing a key field (e.g., user identifier field or test version number field). Events with missing key fields are marked as abnormal events and filtered out. All filtered abnormal events are recorded in the abnormal event log, and the record content includes the original event content, the reason for filtering, and the processing timestamp. The exposure event data stream and the conversion event data stream are output.

[0050] It should be noted that the event type identifier is a metadata field in the original event stream of the A / B test, including two core types: exposure type identifier and conversion type identifier.

[0051] S1.2: Divide the continuous exposure event data stream and conversion event data stream into discrete time intervals of fixed length through time windows, and group them according to the test version number to generate a windowed event set;

[0052] Furthermore, events are assigned to fixed-length time intervals based on the timestamp fields of the exposure event data stream and the conversion event data stream, with each time interval representing a discrete time period. Events within each time interval are grouped, and events with the same test version number are grouped into the same group. Finally, a windowed event set is generated for each time interval and each test version number.

[0053] S1.3: Perform multi-dimensional aggregation operations on the exposure event data stream and conversion event data stream in the windowed event set to obtain the cumulative exposure count and conversion count, and perform correlation matching to generate the current time window indicator statistical record;

[0054] Furthermore, a counting aggregation operation is performed on the exposure event data stream of the windowed event set: based on the user identifier field, a global count of exposure events under each test version number (e.g., "A" or "B") is performed to generate the cumulative exposure count for each test version within the current time window; a deduplication counting aggregation operation is performed on the conversion event data stream: based on the user identifier field, a distributed deduplication count is performed on the conversion events under each test version number to exclude duplicate conversion events of the same user within the same time window, generating the cumulative conversion count for each test version within the current time window; the cumulative exposure count and cumulative conversion count of the same test version number are associated and matched, and combined into a current time window metric statistical record containing the time window start time, time window end time, test version number, cumulative exposure count, and cumulative conversion count.

[0055] S1.4: Combine all time window indicator statistics into a continuous historical time series, and integrate it with the current time window indicator statistics to generate the current window AB test data.

[0056] Furthermore, all time window indicator statistics records are retrieved from persistent storage in chronological order and sorted in ascending order based on the time window start time field to form a continuous historical time series array. The current time window indicator statistics record is appended to the end of the historical time series array to ensure the continuity of the time order. The integrated complete historical time series array is associated with the current time window indicator statistics record, and a mapping relationship is established through the test version number field and the time window identifier field to generate the current window AB test test data containing historical context and real-time data.

[0057] S2: Based on the current window AB test data, retrieve a set of similar final effect size estimates, generate the prior distribution of the test effect through statistical inference methods, update it according to Bayes' theorem, and merge the prior distribution and the current window AB test data to generate the mean and variance of the posterior distribution.

[0058] S2.1: Based on the current window AB test data, select AB test records that meet the similarity threshold from the historical test knowledge base, extract the final effect size estimate, and generate a set of historical test effect size estimates;

[0059] Specifically, based on the type classification metadata in the current window's A / B test data, all A / B test records in the historical test knowledge base are queried. A metadata similarity matching algorithm is used to obtain the similarity score between the current type classification metadata and each historical test record. A similarity threshold is set for filtering, and only historical test records with similarity scores greater than or equal to the similarity threshold are retained. The final effect size estimate field is extracted from the filtered historical test records, and all extracted final effect size estimates are summarized into a data structure to generate a set of historical test effect size estimates.

[0060] It should be noted that type classification metadata refers to the business dimension attribute data of A / B testing experiments, including feature dimensions used to measure the similarity of experiments, such as test function node identifiers and page type classifications.

[0061] The historical test knowledge base is a database that centrally stores and manages detailed records of all completed A / B tests. It is defined based on long-term accumulated test metadata and test result data. By storing and indexing this information in a structured manner, it enables the rapid retrieval of similar historical test records based on the characteristics of the current test.

[0062] The similarity threshold (range: 0.7-0.9) is a numerical threshold set based on a combination of matching accuracy and business risk tolerance. It is used to determine whether the similarity between historical and current experiments is sufficient to provide a reliable reference.

[0063] Meta-information similarity matching algorithm is a mathematical method that quantifies the degree of similarity between current and historical trials by using the cosine value of the spatial angle between their meta-information feature vectors.

[0064] S2.2: The distribution of the historical experimental effect size estimates is fitted using the Bayesian hierarchical inference method to obtain the prior distribution mean and prior distribution variance, and then integrated into a prior distribution.

[0065] Specifically, a Bayesian hierarchical inference method is applied to the set of historical effect size estimates. It is assumed that the historical effect size estimates follow a normal distribution and an uninformative prior distribution is set for the mean and variance parameters of the normal distribution. A Markov chain Monte Carlo sampling method is used to draw samples from the joint posterior distribution. The posterior distribution of hyperparameters is estimated through the sample statistics, and the prior distribution mean and variance parameters describing the distribution characteristics of historical effect sizes are obtained. Finally, the prior distribution mean and variance parameters are integrated into a complete prior distribution.

[0066] S2.3: Based on the cumulative exposure and conversion counts in the current time window's statistical records, calculate the point estimate of the effect size and the naive variance estimate for the current window, and generate the likelihood function representation.

[0067] The formula for calculating the point estimate of the effect size is:

[0068] ;

[0069] in, This represents the point estimate of the effect size. express The number of transformations of the group express Group exposure count, express The number of transformations of the group express Group exposure count, express Group conversion rate estimate express Group conversion rate estimate.

[0070] The formula for calculating the naive variance estimate is:

[0071] ;

[0072] in, This represents the naive variance estimate. express Group conversion rate estimate express Group conversion rate estimate express The variance of the conversion rate of the group express The variance of the conversion rate of the group.

[0073] It should be noted that, based on the cumulative exposure and conversion counts in the current time window's statistical records, the conversion rate estimates for Group A and Group B are calculated separately. The difference between the conversion rate of Group B and the conversion rate of Group A yields the effect size point estimate for the current window. Then, based on the binomial distribution variance formula, the variances of the conversion rates of Group A and Group B are calculated separately. The sum of the two variances yields the naive variance estimate. Finally, based on the effect size point estimate and the naive variance estimate, a likelihood function representation in the form of a normal distribution is constructed.

[0074] S2.4: Based on Bayes' theorem, the prior distribution and the likelihood function are probabilistically fused, and the specific parameter representation of the posterior distribution is obtained through closed-form analytical derivation. The mean and variance of the posterior distribution are then obtained.

[0075] Specifically, based on Bayes' theorem, the prior distribution and the likelihood function are probabilistically fused. The mean and variance parameters of the prior distribution, along with the effect size point estimates and naive variance estimates in the likelihood function, are substituted into the conjugate prior update formula. The mean parameter of the posterior distribution is obtained through weighted averaging, and the variance parameter of the posterior distribution is obtained through harmonic averaging. The complete parameter representation of the posterior distribution is analytically derived, and the mean and variance parameters of the posterior distribution are obtained.

[0076] It should be noted that the conjugate prior update formula is a mathematical tool in Bayesian statistics. When the prior distribution and the likelihood function belong to the family of conjugate distributions, it allows the parameters of the prior distribution to be analytically combined with the sample observation statistics from the likelihood function.

[0077] S3: Perform autoregressive analysis on the historical time series, extract the autoregressive coefficients, obtain the dynamic variance inflation factor according to the time series analysis theory, use the dynamic variance inflation factor to correct the posterior distribution variance, and generate the corrected posterior distribution variance.

[0078] S3.1: Use linear interpolation to fill in missing time point data in historical time series data, use the three-standard-deviation principle to correct outlier data points, and perform stationarity tests to generate standard historical time series data;

[0079] Specifically, preprocessing is performed on the historical time series data. Linear interpolation is used to identify and locate the positions of all missing time points to obtain missing values. The overall mean and standard deviation of the historical time series data are obtained, and outlier data points are identified using the three-standard-deviation principle. The moving average of the preceding and following normal data is used to replace and correct the outlier data points. The Augmented Dickey-Fuller test is performed on the processed complete data series to determine stationarity. If the test fails, first-order differencing is performed until the value is below the stationarity threshold (range: 0.01-0.1, defined according to the conventional setting of significance level α in statistics and the control requirements for the probability of Type I error in hypothesis testing). Finally, standardized historical time series data that meets the analysis requirements are generated.

[0080] It should be noted that the Augmented Dickey-Fuller test is a statistical hypothesis testing method used to determine whether time series data is stationary. The core principle is to determine non-stationarity by testing whether there is a unit root in the time series: if the test result shows that there is a unit root, it indicates that the series is non-stationary and needs to be processed by methods such as differencing until the null hypothesis of the existence of the unit root is rejected, thereby ensuring that the series meets the stationarity requirements.

[0081] S3.2: The standard historical time series data is modeled and analyzed using an autoregressive analysis framework to identify the optimal order of autoregression and obtain the estimated values ​​of the autoregressive coefficients through the least squares estimation method, thereby generating the autoregressive analysis results.

[0082] Specifically, an autoregressive analysis framework is used to model and analyze standard historical time series data. Based on the Akaike information criterion, AIC values ​​are obtained at different orders within a preset order range (range: 1-10, set according to the data sampling frequency and empirical rules in the business scenario). The order that minimizes the AIC value is selected as the optimal order of autoregression. A matrix containing lag terms is constructed, and the current observation value is used as the dependent variable. The least squares estimation method is applied to solve the normal equation system to obtain the estimated values ​​of the autoregressive coefficients. Finally, the autoregressive analysis results, including the optimal order of autoregression, the estimated values ​​of the autoregressive coefficients, and the model fit index, are generated.

[0083] It should be noted that the autoregressive analysis framework is a statistical method for analyzing time series data. Its core idea is to use the historical observations of the time series itself as explanatory variables to construct predictive relationships in order to estimate current or future series values. This includes defining the autoregressive structure, determining the optimal lag order, using the least squares method or maximum likelihood estimation method to obtain the estimated values ​​of the autoregressive coefficients, and performing residual analysis verification steps, aiming to capture the autocorrelation and dynamic patterns in the time series.

[0084] The Akaike information criterion is a criterion for evaluating the quality of statistical methods by using goodness of fit and parameter number penalty. In autoregressive analysis, it is used to select the optimal lag order to balance complexity and the risk of overfitting.

[0085] S3.3: Based on the autoregressive analysis results, obtain the variance inflation factor value for the current time window using the standard formula of time series theory, verify its rationality, and generate a reasonable dynamic variance inflation factor value.

[0086] Specifically, based on the estimated values ​​of the autoregressive coefficients in the autoregressive analysis results, the standard formula in time series theory is applied to obtain the variance inflation factor value for the current time window; the rationality of the variance inflation factor value is verified to ensure that the value meets the theoretical requirements (check whether the variance inflation factor value is ≥1); and a verified and reasonable dynamic variance inflation factor value is generated.

[0087] S3.4: Combine the reasonable dynamic variance inflation factor value with the posterior distribution variance to perform posterior distribution variance correction, and generate the corrected posterior distribution variance through nonnegativity verification and standardized formatting.

[0088] Specifically, the reasonable dynamic variance inflation factor value and the posterior distribution variance parameter are standardized and corrected based on time series theory to obtain the product of the reasonable dynamic variance inflation factor value and the posterior distribution variance parameter. Then, a non-negativity verification process is performed to check whether the non-negativity requirement of the variance parameter is met. The verified value is converted into a standardized floating-point number format, formatted and decimal places are retained to generate the corrected posterior distribution variance parameter.

[0089] It should be noted that the nonnegativity requirement refers to the mathematical constraint that the variance parameter, as an indicator of the degree of data dispersion, must be greater than or equal to zero. This is a fundamental property determined by the mathematical definition of variance itself and the nonnegativity of probability distribution.

[0090] S4: Construct a dual-optimized dynamic confidence interval using the mean of the posterior distribution and the corrected variance of the posterior distribution. Compare the dual-optimized dynamic confidence interval with the preset sequential test boundary and make a decision to generate an evaluation decision.

[0091] S4.1: Based on the mean of the posterior distribution and the corrected variance of the posterior distribution, the upper and lower bounds of the double-optimized dynamic confidence interval are obtained through the standard normal distribution quantile method, and the double-optimized dynamic confidence interval is constructed.

[0092] Furthermore, a dual-optimized dynamic confidence interval is constructed based on the mean parameter and the corrected variance parameter of the posterior distribution. The corresponding standard normal distribution quantile parameter is determined by querying the standard normal distribution table, and the standard deviation of the estimated value is obtained by taking the square root of the corrected variance parameter of the posterior distribution. The upper bound is obtained by summing the product of the mean parameter of the posterior distribution with the standard normal distribution quantile parameter and the standard deviation of the estimated value, and the lower bound is obtained by performing a difference operation. The lower bound and upper bound of the dual-optimized dynamic confidence interval are combined into a complete interval object, thus completing the construction of the dual-optimized dynamic confidence interval.

[0093] It should be noted that the standard normal distribution table lists the cumulative distribution function values ​​of the standard normal distribution. It is set based on the exact probability value obtained by integrating the normal distribution probability density function in probability theory, so that the corresponding precise quantile parameters can be quickly looked up according to any given confidence level.

[0094] S4.2: Compare the dual-optimized dynamic confidence interval object with the preset sequential test boundary value to dynamically generate preliminary decision conclusions;

[0095] Furthermore, the lower and upper bound values ​​stored in the dual-optimized dynamic confidence interval object are read, and the upper and lower bound values ​​in the preset sequential test boundary values ​​are obtained. A triple condition judgment is performed: if the entire interval range of the dual-optimized dynamic confidence interval object is above the upper bound value of the preset sequential test boundary value, a preliminary decision conclusion of "launch version B" is generated; if the entire interval range is below the lower bound value of the preset sequential test boundary value, a preliminary decision conclusion of "rollback" is generated; if the interval range intersects with the preset sequential test boundary value, a preliminary decision conclusion of "continue experiment" is generated; finally, a preliminary decision conclusion object containing the decision type and timestamp is output.

[0096] It should be noted that the preset sequential test boundary values ​​are a set of statistical decision values ​​that are dynamically adjusted according to the sequential analysis theory. They converge as the proportion of experimental information changes and are used to control the overall error rate at multiple mid-term monitoring points in the A / B test, and allow the test to be terminated early when the effect is significant (example values: in the early stage of the test, the upper boundary value is 0.1 and the lower boundary value is -0.1; in the middle stage of the test, the upper boundary value is 0.06 and the lower boundary value is -0.06; at the end of the test, the upper boundary value is 0.03 and the lower boundary value is -0.03).

[0097] S4.3: Based on the preliminary decision conclusions, conduct compliance verification and risk coefficient assessment to generate an assessment decision.

[0098] Furthermore, based on the preliminary decision conclusions, a compliance verification process is executed, comparing the preliminary decision conclusions with the minimum sample size requirements, maximum test duration limits, and key indicator protection thresholds item by item, and recording the items that violate the rules; a risk coefficient assessment process is executed, obtaining the risk coefficient based on the dual-optimized dynamic confidence interval width, the accuracy rate of historical similar decisions, and the risk tolerance threshold of the current business environment; combining the preliminary decision conclusions, compliance verification results, and risk coefficients, a complete assessment decision is generated, including the decision type, confidence level, risk level, and rule exception explanation.

[0099] It should be noted that the minimum sample size requirement refers to the lower limit of the number of samples required for the experimental and control groups to ensure the reliability of statistical conclusions, which is derived based on statistical power analysis and the minimum detectable effect value.

[0100] The maximum test duration limit is the maximum running time limit (e.g., 30 days) preset to prevent excessive consumption of resources by the test. It is set by taking into account business iteration cycle, seasonal factors and opportunity cost.

[0101] The key performance indicator (KPI) protection threshold (range: 3% to 10%) is defined based on the historical fluctuation range of core business indicators and the acceptable business impact boundary. The business environment risk tolerance threshold (range: 0.1 to 0.7) is defined comprehensively based on the current stage of business development, market competition, and the company's risk appetite.

[0102] S5: Based on the evaluation decision, the mean of the posterior distribution and the corrected variance of the posterior distribution are used as recommended parameters to set the initialization parameters for the next stage of the AB test, and the results are integrated to generate an AB test effect evaluation report.

[0103] S5.1: Make dynamic judgments based on evaluation decisions, and extract the posterior distribution mean parameter and the corrected posterior distribution variance parameter as core recommendation parameters to generate recommendation parameters;

[0104] Furthermore, based on the decision type in the evaluation decision, dynamic judgment is made: when the decision type is to terminate the experiment, the parameter extraction process is triggered; the posterior distribution mean parameter and the corrected posterior distribution variance parameter are obtained from the construction process of the dual-optimization dynamic confidence interval on which the evaluation decision depends, and these parameters are marked as core recommended parameters; the core recommended parameters are combined with the experiment identifier, timestamp and confidence level metadata into a structured data object to generate a set of recommended parameters.

[0105] S5.2: Map and convert the recommended parameter set into the prior distribution center value and minimum detectable effect value of the next stage AB test, and adjust the scaling factor of the posterior distribution variance to optimize sensitivity and generate initialization parameter recommendations;

[0106] Furthermore, the posterior distribution mean parameter in the recommended parameter set is directly mapped to the prior distribution center value of the next stage A / B test. At the same time, the minimum detectable effect value is obtained based on the corrected posterior distribution variance parameter in the recommended parameter set, and the corrected posterior distribution variance parameter is adjusted by a scaling factor to optimize the sensitivity of subsequent tests. Finally, the prior distribution center value, the minimum detectable effect value, and the adjusted variance parameter are integrated into a structured data object to generate an initialization parameter recommendation object.

[0107] S5.3: Integrate initialization parameter recommendations, evaluation decisions, dual-optimization dynamic confidence interval data, and recommended parameter sets to generate an A / B test effect evaluation report.

[0108] Furthermore, the prior distribution center value, minimum detectable effect value, and adjusted variance parameter are extracted from the initialization parameter recommendation object. At the same time, the decision type and risk level in the evaluation decision are obtained. Combined with the lower and upper bound values ​​in the dual-optimized dynamic confidence interval data, as well as the original posterior distribution mean parameter and the corrected posterior distribution variance parameter in the recommendation parameter set, all data are merged into a standardized document through the report generation node to generate a complete A / B test effect evaluation report containing text descriptions, data tables, and statistical charts.

[0109] This embodiment also provides a computer device applicable to the method for evaluating the effect of A / B testing based on dynamic confidence intervals, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for evaluating the effect of A / B testing based on dynamic confidence intervals as proposed in the above embodiment.

[0110] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0111] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the method for evaluating the effectiveness of AB testing based on dynamic confidence intervals as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0112] In summary, this invention achieves prior guidance for effect estimation by fusing historical experimental data to construct a Bayesian prior distribution and dynamically updating it with real-time data, thereby improving the stability and reliability of small-sample stage assessments. Furthermore, by extracting the dynamic variance inflation factor through autoregressive analysis and performing time-series correction on the posterior variance, it effectively compensates for the underestimation of variance caused by data autocorrelation, enhancing the accuracy and robustness of confidence intervals.

[0113] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for evaluating the effectiveness of A / B testing based on dynamic confidence intervals, characterized in that: include, The system collects exposure event data streams and conversion event data streams from A / B tests, aggregates them through a rolling window using a time window, obtains the cumulative exposure counts and conversion counts of the A / B tests within the current time window, and stores them in chronological order to generate the current window's A / B test data and historical time series. Based on the current window AB test data, a set of similar final effect size estimates is retrieved. Through statistical inference methods, a prior distribution of the test effect is generated and updated according to Bayes' theorem. The prior distribution and the current window AB test data are then merged to generate the mean and variance of the posterior distribution. Autoregressive analysis is performed on historical time series to extract autoregressive coefficients. Based on time series analysis theory, dynamic variance inflation factor is obtained. The dynamic variance inflation factor is used to correct the posterior distribution variance to generate the corrected posterior distribution variance. A dual-optimized dynamic confidence interval is constructed using the mean of the posterior distribution and the corrected variance of the posterior distribution. The dual-optimized dynamic confidence interval is compared with the preset sequential test boundary and a decision is made to generate an evaluation decision. Based on the evaluation decision, the mean of the posterior distribution and the corrected variance of the posterior distribution are used as recommended parameters to set the initialization parameters for the next stage of the A / B test, and an A / B test effect evaluation report is generated.

2. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 1, characterized in that: The acquisition of exposure event data streams and conversion event data streams in the A / B test refers to the acquisition of the original event streams from the A / B test, the classification of event types, and the filtering and abnormal recording of events with abnormal timestamps and missing key fields, thereby generating exposure event data streams and conversion event data streams.

3. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 2, characterized in that: The process of using a rolling window aggregation through a time window to obtain the cumulative exposure and conversion counts of the A / B test within the current time window, and storing them in chronological order to generate the current window's A / B test data and historical time series, is as follows: The continuous exposure event data stream and conversion event data stream are divided into discrete time intervals of fixed length by time windows, and grouped according to the test version number to generate a windowed event set; Perform multi-dimensional aggregation operations on the exposure event data stream and conversion event data stream in the windowed event set to obtain the cumulative exposure count and conversion count, and perform correlation matching to generate the current time window metric statistical record; All time window indicator statistics are combined into a continuous historical time series, and then linked and integrated with the current time window indicator statistics to generate the current window A / B test data.

4. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 3, characterized in that: The steps for retrieving a set of similar final effect size estimates based on the current window of AB test data, and generating a prior distribution of the test effect using statistical inference methods are as follows. Based on the current window AB test data, AB test records that meet the similarity threshold are selected from the historical test knowledge base, and the final effect size estimate is extracted to generate a set of historical test effect size estimates. The distribution of the historical experimental effect size estimates was fitted using the Bayesian hierarchical inference method to obtain the prior distribution mean and variance, and then integrated into a prior distribution.

5. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 4, characterized in that: The process of updating the posterior distribution based on Bayes' theorem, fusing the prior distribution and the current window's A / B test data, and generating the mean and variance of the posterior distribution involves the following steps: Based on the cumulative exposure and conversion counts in the current time window's statistical records, calculate the point estimate and naive variance estimate of the effect size for the current window, and generate a likelihood function representation. Based on Bayes' theorem, the prior distribution and the likelihood function are probabilistically fused, and the specific parameter representation of the posterior distribution is obtained through closed-form analytical derivation. The mean and variance of the posterior distribution are then obtained.

6. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 5, characterized in that: The steps for performing autoregressive analysis on historical time series, extracting autoregressive coefficients, and obtaining the dynamic variance inflation factor based on time series analysis theory are as follows. Linear interpolation is used to fill in missing time points in historical time series data, outlier data points are corrected using the three-standard-deviation principle, and stationarity is tested to generate standard historical time series data. We use an autoregressive analysis framework to model and analyze standard historical time series data, identify the optimal order of autoregression, and obtain the estimated values ​​of autoregressive coefficients through the least squares estimation method to generate autoregressive analysis results. Based on the autoregressive analysis results, the variance inflation factor value for the current time window is obtained through the standard formula of time series theory, and its rationality is verified to generate a reasonable dynamic variance inflation factor value.

7. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 6, characterized in that: The process of generating the corrected posterior distribution variance involves combining the reasonable dynamic variance inflation factor value with the posterior distribution variance to correct the posterior distribution variance, and then generating the corrected posterior distribution variance through nonnegativity verification and standardized formatting.

8. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 7, characterized in that: The construction of the dual-optimized dynamic confidence interval refers to obtaining the upper and lower bounds of the dual-optimized dynamic confidence interval based on the mean of the posterior distribution and the corrected variance of the posterior distribution using the standard normal distribution quantile method, and thus constructing the dual-optimized dynamic confidence interval.

9. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 8, characterized in that: The steps for generating the evaluation decision are as follows: The dual-optimized dynamic confidence interval object is compared with the preset sequential test boundary value to dynamically generate preliminary decision conclusions; Based on the preliminary decision conclusions, compliance verification and risk factor assessment are conducted to generate an assessment decision.

10. The method for evaluating the effect of A / B testing based on dynamic confidence intervals as described in claim 9, characterized in that: The evaluation-based decision-making process uses the mean of the posterior distribution and the corrected variance of the posterior distribution as recommended parameters to set the initialization parameters for the next stage of the A / B test, and integrates these parameters to generate an A / B test performance evaluation report. The steps are as follows: Based on the evaluation decision, dynamic judgment is made, and the posterior distribution mean parameter and the corrected posterior distribution variance parameter are extracted as core recommendation parameters to generate recommendation parameters; The recommended parameter set is mapped and converted into the prior distribution center value and minimum detectable effect value of the next stage AB test, and the scaling factor of the posterior distribution variance is adjusted to optimize the sensitivity, generating the initialization parameter recommendations; By integrating the initialization parameter recommendations, evaluation decisions, dual-optimization dynamic confidence interval data, and recommended parameter sets, an A / B test effect evaluation report is generated.

Citation Information

Patent Citations

  • Inertial navigation test calibration frequency optimization method based on confidence theory

    CN114383630A

  • SINS / DVL dynamic alignment method and system based on variational Bayes and medium

    CN118067151A

  • Clinical examination equipment remote calibration method and system based on edge calculation

    CN120727238A

  • Method and apparatus for evaluating fraud risk in an electronic commerce transaction

    US20020194119A1

  • Stochastic variable selection method for model selection

    US20050086010A1