A data processing method and device, computer equipment and storage medium

By identifying experimental and control groups in the information search strategy, calculating the target pooled variance and hypothesis test results, and selecting the target strategy scheme, the problems of search results not meeting user needs and large computational load caused by relying on personal experience are solved, achieving more efficient data processing and more accurate search results.

CN116483882BActive Publication Date: 2026-05-19BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2023-04-26
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

When selecting the optimal strategy from multiple information search strategies, relying on personal experience may result in search results that do not meet user needs, leading to a waste of information resources, as well as high computational load and low efficiency.

Method used

By identifying experimental and control groups from multiple test groups, calculating the pooled variance of the target and the hypothesis test results, selecting the target strategy, and performing data processing.

Benefits of technology

It improves the accuracy of selecting the optimal strategy, reduces the amount of computation, improves processing efficiency, and ensures that search results better meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116483882B_ABST
    Figure CN116483882B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method and device, computer equipment and a storage medium, wherein the method comprises: determining an experimental group and a control group from a plurality of test groups to obtain a target test group containing the experimental group and the control group; the plurality of test groups correspond to different policy schemes respectively; using observation data corresponding to each sample in the experimental group and the control group respectively, determining a target combined variance reflecting overall difference between the plurality of test groups, to determine a hypothesis test result of the target test group based on the target combined variance; based on the hypothesis test result, determining a target policy scheme from a plurality of policy schemes associated with the target test group, to perform data processing based on the target policy scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer processing technology, and more specifically, to a data processing method, apparatus, computer equipment, and storage medium. Background Technology

[0002] In various application scenarios, such as search, there may be a need to select one strategy from multiple different options for practical application. For example, in a search scenario, there are several available information search strategies, and the search results obtained using each strategy may differ for the same search information. In practice, one information search strategy will be selected to complete the information search.

[0003] In this situation, if one relies on personal experience to select an information search strategy from multiple different information search strategies for use in the search application, the selected information search strategy may not be the optimal one among multiple information search strategies. The search results obtained using such an information search strategy are more likely to fail to meet the user's search needs and are more likely to lead to a waste of information resources. Summary of the Invention

[0004] This disclosure provides at least one data processing method, apparatus, computer device, and storage medium.

[0005] In a first aspect, embodiments of this disclosure provide a data processing method, comprising: determining an experimental group and a control group from multiple test groups to obtain a target experimental group including the experimental group and the control group; the multiple test groups each corresponding to different strategy schemes; using the observation data corresponding to each sample in the experimental group and the control group respectively to determine a target pooled variance reflecting the overall differences among the multiple test groups, so as to determine the hypothesis test result of the target experimental group based on the target pooled variance; and based on the hypothesis test result, determining a target strategy scheme from multiple strategy schemes associated with the target experimental group, so as to perform data processing based on the target strategy scheme.

[0006] In one optional implementation, each of the strategy schemes corresponds to a different information search strategy. The test group under each strategy scheme includes multiple samples, and each sample includes a search result determined based on the information search strategy under the strategy scheme. The observation data corresponding to the sample includes the consumption data corresponding to the search result. When the target strategy scheme performs data processing, it is used to determine the search results to be displayed for the search information by adopting the information search strategy corresponding to the target strategy scheme based on the acquired search information.

[0007] In one optional implementation, the target pooled variance is determined as follows: the pooled variance corresponding to the target test group is determined; in response to the difference in sample size among the multiple test groups being less than a first threshold and the difference in variance being less than a second threshold, the target pooled variance corresponding to the multiple test groups is determined based on the pooled variance.

[0008] In one optional implementation, determining the hypothesis test result of the target experimental group based on the target pooled variance includes: determining the mean difference corresponding to the target experimental group based on the observation data corresponding to each sample in the experimental group and the control group of the target experimental group; determining the ratio between the mean difference and the normalized target pooled variance as a statistic of the target experimental group, and determining the hypothesis test result corresponding to the range distribution under multiple tests based on the statistic.

[0009] In one optional implementation, determining the pooled variance corresponding to the target experimental group includes: determining the first sample size and first variance corresponding to the experimental group of the target experimental group, and the second sample size and second variance corresponding to the control group, and determining the pooled variance corresponding to the target experimental group.

[0010] In one optional implementation, the mean difference corresponding to the target experimental group is determined in the following manner: in the target experimental group, a first mean corresponding to the experimental group is determined based on the observation data corresponding to each first sample in the experimental group; and a second mean corresponding to the control group is determined based on the observation data corresponding to each second sample in the control group; the mean difference corresponding to the target experimental group is determined based on the first mean and the second mean.

[0011] In an optional implementation, before determining the target strategy from multiple strategy options associated with the target experimental group based on the hypothesis test results, the method further includes: determining a null hypothesis based on the strategy options corresponding to the experimental group and the control group in the target experimental group; the null hypothesis is used to select and predict a strategy option from multiple strategy options corresponding to the experimental group and the control group; determining the target strategy from multiple strategy options associated with the target experimental group based on the hypothesis test results includes: comparing the hypothesis test results with a preset significance level, and verifying whether the null hypothesis is valid based on the comparison results, so that in response to the null hypothesis being valid, the predicted strategy option is taken as the target strategy option.

[0012] Secondly, embodiments of this disclosure also provide a data processing apparatus, comprising: a first determining module, configured to determine an experimental group and a control group from multiple test groups to obtain a target experimental group including the experimental group and the control group; wherein the multiple test groups correspond to different strategy schemes; a second determining module, configured to use observation data corresponding to each sample in the experimental group and the control group to determine a target pooled variance reflecting the overall differences among the multiple test groups, so as to determine the hypothesis test result of the target experimental group based on the target pooled variance; and a third determining module, configured to determine a target strategy scheme from multiple strategy schemes associated with the target experimental group based on the hypothesis test result, so as to perform data processing based on the target strategy scheme.

[0013] In one optional implementation, each of the strategy schemes corresponds to a different information search strategy. The test group under each strategy scheme includes multiple samples, and each sample includes a search result determined based on the information search strategy under the strategy scheme. The observation data corresponding to the sample includes the consumption data corresponding to the search result. When the target strategy scheme performs data processing, it is used to determine the search results to be displayed for the search information by adopting the information search strategy corresponding to the target strategy scheme based on the acquired search information.

[0014] In one optional embodiment, the apparatus further includes a processing module for determining the target pooled variance in the following manner: determining the pooled variance corresponding to the target test group; and, in response to the fact that the difference in sample size among the plurality of test groups is less than a first threshold and the difference in variance is less than a second threshold, determining the target pooled variance corresponding to the plurality of test groups based on the pooled variance.

[0015] In one optional implementation, when determining the hypothesis test result of the target experimental group based on the target pooled variance, the second determining module is configured to: determine the mean difference corresponding to the target experimental group based on the observation data corresponding to each sample in the experimental group and the control group of the target experimental group; determine the ratio between the mean difference and the normalized target pooled variance as the statistic of the target experimental group, so as to determine the hypothesis test result corresponding to the range distribution under multiple tests based on the statistic.

[0016] In one optional implementation, when determining the pooled variance corresponding to the target experimental group, the processing module is used to: determine the first sample size and first variance corresponding to the experimental group of the target experimental group, and the second sample size and second variance corresponding to the control group, and determine the pooled variance corresponding to the target experimental group.

[0017] In one optional implementation, the mean difference corresponding to the target experimental group is determined in the following manner: in the target experimental group, a first mean corresponding to the experimental group is determined based on the observation data corresponding to each first sample in the experimental group; and a second mean corresponding to the control group is determined based on the observation data corresponding to each second sample in the control group; the mean difference corresponding to the target experimental group is determined based on the first mean and the second mean.

[0018] In an optional implementation, before determining the target strategy scheme from multiple strategy schemes associated with the target experimental group based on the hypothesis test results, the third determining module is further configured to: determine the null hypothesis based on the strategy schemes corresponding to the experimental group and the control group in the target experimental group; the null hypothesis is used to select and predict a strategy scheme from multiple strategy schemes corresponding to the experimental group and the control group; when determining the target strategy scheme from multiple strategy schemes associated with the target experimental group based on the hypothesis test results, the third determining module is configured to: compare the hypothesis test results with a preset significance level, and verify whether the null hypothesis is valid based on the comparison results, so that in response to the null hypothesis being valid, the predicted strategy scheme is taken as the target strategy scheme.

[0019] Thirdly, an optional implementation of this disclosure also provides a computer device, a processor, and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.

[0020] Fourthly, an optional implementation of this disclosure also provides a computer-readable storage medium storing a computer program that, when run, performs the steps of the first aspect or any possible implementation of the first aspect.

[0021] This disclosure provides a data processing method, apparatus, computer device, and storage medium that can identify multiple test groups corresponding to multiple strategy schemes to be screened, and select a target experimental group including an experimental group and a control group. Based on the observation data corresponding to each sample, a target pooled variance reflecting the overall differences among the multiple test groups is determined, and the target pooled variance is used to determine the hypothesis test results for the target experimental group, thereby determining the target strategy scheme. This approach uses observation data that actually reflects the superiority or inferiority of each strategy scheme to perform multiple comparisons, making the selection of the target strategy scheme more accurate compared to methods relying on human experience.

[0022] Furthermore, the data processing method provided in this embodiment of the present disclosure determines the hypothesis test results by using the target merging scheme determined by the selected experimental group and control group when performing multiple comparisons of multiple strategy schemes. This avoids the problem of excessive computation when determining the pooled variance for all test groups under multiple comparisons. Therefore, the time required for data processing is shorter and the efficiency is higher.

[0023] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0024] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart of a data processing method provided by an embodiment of this disclosure is shown;

[0026] Figure 2 A schematic diagram of a data processing apparatus provided in an embodiment of this disclosure is shown;

[0027] Figure 3 A schematic diagram of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown herein can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0029] Research has shown that in application scenarios such as search, different information search strategies can yield different results for the same search information. In practical applications, it's necessary to select the optimal strategy from multiple options to complete the search task. Relying solely on personal experience to choose a search strategy may not be optimal, potentially leading to search results that don't meet the user's needs and wasting information resources.

[0030] Based on the above research, this disclosure provides a data processing method that can identify multiple test groups corresponding to various strategy schemes to be screened, and select a target experimental group including experimental and control groups. Based on the observation data corresponding to each sample, a target pooled variance reflecting the overall differences among the multiple test groups is determined. This target pooled variance is then used to determine the hypothesis test results for the target experimental group, thereby determining the target strategy scheme. This method uses observation data that actually reflects the superiority or inferiority of each strategy scheme to perform multiple comparisons, making the selection of the target strategy scheme more accurate compared to methods relying on human experience.

[0031] In addition, when performing multiple comparisons of selected strategy options, determining the hypothesis test results by using the target pooling scheme determined by the selected experimental and control groups can avoid the problem of excessive computation when determining the pooled variance for all test groups under multiple comparisons. Therefore, the time required for data processing is shorter and the efficiency is higher.

[0032] The shortcomings of the above solutions are the result of the inventor's practical experience and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below should be considered as the inventor's contribution to this disclosure.

[0033] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0034] To facilitate understanding of this embodiment, a data processing method disclosed in this disclosure will first be described in detail. The execution subject of the data processing method provided in this disclosure is generally a computer device with certain computing capabilities. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, an in-vehicle device, a wearable device, etc. In some possible implementations, this data processing method can be implemented by a processor calling computer-readable instructions stored in memory.

[0035] The data processing method provided by the embodiments of this disclosure is described below. Specifically, the data processing method provided by the embodiments of this disclosure can be used to select the optimal solution from multiple solutions, and therefore can be applied to various different fields. These fields include, for example, selecting the optimal information search strategy in a search scenario as described above, or determining a better strategy threshold from multiple selectable strategy thresholds, or selecting the optimal image processing algorithm in the field of image processing, etc. When selecting the optimal solution from multiple solutions, the data processing method provided by the embodiments of this disclosure uses observation data obtained under each solution to perform an evidence-based solution selection. Therefore, compared to screening methods that rely on human experience, it can more easily and accurately select the optimal solution.

[0036] See Figure 1 The diagram shows a flowchart of a data processing method provided in an embodiment of this disclosure. The method includes steps S101 to S103, wherein:

[0037] S101: Determine the experimental group and the control group from multiple test groups to obtain a target experimental group that includes the experimental group and the control group; the multiple test groups correspond to different strategy schemes respectively;

[0038] S102: Using the observation data corresponding to each sample in the experimental group and the control group, determine the target pooled variance that reflects the overall difference between the multiple test groups, and determine the hypothesis test result of the target experimental group based on the target pooled variance;

[0039] S103: Based on the hypothesis test results, determine the target strategy from the multiple strategy options associated with the target experimental group, and perform data processing based on the target strategy.

[0040] The following section will take the search scenario described above as an example to explain S101 to S103 in detail.

[0041] Regarding S101 above, multiple test groups can first be identified, each corresponding to a strategy scheme, and each strategy scheme differs from the others. Specifically, each strategy scheme can correspond to an information search strategy, which is used to complete the information search task. Specifically, after receiving the search information sent by the user, the information search strategy determines how to obtain relevant search results.

[0042] Here, different strategy schemes correspond to different information search strategies. Specifically, this can be reflected in the different search dimensions indicated by the information search strategy. For example, one information search strategy might involve directly performing related searches based on keywords extracted from the search information, while another might involve combining trending words with extracted keywords for related searches. Alternatively, the difference might lie only in certain thresholds. For instance, one information search strategy might involve identifying multiple semantic words with a similarity exceeding 70% to the aforementioned keywords to obtain relevant search results, while another might involve identifying multiple semantic words with a similarity exceeding 80% to the aforementioned keywords to obtain relevant search results. Taking the latter as an example, it is foreseeable that with a 70% threshold, the search results may have low relevance to the keywords, while with an 80% threshold, the number of search results may be small, providing insufficient information to the user. Therefore, it is necessary to select a target strategy scheme from these different strategy schemes.

[0043] For each strategy, a corresponding test group can be determined, and each test group contains multiple samples, such as multiple search results obtained under an information search strategy. The quality of the information search strategy can be expressed by the quality of the search results, which can be evaluated by various indicators, such as user viewing time, number of user likes and comments, number of viewers, etc. Specifically, in this embodiment, these are referred to as the consumption data corresponding to the search results and are used as the observation data of the samples.

[0044] In this embodiment of the disclosure, it can be determined that if the search results under an information search strategy correspond to consumption data indicating that more users have viewed, liked, or commented on them, then this information search strategy is superior. Therefore, by using the observation data of samples under various information search strategies, multiple different information search strategies can be compared to select a superior one as the target search strategy. Correspondingly, this process is also known as determining the target strategy scheme under multiple strategy schemes. For the finally selected target strategy scheme, its corresponding target search strategy can be used for data processing. After receiving a search request sent by the user, the selected target search strategy can be used to obtain search results that better meet the user's search needs, and these results will be displayed to the user as the search results to be shown.

[0045] After identifying multiple test groups under different strategy schemes, experimental and control groups can be selected from them. For multiple strategy schemes, in some cases, there may be several strategies that require particular attention over a period of time. For example, after updating an information search strategy, the focus might be on whether the updated strategy improves upon the previous one. In this case, the corresponding test groups under these two information search strategies would be used as the experimental and control groups. Alternatively, experimental and control groups can be selected from multiple test groups based on the actual situation, which will not be elaborated further here.

[0046] Here, the selected experimental group and control group are specifically referred to as the target experimental group in this embodiment of the disclosure. The experimental group and control group under the target experimental group are independent of each other.

[0047] Regarding S102 above, for the selected target experimental group, the target pooled variance reflecting the overall difference between the multiple test groups can be determined by using the observation data corresponding to each sample in the experimental group and the control group, so as to determine the hypothesis test result of the target experimental group based on the target pooled variance.

[0048] Here, we can first determine the null hypothesis (H0) based on the strategies corresponding to the experimental and control groups in the target experimental group. For example, the null hypothesis could be: the user dwell time X0 in the consumption data of the experimental group is not significantly different from the user dwell time X1 in the consumption data of the control group, that is: H0: X0 = X1; correspondingly, the alternative hypothesis is: the user dwell time X0 in the consumption data of the experimental group is longer than the user dwell time X1 in the consumption data of the control group, that is: H1: X0 > X1. The strategy in the control group is usually the current strategy. Therefore, under the above-mentioned setting of null hypothesis H0 and alternative hypothesis H1, if the null hypothesis cannot be rejected, the current strategy in the control group will continue to be used in actual application.

[0049] When determining the target strategy for the experimental and control groups, this can be done by determining the pooled variance of the target to ascertain the hypothesis test results. In practice, by comparing the hypothesis test results with the preset significance level, it can be determined whether the null hypothesis holds, thereby determining the target strategy.

[0050] Here, for two experimental groups, a Type I error may occur during the screening process. A Type I error specifically refers to incorrectly rejecting the null hypothesis as the screening result. However, the embodiments of this disclosure specifically involve the comparison and screening of multiple test groups, i.e., multiple comparisons, which further leads to the Type I error inflation problem in multiple comparisons, i.e., it is more likely to occur.

[0051] In this scenario, for multiple comparisons, a Tukey's test can be used. The Tukey's test requires using the variance information of all test groups when comparing two groups. However, in practice, multiple comparisons often involve a large number of test groups. If only two test groups are selected each time to determine the variance information, the amount of data required for computation would be enormous.

[0052] Therefore, in this embodiment of the disclosure, the following method is specifically used to determine the target pooled variance: determine the pooled variance corresponding to the target test group; in response to the fact that the difference in sample size among the multiple test groups is less than a first threshold and the difference in variance is less than a second threshold, determine the target pooled variance corresponding to the multiple test groups based on the pooled variance. Here, the target pooled variance corresponding to the multiple test groups obtained is also the variance information of all test groups under the graph basis test as described above. For ease of description, the pooled variance described above is denoted as S, and the target pooled variance determined using the pooled variance S is denoted as S'.

[0053] First, the method used to determine the pooled variance S corresponding to the target experimental group is explained. Specifically, the pooled variance S can be determined as follows: The first sample size and first variance corresponding to the experimental group of the target experimental group, and the second sample size and second variance corresponding to the control group are determined, thereby determining the pooled variance corresponding to the target experimental group.

[0054] The following example illustrates this. For the experimental group and the control group, i... n Specifically, for example, i0 can represent the experimental group and i1 can represent the control group. Experimental group i n The corresponding sample size can be expressed as n in Therefore, the first sample size corresponding to the above experimental group is expressed as n. i0 The second sample size corresponding to the control group is denoted as n.i1 Experimental group i n The corresponding variance is denoted as σ. in Similarly, the first variance corresponding to the above experimental group is denoted as σ. i0 The second variance corresponding to the control group is denoted as σ. i1 .

[0055] Based on the above representation, when determining the pooled variance S of the target experimental group, it specifically satisfies the following formula (1):

[0056]

[0057] As for the selected target experimental group and the original multiple test groups, in fact, when obtaining test groups through random distribution, no matter how many test groups are obtained, the sample size of the test groups is similar, and the distribution of the observation data of each test group is similar before the experiment.

[0058] The specific reasons are explained below: First, when dividing the samples into test groups, if there is a large difference in the sample size between the test groups, then the random allocation method itself is problematic. Therefore, no matter what method is used, the reliability of the selection result obtained after determining the target strategy through the test groups cannot be guaranteed to be correct. Thus, when determining the test groups through random allocation, the sample size of each test group should be similar.

[0059] Secondly, considering that control variables are needed when comparing test groups, and that the differences in strategy schemes between test groups are generally small, especially when there are a large number of test groups, such as in the application scenario described in the embodiments of this disclosure, the differences between test groups may only be in the threshold. Therefore, it can be foreseen that the changes in the observed data of the samples in each test group are small, so the variances between test groups remain in a relatively close state.

[0060] Based on the above explanation of sample size and variance, the following two assumptions can be obtained, specifically (2-1) and (2-2):

[0061]

[0062]

[0063] Based on this, when determining the target pooled variance S' for multiple test groups, the following formula (3) can be satisfied.

[0064]

[0065] Therefore, the pooled variance S of the target test group and the target pooled variance S' of the multiple test groups satisfy the following formula (4):

[0066]

[0067] That is, after determining the pooled variance corresponding to the target test group, the target pooled variance corresponding to multiple test groups can be approximately determined by using the above formula (4). This method can effectively reduce the computational load when calculating the target pooled variance S' using the above formula (3) when there are multiple test groups, and can effectively improve efficiency.

[0068] After determining the target pooled variance, the mean difference of the target experimental group can be determined based on the observation data corresponding to each sample in the experimental group and the control group of the target experimental group. The ratio between the mean difference and the target pooled variance after normalization is determined as the statistic of the target experimental group, and the hypothesis test result corresponding to the range distribution under multiple tests is determined based on the statistic.

[0069] Here, for example, in the consumption data of the experimental group listed in the above example, the user dwell time is X0, and the user dwell time in the consumption data of the control group is X1. When determining the mean difference Δ, it can be determined that the mean difference Δ = X0 - X1.

[0070] Here, when determining the mean difference, the user dwell time X1 corresponding to the experimental group and the user dwell time X2 corresponding to the control group are selected, both of which can be expressed as means. Therefore, the mean difference can be determined as follows: In the target experimental group, based on the observation data corresponding to each first sample in the experimental group, a first mean X0 corresponding to the experimental group is determined; and based on the observation data corresponding to each second sample in the control group, a second mean X1 corresponding to the control group is determined; based on the first mean and the second mean, the mean difference corresponding to the target experimental group is determined.

[0071] By using the mean difference Δ and normalizing the pooled variance of the target, the statistic t can be determined, where t satisfies t = Δ / S'. The statistic t specifically follows a range distribution under multiple tests, which can be a studentized range distribution (t-formed range). Therefore, the hypothesis test result (p-value) corresponding to the statistic t can be obtained by looking up a table.

[0072] Regarding S103 above, given the hypothesis test results, the target strategy can be determined from among the multiple strategy options associated with the target experimental group.

[0073] Specifically, based on the relevant explanation in S101 above, the null hypothesis H0 can be determined. The p-value of the determined hypothesis test result can be compared with the preset significance level α to determine whether the null hypothesis H0 is valid. In this case, if the null hypothesis H0 is not valid, the strategy predicted in the alternative hypothesis H1 is taken as the target strategy. Here, the preset significance level α can be set to 0.05.

[0074] In the specific comparison, if p < α, then the null hypothesis H0 is rejected, meaning that the user dwell time X0 in the consumption data of the experimental group is longer than the user dwell time X1 in the consumption data of the control group, and the corresponding strategy in the experimental group is superior. Conversely, if p < α, then the strategy of the control group is maintained, meaning that the user dwell time X0 in the consumption data of the experimental group is similar to the user dwell time X1 in the consumption data of the control group, and therefore the original strategy of the control group is not changed and is maintained.

[0075] In this way, for the two different strategy schemes corresponding to the experimental group and the control group under the target experimental group, it is possible to determine one of the better strategy schemes, or decide to maintain the original strategy scheme. In practical application scenarios, a more suitable information search strategy can be determined and implemented in the actual search scenario to process the received search information and obtain the search results to be displayed to the user.

[0076] In another embodiment of this disclosure, after determining the target strategy scheme from the multiple strategy schemes associated with the target test group, the following method can be used to determine the final selected target strategy scheme from the strategy schemes corresponding to the multiple test groups respectively: in response to the existence of an unselected test group among the multiple test groups, the test group and the test group corresponding to the target strategy are determined as a new target test group, and a new target strategy scheme is determined based on the new target test group, so as to determine the target strategy scheme for obtaining the search results corresponding to the search information from the multiple strategy schemes.

[0077] In the steps described above, two key test groups were selected from multiple test groups as target experimental groups to determine the optimal target strategy. However, in reality, there are multiple selectable test groups, and correspondingly, multiple selectable strategy schemes. Based on this, all test groups can be traversed to determine the optimal target strategy from the multiple strategy schemes corresponding to the multiple test groups using the observation data of the samples in each test group.

[0078] Since the above steps have identified a superior strategy for one of the target test groups, this test group can be combined with another untested test group to form a new target test group. This process is repeated to continuously compare and select a superior test group from multiple test groups. Specifically, the strategy under this superior test group results in longer user dwell time in the observed data, meaning the corresponding information search strategy yields search results that better meet user needs and effectively reduces wasted information display. This type of strategy is also more suitable for processing received search requests.

[0079] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0080] Based on the same inventive concept, this disclosure also provides a data processing device corresponding to the data processing method. Since the principle of the device in this disclosure for solving the problem is similar to that of the data processing method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0081] Reference Figure 2 The diagram shown is a schematic representation of a data processing apparatus provided in an embodiment of this disclosure. The apparatus includes: a first determining module 21, a second determining module 22, and a third determining module 23; wherein,

[0082] The first determining module 21 is used to determine the experimental group and the control group from multiple test groups to obtain a target experimental group that includes the experimental group and the control group; the multiple test groups correspond to different strategy schemes respectively;

[0083] The second determining module 22 is used to determine the target pooled variance, which reflects the overall difference between the multiple test groups, using the observation data corresponding to each sample in the experimental group and the control group, so as to determine the hypothesis test result of the target experimental group based on the target pooled variance.

[0084] The third determining module 23 is used to determine a target strategy scheme from multiple strategy schemes associated with the target experimental group based on the hypothesis test results, so as to perform data processing based on the target strategy scheme.

[0085] In one optional implementation, each of the strategy schemes corresponds to a different information search strategy. The test group under each strategy scheme includes multiple samples, and each sample includes a search result determined based on the information search strategy under the strategy scheme. The observation data corresponding to the sample includes the consumption data corresponding to the search result. When the target strategy scheme performs data processing, it is used to determine the search results to be displayed for the search information by adopting the information search strategy corresponding to the target strategy scheme based on the acquired search information.

[0086] In an optional embodiment, the apparatus further includes a processing module 24, configured to determine the target pooled variance in the following manner: determine the pooled variance corresponding to the target test group; in response to the difference in sample size among the plurality of test groups being less than a first threshold and the difference in variance being less than a second threshold, determine the target pooled variance corresponding to the plurality of test groups based on the pooled variance.

[0087] In one optional implementation, when determining the hypothesis test result of the target experimental group based on the target pooled variance, the second determining module 22 is configured to: determine the mean difference corresponding to the target experimental group based on the observation data corresponding to each sample in the experimental group and the control group of the target experimental group; determine the ratio between the mean difference and the normalized target pooled variance as the statistic of the target experimental group, so as to determine the hypothesis test result corresponding to the range distribution under multiple tests based on the statistic.

[0088] In one optional implementation, when determining the pooled variance corresponding to the target experimental group, the processing module 24 is used to: determine the first sample size and first variance corresponding to the experimental group of the target experimental group, and the second sample size and second variance corresponding to the control group, and determine the pooled variance corresponding to the target experimental group.

[0089] In one optional implementation, the mean difference corresponding to the target experimental group is determined in the following manner: in the target experimental group, a first mean corresponding to the experimental group is determined based on the observation data corresponding to each first sample in the experimental group; and a second mean corresponding to the control group is determined based on the observation data corresponding to each second sample in the control group; the mean difference corresponding to the target experimental group is determined based on the first mean and the second mean.

[0090] In an optional implementation, before determining the target strategy scheme from multiple strategy schemes associated with the target experimental group based on the hypothesis test results, the third determining module 23 is further configured to: determine the null hypothesis based on the strategy schemes corresponding to the experimental group and the control group in the target experimental group; the null hypothesis is used to select and predict a strategy scheme from multiple strategy schemes corresponding to the experimental group and the control group; when determining the target strategy scheme from multiple strategy schemes associated with the target experimental group based on the hypothesis test results, the third determining module 23 is configured to: compare the hypothesis test results with a preset significance level, and verify whether the null hypothesis is valid based on the comparison results, so that in response to the null hypothesis being valid, the predicted strategy scheme is taken as the target strategy scheme.

[0091] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0092] This disclosure also provides a computer device, such as... Figure 3 The diagram shown is a schematic representation of a computer device structure provided in an embodiment of this disclosure, including:

[0093] Processor 10 and memory 20; the memory 20 stores machine-readable instructions executable by processor 10, and processor 10 executes the machine-readable instructions stored in memory 20. When the machine-readable instructions are executed by processor 10, processor 10 performs the following steps:

[0094] An experimental group and a control group are determined from multiple test groups to obtain a target experimental group that includes the experimental group and the control group; each of the multiple test groups corresponds to a different strategy scheme; using the observation data corresponding to each sample in the experimental group and the control group, a target pooled variance reflecting the overall difference among the multiple test groups is determined, and the hypothesis test result of the target experimental group is determined based on the target pooled variance; based on the hypothesis test result, a target strategy scheme is determined from the multiple strategy schemes associated with the target experimental group, and data processing is performed based on the target strategy scheme.

[0095] The aforementioned memory 20 includes a main memory 210 and an external memory 220; the main memory 210, also known as internal memory, is used to temporarily store the computational data in the processor 10, as well as the data exchanged with external memory 220 such as a hard disk. The processor 10 exchanges data with the external memory 220 through the main memory 210.

[0096] The specific execution process of the above instructions can be referred to the steps of the data processing method described in the embodiments of this disclosure, and will not be repeated here.

[0097] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the data processing method described in the above method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0098] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the data processing method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0099] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0100] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0101] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0102] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0103] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A data processing method, characterized in that, include: Experimental and control groups are determined from multiple test groups to obtain a target experimental group that includes the experimental and control groups; the multiple test groups correspond to different strategy schemes. Using the observation data corresponding to each sample in the experimental group and the control group, a target pooled variance reflecting the overall difference among the multiple test groups is determined, and the hypothesis test result of the target experimental group is determined based on the target pooled variance. Based on the hypothesis test results, a target strategy is determined from multiple strategy options associated with the target experimental group, and data processing is performed based on the target strategy. The hypothesis testing results for determining the target experimental group based on the target pooled variance include: Based on the observation data corresponding to each sample in the experimental group and the control group of the target experimental group, the mean difference corresponding to the target experimental group is determined; The ratio between the mean difference and the normalized target pooled variance is determined as the statistic for the target experimental group, and the hypothesis test result corresponding to the range distribution under multiple tests is determined based on the statistic.

2. The method according to claim 1, characterized in that, Each of the aforementioned strategy schemes corresponds to a different information search strategy. The test group under each strategy scheme includes multiple samples. Each sample includes a search result determined based on the information search strategy under the strategy scheme. The observation data corresponding to the sample includes the consumption data corresponding to the search result. When processing data, the target strategy scheme is used to determine the search results to be displayed for the search information by adopting an information search strategy corresponding to the target strategy scheme based on the acquired search information.

3. The method according to claim 1 or 2, characterized in that, The target pooled variance is determined as follows: Determine the pooled variance corresponding to the target experimental group; In response to the fact that the difference in sample size among the multiple test groups is less than a first threshold and the difference in variance is less than a second threshold, a target pooled variance corresponding to the multiple test groups is determined based on the pooled variance.

4. The method according to claim 3, characterized in that, Determining the pooled variance corresponding to the target experimental group includes: Determine the first sample size and first variance of the experimental group corresponding to the target experimental group, and the second sample size and second variance of the control group, and determine the pooled variance corresponding to the target experimental group.

5. The method according to claim 1, characterized in that, The mean difference corresponding to the target experimental group was determined in the following manner: In the target experimental group, a first mean is determined based on the observation data corresponding to each first sample in the experimental group; and a second mean is determined based on the observation data corresponding to each second sample in the control group. Based on the first mean and the second mean, the mean difference corresponding to the target experimental group is determined.

6. The method according to claim 1, characterized in that, Before determining the target strategy from multiple strategy options associated with the target experimental group based on the hypothesis test results, the method further includes: Based on the strategy schemes corresponding to the experimental group and the control group in the target experimental group, the null hypothesis is determined; the null hypothesis is used to select and predict one strategy scheme from the multiple strategy schemes corresponding to the experimental group and the control group, respectively. The step of determining the target strategy from multiple strategy options associated with the target experimental group based on the hypothesis test results includes: The hypothesis test results are compared with a preset significance level, and the null hypothesis is verified based on the comparison results. If the null hypothesis is true, the predicted strategy is taken as the target strategy.

7. A data processing apparatus, characterized in that, include: The first determining module is used to determine the experimental group and the control group from multiple test groups to obtain a target experimental group that includes the experimental group and the control group; the multiple test groups correspond to different strategy schemes respectively; The second determining module is used to determine the target pooled variance, which reflects the overall difference between the multiple test groups, by using the observation data corresponding to each sample in the experimental group and the control group, so as to determine the hypothesis test result of the target experimental group based on the target pooled variance. The third determination module is used to determine a target strategy from multiple strategy options associated with the target experimental group based on the hypothesis test results, and to perform data processing based on the target strategy. The hypothesis testing results for determining the target experimental group based on the target pooled variance include: Based on the observation data corresponding to each sample in the experimental group and the control group of the target experimental group, the mean difference corresponding to the target experimental group is determined; The ratio between the mean difference and the normalized target pooled variance is determined as the statistic for the target experimental group, and the hypothesis test result corresponding to the range distribution under multiple tests is determined based on the statistic.

8. A computer device, characterized in that, include: A processor and a memory, the memory storing machine-readable instructions executable by the processor, the processor executing the machine-readable instructions stored in the memory, wherein when the machine-readable instructions are executed by the processor, the processor performs the steps of the data processing method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer device, performs the steps of the data processing method as described in any one of claims 1 to 6.