Traffic distribution method and device, equipment and storage medium
By matching feature sets based on the historical data of the use objects in AB testing and allocating them to test groups with high similarity, the problem of ignoring exception feedback in the prior art is solved, and more accurate test results and decision support are achieved.
Patent Information
- Application Number
- CN202510406307.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art ignores abnormal feedback from different groups of users in AB testing, resulting in the overall feedback results ignoring potential vulnerabilities and affecting the decision accuracy of the service platform.
By determining its attribute feature set based on the historical data of the usage object and matching the preset reference feature set, the object is assigned to a test group with high similarity, ensuring the accuracy of the AB test results.
It improves the accuracy of AB test results, can identify the feedback of groups of users of different attribute characteristics, and enhances the accuracy of subsequent decisions.
Smart Images

Figure CN120281719A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a traffic allocation method, apparatus, device, and storage medium. Background Technique
[0002] For a service platform that provides Internet services, whenever there is an update to the service function (such as modifying the original function, adding new functions, etc.), before pushing the updated function to all users of the service platform (i.e., before full-scale rollout), the updated function will first be pushed to a part of the users for a small-scale A / B test to determine whether the updated function can provide effective gain. The so-called A / B test means randomly dividing all users who log in to the service platform into at least two experimental groups, and letting users in different experimental groups experience different versions of the updated function (such as the original version and the new version). Then, feedback data (such as click-through rate, conversion rate, etc.) from different experimental groups is collected to determine which version is better.
[0003] In the above A / B test process, how to allocate traffic to each user who logs in to the service platform is a key link to improve the accuracy of the test results. The so-called traffic allocation means allocating each user participating in the A / B test (i.e., traffic) to different experimental groups according to certain rules. The core goal of traffic allocation is to ensure the comparability of users between different experimental groups, so as to accurately evaluate the effect differences of different versions of the updated function.
[0004] Under the related technology, when allocating traffic to each user, usually some users are randomly selected from all users of the service platform, and when these users log in to the service platform for network activities, they are respectively allocated to different experimental groups.
[0005] However, the test results obtained in this way can only represent the overall feedback of each user on the updated function, but ignore the abnormal feedback of different user groups on the updated function. In other words, even if a first user group shows relatively obvious and important negative feedback on the new version of the updated function, due to the large number of positive feedbacks provided by other users, the overall feedback after comprehensive averaging shows a positive feedback result, which in turn causes the service platform to ignore the feedback data provided by this user group.
[0006] For example, assume that for an update function, only the user group that heavily uses the update function (e.g., the continuous usage duration reaches more than 10 hours) can trigger the vulnerability of the update function. Therefore, the user group that heavily uses the update function provides negative feedback for this reason. However, since other users do not trigger the corresponding vulnerability, the overall feedback for this update function is a positive feedback result, causing the service platform to ignore the vulnerability existing in the update function, which may lead to relatively serious impacts.
[0007] Therefore, there is an urgent need for a new traffic allocation method to improve the accuracy in the traffic allocation process. Summary of the Invention
[0008] The embodiments of the present application provide a traffic allocation method, device, equipment and storage medium to improve the accuracy in the traffic allocation process.
[0009] In a first aspect, a traffic allocation method is provided, which is applied to a target service platform and includes:
[0010] When each first user logs in to the target service platform, based on the historical data of each first user in the target service platform, the attribute feature set of each first user is obtained respectively;
[0011] Among them, for each first user, the following are performed respectively:
[0012] Based on the reference feature sets of each preset test group, the set similarity between each reference feature set and the attribute feature set of a first user is determined respectively; where each reference feature is an attribute feature with a feature homogeneity degree reaching a first set threshold among the first users included in a test group; each of the test groups includes at least two experimental groups for A / B testing;
[0013] Based on the set similarity of each obtained reference feature set, the first user is allocated to an experimental group in a target test group whose set similarity reaches a second set threshold.
[0014] In a second aspect, a traffic allocation device is provided, which is applied to a target service platform and includes:
[0015] An acquisition module: used to obtain the attribute feature set of each first user respectively when each first user logs in to the target service platform, based on the historical data of each first user in the target service platform;
[0016] Processing module: For each of the first usage objects, perform respectively: Based on the reference feature sets of each preset test group, determine the set similarity between each reference feature set and the attribute feature set of a first usage object; wherein, each reference feature is: an attribute feature with a feature homogeneity degree reaching a first set threshold among the first usage objects included in a test group; each of the test groups includes: at least two experimental groups for A / B testing.
[0017] Allocation module: Based on the set similarity of each obtained reference feature set, allocate the first usage object to one of the experimental groups in the target test group where the set similarity reaches a second set threshold.
[0018] Optionally, the processing module is further configured to:
[0019] Based on the historical data of each second usage object in the target service platform, obtain the attribute feature set of each second usage object respectively; wherein, the second usage object is: other usage objects in the target service platform except the first usage objects.
[0020] Based on the set similarity between the attribute feature sets of the second usage objects, perform feature clustering on the second usage objects to obtain the test groups; wherein, in each test group, the set similarity between the attribute feature sets of any two second usage objects reaches a group threshold.
[0021] For each of the test groups, perform respectively: Based on the feature homogeneity degree among the attribute features of the second usage objects in a test group, obtain the reference feature set of the test group.
[0022] Optionally, when the acquisition module is used to obtain the reference feature set of a test group based on the feature homogeneity degree among the attribute features of the second usage objects in the test group, it is specifically configured to:
[0023] In the attribute feature sets of the second usage objects included in the test group, screen out at least one candidate feature; wherein, each candidate feature is: an attribute feature with a difference degree from the attribute features of other test groups among the test groups reaching a preset difference threshold.
[0024] For each of the candidate features, perform respectively:
[0025] Based on a candidate feature, obtain the feature homogeneity degree corresponding to the candidate feature based on the feature similarity between every two second usage objects in the test group; wherein, the value of the feature homogeneity degree is positively correlated with the value obtained after fusing the obtained feature similarities.
[0026] When the feature homogeneity degree corresponding to the one candidate feature reaches the first set threshold, the one candidate feature is used as a reference feature of the one test group.
[0027] Optionally, after obtaining the test groups and before the first users log in to the target service platform, the processing module is further configured to:
[0028] For each of the test groups, perform respectively:
[0029] Respectively obtain the identity identifier of each second user included in a test group in the target service platform;
[0030] In the one test group, create at least two blank experimental groups;
[0031] Based on the obtained identity identifiers, randomly assign each second user to one of the at least two experimental groups.
[0032] Optionally, when the obtaining module is used to respectively obtain the attribute feature set of each first user based on the historical data of each first user in the target service platform, it is specifically configured to:
[0033] For each of the first users, perform respectively:
[0034] Based on the historical data of a first user in the target service platform, obtain the generation time of each interaction data generated after the interaction between the first user and the target service platform;
[0035] Respectively based on the obtained generation times, obtain the time weights of the corresponding interaction data; wherein, the time weight is negatively correlated with the duration between the generation time and the current moment;
[0036] Based on the obtained time weights, perform weighted adjustment on the interaction data, and perform comprehensive feature extraction on the weighted-adjusted interaction data and other data in the historical data except the interaction data, to obtain the attribute feature set of the first user.
[0037] Optionally, when the obtaining module is used to respectively obtain the attribute feature set of each first user based on the historical data of each first user in the target service platform, it is specifically configured to:
[0038] Based on a preset extraction time range, obtain the historical data of each of the first usage objects from the target service platform; wherein, the generation time of each interaction data included in each historical data is within the extraction time range; and the interaction data is generated after the corresponding first usage object interacts with the target service platform.
[0039] For the historical data of each of the first usage objects respectively, perform feature extraction to obtain the attribute feature set of each of the first usage objects.
[0040] Optionally, the target test group further includes: a retention group not used for A / B testing.
[0041] When the allocation module is used to allocate a first usage object to one of the experimental groups in a target test group where the set similarity of each of the obtained reference feature sets reaches a second set threshold, specifically:
[0042] Based on the set similarity of each of the obtained reference feature sets, allocate the first usage object to a target test group where the set similarity reaches the second set threshold.
[0043] For the retention group, randomly allocate the first usage object based on the identity identifier of the first usage object in the target service platform.
[0044] When the first usage object is not randomly allocated to the retention group, allocate the first usage object to one of the at least two experimental groups included in the target test group; wherein, for one experimental group, the change range of the set homogeneity degree between the attribute feature sets of the allocated usage objects in the experimental group before and after allocating the first usage object conforms to a preset amplitude condition.
[0045] Optionally, when the allocation module is used to allocate the first usage object to one of the at least two experimental groups included in the target test group, specifically:
[0046] For each of the at least two experimental groups, respectively execute:
[0047] For the set similarity between the attribute feature sets of every two allocated usage objects in an experimental group, perform fusion processing to obtain the corresponding first set homogeneity degree.
[0048] Take the first usage object as a new assigned object in the experimental group, and re - perform the fusion process on the set similarity between the attribute feature sets of every two assigned usage objects in the experimental group to obtain the corresponding second set of homogeneity degree;
[0049] Based on the first set of homogeneity degree and the second set of homogeneity degree, determine the change range of the set homogeneity degree after the first usage object is assigned to the experimental group;
[0050] Assign the first usage object to an experimental group where the change range of the set homogeneity degree does not exceed the amplitude threshold.
[0051] Optionally, when the processing module is used to determine the set similarity between each reference feature set of each preset test group and the attribute feature set of a first usage object respectively, it is specifically used for:
[0052] For each reference feature set in each reference feature set of each test group, perform respectively:
[0053] Based on a preset dimension threshold, perform dimensionality reduction processing on a reference feature set to obtain a dimensionality - reduced reference feature set;
[0054] Based on the dimension threshold, perform dimensionality reduction processing on each associated attribute feature associated with the reference feature set in the attribute feature set of the usage object to obtain the dimensionality - reduced associated attribute features;
[0055] Based on the dimensionality - reduced reference feature set and the dimensionality - reduced associated attribute features, obtain the set similarity between the reference feature set and the first usage object.
[0056] Optionally, when the processing module is used to determine the set similarity between each reference feature set of each preset test group and the attribute feature set of a first usage object respectively, it is specifically used for:
[0057] For each reference feature set in each reference feature set of each test group, perform respectively:
[0058] Obtain each associated attribute feature associated with a reference feature set in the attribute feature set of the first usage object;
[0059] Respectively determine the association similarity between each associated attribute feature included in each other first usage object except the first usage object among the first usage objects and the associated feature attributes of the first usage object;
[0060] Based on the obtained correlation similarities, the associated attribute features, and the one reference feature set, determine the set similarity between the one reference feature set and the one first usage object; wherein, the value of the set similarity is positively correlated with the values of the correlation similarities.
[0061] In a third aspect, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method as described in the first aspect.
[0062] In a fourth aspect, there is provided a computer device comprising:
[0063] a memory for storing program instructions;
[0064] a processor for invoking the program instructions stored in the memory and executing the method as described in the first aspect according to the obtained program instructions.
[0065] In a fifth aspect, there is provided a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the method as described in the first aspect.
[0066] In the embodiments of the present application, when a first usage object logs in to the target service platform, the traffic allocation process for the first usage object is started. First, the historical data of the first usage object in the target service platform is used to obtain the attribute feature set of each first usage object.
[0067] Secondly, for each first usage object among the first usage objects, the allocation operation can be separately performed, that is: based on the reference feature sets of the preset test groups, the similarity between each reference feature set and the attribute feature set of a first usage object is determined respectively. Since the reference features in the reference feature set are part of the attribute features selected from all the attribute feature sets and the degree of homogeneity reaches the first set threshold, therefore, by only comparing the similarity between the reference feature set and the attribute feature set, the accuracy of the similarity of the reference feature set between the two can be ensured, and the computing resources required for similarity calculation can be reduced, improving the efficiency and effect of the process of allocating the first usage object to the target test group.
[0068] Further, after determining the similarity between a first user object and each test group, a test group whose similarity meets the second set threshold can be selected as the target test group, and a first user object can be assigned to the experimental group in the target test group that meets the conditions. In this way, a test group with a relatively high similarity is selected as the target test group for a first user object, ensuring that the similarity of the reference features among the first user objects included in the target test group can present a relatively high degree that meets the second set threshold. In other words, it ensures the degree of homogeneity among the property features related to the reference attribute set among the first user objects in the target test group. Therefore, when setting up the corresponding experimental group for AB testing in this target test group, it can first ensure the degree of homogeneity of the user objects participating in the AB testing within the range of the reference feature set, so that the AB test result can be strongly correlated with the reference feature set in the target test group.
[0069] On the other hand, corresponding AB tests are set for different test groups, taking advantage of the differences between the reference feature sets of each test group, ensuring that each AB test can conduct corresponding tests for user object groups with different attribute features, ensuring that the obtained feedback data can all represent the usage experiences of user object groups with different attribute features, and further enabling developers to obtain the key content masked in the overall data based on the detailed differences in different feedback data, thereby improving the accuracy of subsequent decisions. Description of the Drawings
[0070] Figure 1 An application scenario of the traffic allocation method provided by the embodiment of the present application;
[0071] Figure 2 A flowchart of the traffic allocation method provided by the embodiment of the present application Figure 1 ;
[0072] Figure 3 A schematic diagram of the principle of the traffic allocation method provided by the embodiment of the present application Figure 1 ;
[0073] Figure 4 A schematic diagram of the principle of the traffic allocation method provided by the embodiment of the present application Figure 2 ;
[0074] Figure 5 A flowchart of the traffic allocation method provided by the embodiment of the present application Figure 2 ;
[0075] Figure 6 A schematic diagram of the principle of the traffic allocation method provided by the embodiment of the present application Figure 3 ;
[0076] Figure 7 Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 4 ;
[0077] Figure 8 Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 5 ;
[0078] Figure 9 Schematic diagram of the process of the traffic allocation method provided by an embodiment of the present application Figure 3 ;
[0079] Figure 10 Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 6 ;
[0080] Figure 11 Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 7 ;
[0081] Figure 12 Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 8 ;
[0082] Figure 13A Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 9 ;
[0083] Figure 13B Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 10 ;
[0084] Figure 14 Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 10 One;
[0085] Figure 15 Schematic diagram of the principle of the traffic allocation method provided by an embodiment of the present application Figure 10 Two;
[0086] Figure 16 Schematic diagram of the structure of the traffic allocation device provided by an embodiment of the present application Figure 1 ;
[0087] Figure 17 Schematic diagram of the structure of the traffic allocation device provided by an embodiment of the present application Figure 2 。 Detailed implementation manners
[0088] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.
[0089] The following explains some terms in the embodiments of the present application to facilitate understanding by those skilled in the art.
[0090] (1) A / B Testing:
[0091] A / B Testing refers to randomly dividing each user object logging in to the service platform into at least two experimental groups, and allowing the user objects in different experimental groups to experience different versions of the updated function (such as the original version and the new version) respectively, so as to evaluate the effect differences of each version, and then determine the best version.
[0092] (2) Traffic allocation:
[0093] Traffic allocation refers to the process of allocating each user object (i.e., traffic) participating in the A / B test to different experimental groups according to certain rules, so as to scientifically compare the effect differences of different versions (such as functions, interfaces, algorithms, etc.), and at the same time ensure the reliability and effectiveness of the test results.
[0094] (3) Target service platform:
[0095] The target service platform refers to any one of the platforms that provide Internet services for each user object, and its types can include: social platforms, instant messaging platforms, multimedia display platforms, etc. On the other hand, the manifestation form of the target service platform can be a web page, an application program, a small program inside the application program, etc.
[0096] (4) Historical data:
[0097] Historical data refers to: the identity information, interaction data, etc. recorded in the target service platform when the user object uses the target service platform.
[0098] It should be noted that in the embodiments of the present application, operations such as obtaining historical data and attribute feature sets are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0099] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0100] The application fields of the traffic allocation method provided by the embodiments of the present application will be briefly introduced below.
[0101] After developers on the service platform develop new service functions for the service platform, before pushing these updated functions to all users of the service platform (i.e., before full-scale rollout), corresponding A / B tests will be conducted on these updated functions first.
[0102] Under the related technology, in the process of traffic allocation for A / B testing, usually, some users are randomly selected from all users of the service platform, and when these users log in to the service platform for network activities, they are respectively assigned to different experimental groups.
[0103] However, the test results obtained in this way can only represent the overall feedback of each user on the updated function, but ignore the abnormal feedback of different user groups on the updated function. In other words, even if a first user group shows obvious and important negative feedback on the new version of the updated function, due to the numerous positive feedbacks provided by other users, the overall feedback after comprehensive averaging shows a positive feedback result, which in turn causes the service platform to ignore the feedback data provided by this user group.
[0104] For example, assume that for an updated function, only the user group that heavily uses this updated function (e.g., the continuous usage duration reaches more than 10 hours) can trigger the vulnerability of this updated function. Therefore, the user group that heavily uses this updated function provides negative feedback for this reason. However, since other users do not trigger the corresponding vulnerability, the overall feedback for this updated function is a positive feedback result, causing the service platform to ignore the vulnerability existing in this updated function, which may in turn lead to more serious impacts.
[0105] In view of this, the embodiments of the present application provide a traffic allocation method. In this method: when each first user logs in to the target service platform, based on the historical data of these first users in the target service platform, the attribute feature set of each first user is respectively obtained, and the feature information of each first user in different dimensions is obtained.
[0106] In this way, for each first user among the first users, the following operations can be respectively performed: First, based on the reference feature sets of each preset test group, the similarity between each reference feature set and the attribute feature set of a first user is respectively determined; where each reference feature is an attribute feature with a homogeneity degree reaching a first set threshold among the first users included in a test group; at least two experimental groups in the A / B test are respectively set for each test group.
[0107] Thus, for a first user object, it can first determine the similarity between itself and a plurality of pre-set test groups, and then select a suitable test group as the target test group according to the similarity. Moreover, for different test groups, only the similarity between a first user object and the reference feature set of the test group needs to be determined to complete the classification process for a first user object, improving the efficiency of the classification process for a first user object.
[0108] After obtaining the target test group for a first user object, after allocating the first user object to the experimental group corresponding to the target test group, the traffic allocation for the first user object can be completed.
[0109] Thus, in the traffic allocation process provided in this application, when the first user object logs in to the target service platform, the traffic allocation process for the first user object is started. First, the attribute feature set of each first user object is obtained by using the historical data of the first user object in the target service platform.
[0110] Secondly, for each first user object among the first user objects, the allocation operation can be separately executed, that is: based on the reference feature sets of the pre-set test groups respectively, the similarity between each reference feature set and the attribute feature set of a first user object is determined respectively. Since the reference features in the reference feature set are part of the attribute features whose homogeneity reaches the first set threshold selected from all the attribute feature sets, only comparing the similarity between the reference feature set and the attribute feature set can ensure the accuracy of the similarity of the reference feature set between the two, and can also reduce the computing resources consumed during the similarity calculation, improving the efficiency and effect of the process of allocating the first user object to the target test group.
[0111] Further, after determining the similarity between a first user object and each test group, a test group whose similarity meets the second set threshold can be selected as the target test group, and a first user object can be assigned to the experimental group in the target test group that meets the conditions. In this way, a test group with a higher similarity is selected as the target test group for a first user object, ensuring that the similarity of the reference features among the first user objects included in the target test group can present a relatively high degree that meets the second set threshold. In other words, it ensures the degree of homogeneity among the attribute features related to the reference attribute set among the first user objects in the target test group. Therefore, when corresponding experimental groups are set up in the target test group for A / B testing, it can first ensure the degree of homogeneity among the user objects participating in the A / B testing within the range of the reference feature set, and further make the A / B test result strongly correlated with the reference feature set in the target test group.
[0112] On the other hand, corresponding A / B tests are set for different test groups, taking advantage of the differences between the reference feature sets of each test group, ensuring that each A / B test can perform corresponding tests on user object groups with different attribute features, ensuring that the obtained feedback data can all represent the usage experiences of user object groups with different attribute features, and further enabling developers to obtain the key content hidden in the overall data based on the detailed differences in different feedback data, thereby improving the accuracy of subsequent decisions.
[0113] The application scenario of the traffic allocation method provided by the present application will be described below.
[0114] Please refer to Figure 1 , which is a schematic diagram of an application scenario of the traffic allocation method provided by the present application. This application scenario includes a server 101 and a client 102. Communication can be carried out between the server 101 and the client 102. The communication method can be communication using wired communication technology. For example, communication can be carried out by connecting a network cable or a serial cable; it can also be communication using wireless communication technology. For example, communication can be carried out through technologies such as Bluetooth or wireless fidelity (WIFI), and specific details are not limited.
[0115] The client 102 generally refers to a device that can execute downstream tasks and the like based on the result of the allocation of user objects by the traffic allocation method. For example, a terminal device, a third-party application program that can be accessed by the terminal device, or a web page that can be accessed by the terminal device, etc. The server 101 generally refers to a device that can perform traffic allocation and the like for user objects. For example, a terminal device or a server, etc.
[0116] The terminal device includes, but is not limited to, mobile phones, computers, intelligent medical devices, smart home appliances, vehicle-mounted terminals, or aircraft, etc. The server includes, but is not limited to, cloud servers, local servers, or associated third-party servers, etc. Both the client 101 and the server 102 can adopt cloud computing to reduce the occupation of local computing resources; similarly, cloud storage can also be adopted to reduce the occupation of local storage resources.
[0117] As an embodiment, the server 101 and the client 102 can be the same device, or they can be different devices respectively, or they can be different devices with some modules shared, etc., and no specific restrictions are made.
[0118] The execution subject of the embodiment of the present application can be the server 101. During the process of traffic allocation, when the first user object logs in to the target service platform through the client 102, the server 101 can, according to the traffic allocation method provided by the embodiment of the present application, allocate the first user object to the target experimental group in the target test group, so as to start the experience of the first user object for the update function and collect the corresponding feedback data.
[0119] For the traffic allocation method provided by the embodiment of the present application, it can be applied to different AB test tasks.
[0120] For example, for an item trading platform, when the platform hopes to optimize the recommendation algorithm and improve the success rate of item trading, it can adopt the traffic allocation method provided by the embodiment of the present application. After allocating the user objects logging in to the item trading platform to different test groups, the user objects allocated to the corresponding target test group are set with corresponding AB experiments, so that some user objects experience the traditional recommendation algorithm, and some user objects experience the optimized personalized recommendation algorithm. Then, based on the feedback data of different versions, the effect of the optimized recommendation algorithm is tested.
[0121] Another example is that for an online education platform, when the platform hopes to optimize the course content and display interface, it can also adopt the traffic allocation algorithm provided by the embodiment of the present application. The user objects logging in to the online education platform are divided into different test groups based on age or learning stage, and AB tests are respectively carried out in each test group, so that some user objects experience the standard courses and display interfaces applicable to all user objects, and some user objects experience the course content and display interfaces customized according to age, learning stage, etc., so as to evaluate the effects of the new and old versions in different experimental groups.
[0122] For another example, for a content distribution platform, when it hopes to optimize its information push strategy and increase the corresponding click-through rate and reading duration, it can adopt the traffic allocation method provided in the embodiments of the present application. The usage objects logged in to the content distribution platform are divided into test groups with different reference feature sets, and then corresponding A / B tests are set in each test group to experiment with the push strategies of the old and new versions and their impacts on the feedback data of the usage objects.
[0123] The above has given examples of the application fields of the traffic allocation method provided in the embodiments of the present application. In actual applications, this traffic allocation method can also be applied to other A / B test tasks, and the present application does not limit this.
[0124] Next, based on Figure 1 , a specific introduction to the traffic allocation method provided in the embodiments of the present application will be given. Please refer to Figure 2 , which is a flowchart illustration of a traffic allocation method provided in the embodiments of the present application. Figure 1 .
[0125] S21: When each first usage object logs in to the target service platform, based on the historical data of each first usage object in the target service platform, the attribute feature set of each first usage object is obtained respectively.
[0126] Among them, the first usage object is a usage object that has used the network service provided by the target service platform in the target service platform. Since these usage objects have already performed interaction behaviors in the target service platform, corresponding historical data has been recorded and saved in the target service platform.
[0127] For the obtained historical data, before obtaining the corresponding attribute feature set, corresponding data cleaning and preprocessing can be performed on these historical data to ensure the accuracy, consistency, and integrity of the data, providing a reliable basis for subsequent traffic allocation processing. Exemplarily, corresponding operations such as missing value processing, outlier detection, and normalization processing can be performed on the historical data to obtain standard structured data.
[0128] When the traffic allocation method provided in the embodiments of the present application starts to be implemented, whenever a first usage object logs in to the target service platform, the attribute feature set of the corresponding first usage object can be obtained from the historical data of the first usage object in the target service platform.
[0129] Exemplarily, feature extraction can be performed from the identity information and interaction data recorded by the first usage object in the target service platform to obtain the corresponding attribute feature set, and each attribute feature therein represents different feature information of the first usage object, such as age, login frequency, browsing duration, click frequency, etc.
[0130] Thus, after obtaining the attribute feature sets of each first usage object, the following operations can be separately performed for each first usage object:
[0131] S22: Based on the reference feature sets of each preset test group, respectively determine the set similarity between each reference feature set and the attribute feature set of a first usage object; wherein, each reference feature is an attribute feature with a feature homogeneity degree reaching a first set threshold among the first usage objects included in a test group.
[0132] First, an introduction to each preset test group is given. For the target service platform, in each preset test group, a certain number of usage objects are included. As Figure 3 shown, in a test group, among the attribute feature sets of the included usage objects, there are some attribute features whose feature homogeneity degree with the corresponding attribute features of other usage objects in this test group reaches the first set threshold, and these attribute features that reach the first set threshold are the reference features of this test group. Exemplarily, assume there is a test group where the values of the attribute features corresponding to the ages of the included usage objects are all between 24 and 26 years old. Then, for this test group, one of the reference features is the age feature of the usage object.
[0133] Thus, for the target service platform, each different test group represents a group of users of this target service platform; furthermore, through these test groups, the target service platform can relatively clearly clarify the feature distribution of its usage objects, so as to set corresponding AB tests for different types of usage objects.
[0134] Therefore, when the target service platform hopes to test any function or content, at least two experimental groups in the corresponding AB test can be set in different test groups according to the test requirements.
[0135] Optionally, when the target service platform starts the AB test, when allocating traffic to each first usage object logging in to the target service platform, as Figure 4 shown, each preset test group can be a part of the test groups screened from all the test groups that involve the AB test requirements. Exemplarily, when the target service platform needs to test a game function, then some test groups whose reference feature sets are related to age, login duration, and activity level can be pre-screened from all the test groups as the test groups participating in this AB test, so as to start the corresponding AB test. In this way, the test results obtained from the AB test can relatively intuitively reflect the correlation degree between the reference feature set of the test group and the feedback data of this game function to be tested.
[0136] The above introduced the meaning of the test groups for the traffic allocation method provided by the embodiments of the present application. For the acquisition of these test groups, the following methods can be used:
[0137] In a possible implementation manner, when the target service platform needs to obtain test groups, corresponding clustering processing can be performed based on the historical data of the usage objects it contains.
[0138] Please refer to Figure 5 , which is a process schematic of the traffic allocation method provided by the embodiments of the present application Figure 2 , as Figure 5 shown. The specific implementation steps of this method are as follows:
[0139] S301: Based on the historical data of each second usage object in the target service platform, obtain the attribute feature set of each second usage object respectively; where the second usage object is: other usage objects in the target service platform except each first usage object.
[0140] Since the essence of the process of obtaining test groups is actually to obtain a group of usage objects with a relatively high degree of homogenization of some attribute features, therefore, similar to step S21 above, at the beginning of the acquisition of test groups, it is first necessary to obtain the attribute feature set of each second usage object based on the historical data of each second usage object in the target service platform respectively.
[0141] On the other hand, the second usage object is other usage objects in the target service platform except each first usage object. This content can be understood in combination with the process of obtaining test groups: as Figure 6 shown, in the process of obtaining test groups, scattered usage objects will gradually aggregate into different groups through the similarity between their respective attribute features. These aggregated groups can be used as the corresponding test groups. Therefore, for the second usage object, the process of dividing it into test groups has been completed.
[0142] The first usage object needs to be divided into different test groups based on its own attribute feature set after logging in to the target service platform. Therefore, the first usage object is a usage object that has not been divided based on these test groups yet. So, each first usage object does not overlap with each second usage object among all the usage objects in the target service platform.
[0143] S302: Based on the set similarity between the attribute feature sets of each second usage object, perform feature clustering on the second usage objects to obtain each test group; where, for each test group, the set similarity between the attribute feature sets of any two second usage objects reaches the group threshold.
[0144] In this step, when performing feature clustering on each second usage object through the set similarity between the attribute feature sets of the second attribute objects, it can be achieved through different offline clustering methods. For example, the K-Means Clustering algorithm, the Hierarchical Clustering algorithm, or the density-based clustering algorithm, etc.
[0145] Exemplarily, taking K-Means clustering as an example, the feature clustering process for each second usage object can be achieved based on the following steps:
[0146] First, determine the number of clusters K in the clustering process. This number of clusters K represents the number of test groups that each second usage object needs to be divided into during this clustering process.
[0147] The determination of this K value can be carried out through the Elbow Method or the Silhouette Coefficient. Among them, the Elbow Method is a method for determining the optimal K value by analyzing the law of the within-cluster sum of squares of the clustering model changing with the number of clusters K. In this method, it is considered that when K is less than the true number of clusters, the within-cluster sum of squares will decrease rapidly, and after exceeding the true number of clusters, the improvement of the within-cluster sum of squares will become insignificant. Therefore, by selecting the K value at the inflection point of these two changes, the optimal number of test groups for each second usage object to be clustered can be determined. Specifically, the execution steps of this method can be: calculate the within-cluster sum of squares of the clustering model in each case from K = 2 to K = 10, and then draw the corresponding graph based on the obtained results, observe the change trend, and determine the optimal K value.
[0148] Next, after determining the K value, the following operations can be continued: First, randomly select K usage objects from the second usage objects as the initial clustering centers, and then calculate the distance (such as the Euclidean distance) between the attribute feature set of each second usage object and the attribute feature sets of each clustering center, and assign the second usage object to the cluster corresponding to the nearest clustering center. After completing the assignment of the second usage objects, calculate the mean value of all second usage objects in each cluster and use it as the new clustering center.
[0149] In this way, repeat the above process for the assignment operation of the second usage objects until the clustering centers no longer change or reach the maximum number of iterations.
[0150] In this way, for the second usage objects after division, the corresponding clusters are the respective test groups. And in each test group, the set similarity between the attribute feature sets of the second usage objects it contains can reach the group threshold.
[0151] After obtaining each test group, the following operations can be continued:
[0152] S303: For each test group, perform the following respectively: Based on the feature homogeneity degree among the attribute features of each second usage object in a test group, obtain the reference feature set of the test group.
[0153] Among the second usage objects included in each obtained test group, not all the attribute features in their corresponding attribute feature sets have a high feature similarity. Instead, for some of the features, the homogeneity degree in the test group is relatively high.
[0154] Therefore, for each test group, it is also necessary to screen out the reference feature set corresponding to the test group from all the attribute features included in the attribute feature set, so as to determine the same features of each second usage object included in each test group.
[0155] In a possible way, for the screening of the reference feature set, each attribute feature in the attribute feature set can be traversed, and its feature homogeneity degree in the test group can be calculated respectively, and then the attribute features whose feature homogeneity degree reaches the first set threshold are selected as the corresponding reference features.
[0156] In another possible implementation manner, the following method can also be adopted to obtain the reference feature set:
[0157] First, in the attribute feature sets of each second usage object included in a test group, screen out at least one candidate feature; among them, each candidate feature is an attribute feature whose difference degree from the attribute features of other test groups in each test group reaches a preset difference threshold.
[0158] In this process, what needs to be determined first is the relatively important attribute features in each test group. For this purpose, taking the above clustering method as an example, for the clustering center corresponding to each test group, the feature values of each cluster center can be compared horizontally to find the feature with the largest difference. For example, the "login duration" of cluster A is significantly higher than that of cluster B, indicating that this feature is the key to distinguishing the two.
[0159] Therefore, based on the difference degree between attribute features, at least one candidate feature can be screened out from the attribute feature set relatively quickly. Among them, the difference threshold can be set in advance through practical experience or obtained based on model training, and this application does not limit this.
[0160] In this way, after obtaining at least one candidate feature, the following operations can be performed for each candidate feature respectively:
[0161] Based on a candidate feature, obtain the feature homogeneity degree corresponding to a candidate feature according to the feature similarity between every two second usage objects in a test group; wherein, the value of the feature homogeneity degree is positively correlated with the value obtained after the fusion processing of each obtained feature similarity.
[0162] As Figure 7 shown, for a candidate feature, its corresponding feature homogeneity degree is determined by the similarity between the values of the same attribute feature of every two second usage objects among the second usage objects. In this test group, the more similar the values among these attribute features are, the higher the feature homogeneity degree of this candidate feature in this test group. In terms of the specific relationship, the value obtained after the fusion of the similarity between every two second usage objects is positively correlated with the value of the feature homogeneity degree.
[0163] Therefore, when the feature homogeneity degree corresponding to this candidate feature reaches the first set threshold, this candidate feature can be used as a reference feature of a test group.
[0164] By analogy, after performing the above operations for other candidate features respectively, a corresponding reference feature set can be obtained based on the feature homogeneity degree corresponding to each candidate feature.
[0165] In the above method, by first screening candidate features and then screening reference features, the reference feature set of the test group is quickly determined from each attribute feature included in the test group, which can not only ensure that these reference feature sets are important features with greater differences from other test groups, but also ensure that the homogeneity degree of these reference features within the same test group is relatively high. Furthermore, it ensures that when conducting the AB test for the test group subsequently, the homogeneity degree among the usage objects participating in the test is relatively high.
[0166] Through the above method, the acquisition of the reference feature set corresponding to each test group is completed, and then the acquisition of different test groups is completed. For these test groups, among the usage objects included therein, the feature homogeneity degree of some attribute features is relatively high. Therefore, the usage objects in this test group can be determined to be a usage group with a certain feature. Further, by independently setting an AB test for each test group, the test result can have a stronger correlation with the reference feature set in the test group, making it easier to analyze the relationship between its advantages or disadvantages and the attribute features of the usage objects from the test result, and then optimizing the updated function in the AB test.
[0167] Based on these test groups, their primary function is to divide the various users in the target service platform into different groups. Due to the differences in the reference feature sets among different groups, the same user can be classified into different test groups based on their own attribute feature set. Exemplarily, assume that in test group A, the reference feature set represents: left-handed, highly active users; in test group B, the reference feature set represents: aged 20 - 25, with a large number of transactions. Then a user with all these 4 features can be classified into both test group A and test group B.
[0168] When the target service platform is about to start an A / B test, based on its expected target group, it can select some test groups whose reference feature sets meet the expectations from all the test groups as the test groups participating in this A / B test, and for each selected test group, set at least two corresponding experimental groups. Then these test groups containing at least two experimental groups can be used as the preset test groups to provide a basis for traffic allocation for each first user.
[0169] In a possible implementation manner, when setting at least two corresponding experimental groups for each test group, the following method can be adopted for each test group:
[0170] First, obtain the identity identifiers of each second user included in a test group in the target service platform respectively.
[0171] For each user, the target service platform assigns an identity identifier to distinguish the users in the target service platform. Therefore, in this process, these identity identifiers can be used as the basis for random allocation.
[0172] Next, create at least two blank experimental groups in a test group; where each experimental group corresponds to a different version of the function to be tested. For example, experimental group A is used to guide the users in it to experience version 1 of the function to be tested, and experimental group B is used to guide the users in it to experience version 2 of the function to be tested, and so on. More experimental groups can also be set so that users can experience other versions of the test function.
[0173] Finally, based on the obtained identity identifiers, randomly assign each second user to one of the at least two experimental groups.
[0174] Since the identity identifiers of the users are usually combinations of numbers or characters, it is possible to directly perform random partitioning using the identity identifiers. For the second users with identity identifiers ranging from 0000001 to 9999999, a threshold of 5555555 can be selected. The second users with identity identifiers before the threshold are assigned to experimental group 1, and the second users with identity identifiers after the threshold are assigned to experimental group 2.
[0175] Furthermore, in order to better improve the randomness among the users assigned to different experimental groups, the following method can also be used to assign the second users.
[0176] As Figure 8 shown, first, the hash algorithm and the preset hash seed can be used to perform hash operations on the identity identifiers of each second user to obtain the corresponding hash values. Among them, the hash algorithm can be the MD5 algorithm or the MurmurHash algorithm, etc.
[0177] Secondly, determine the corresponding number of buckets. For example, it can be 1000 or 10000 to facilitate ratio adjustment (such as 10% traffic corresponding to 100 or 1000 buckets).
[0178] Then, based on the hash value corresponding to the identity identifier of each second user, take the modulus of the number of buckets to obtain the bucket number corresponding to each second user (for example, 0 - 999), and assign the second user to the corresponding bucket.
[0179] In this way, when dividing the experimental groups, different buckets can be selected for partitioning. For example, assuming there are only two experimental groups, buckets 0 - 555 can be assigned to experimental group 1, and buckets 556 - 999 can be assigned to experimental group 2. Among them, the number of experimental groups is only used for illustrative purposes.
[0180] In this way, for each preset test group, among the at least two experimental groups it contains, a certain number of second users have already been assigned to support the basic test of the AB test. And through the pre - assignment of each second user during the process of obtaining the test group, the corresponding experimental group can be quickly obtained, enabling the AB test to be quickly started and improving the acquisition efficiency of the AB test results.
[0181] Optionally, for the second users assigned to the experimental groups in this way, they can be the users predicted to be likely to log in to the target service platform during the AB test process.
[0182] After clarifying the process of setting up the test group, the reference feature set, and each experimental group, the following operations can be continued for each first user logging in to the target service platform:
[0183] Step S23: Based on the set similarity of each obtained reference feature set, assign a first user to one of the experimental groups in the target test group whose set similarity reaches the second set threshold.
[0184] For each first user logging in to the target service platform, by using the set similarity between the attribute feature set of the first user and each reference feature set, the first user can be assigned to the target test group whose set similarity meets the second set threshold, and then assigned to one of the experimental groups in the target test group.
[0185] In the embodiment of the present application, when the first user logs in to the target service platform, the traffic allocation process for the first user is started. First, use the historical data of the first user in the target service platform to obtain the attribute feature set of each first user.
[0186] Secondly, for each first user among the first users, the allocation operation can be performed separately, that is: based on the reference feature sets of each preset test group, determine the similarity between each reference feature set and the attribute feature set of a first user respectively. Since each reference feature in the reference feature set is a part of the attribute features whose homogeneity degree reaches the first set threshold selected from all the attribute feature sets, therefore, only comparing the similarity between the reference feature set and the attribute feature set can ensure the accuracy of the similarity of the reference feature set between the two, and can also reduce the computing resources consumed during the similarity calculation, improving the efficiency and effect of the process of assigning the first user to the target test group.
[0187] Furthermore, after determining the similarity between a first user and each test group, the test group whose similarity meets the second set threshold can be selected as the target test group, and a first user can be assigned to one of the experimental groups in the target test group that meets the conditions. In this way, the test group with a higher similarity is selected as the target test group for a first user, ensuring that the similarity of the reference features among the first users included in the target test group can reach a relatively high degree that meets the second set threshold. In other words, it ensures the homogeneity degree of the attribute features related to the reference attribute set among the first users in the target test group. Therefore, when an AB test is set up with a corresponding experimental group in this target test group, it can first ensure the homogeneity degree of the users participating in the AB test within the range of the reference feature set, and then make the AB test result strongly correlated with the reference feature set in the target test group.
[0188] On the other hand, corresponding A / B tests are set for different test groups, taking advantage of the differences between the reference feature sets of each test group, ensuring that each A / B test can target user groups with different attribute features for corresponding tests, ensuring that the feedback data obtained can represent the usage experiences of user groups with different attribute features, and further enabling developers to obtain key content hidden in the overall data based on the detailed differences in different feedback data, thereby improving the accuracy of subsequent decisions.
[0189] The above describes the process of allocating each first user to an experimental group in a target test group using the historical data of each first user. To further improve the accuracy of the allocation process, the time attribute of the historical data of each first user also needs to be considered. For this purpose, the embodiments of the present application provide the following possible implementation manners:
[0190] In a possible implementation manner, a time decay mechanism can be introduced to process the historical data.
[0191] Specifically, for each first user, the following operations can be performed separately:
[0192] Please refer to Figure 9 , which is a schematic flow of the traffic allocation method provided by the embodiments of the present application Figure 3 .
[0193] As Figure 9 shown, the specific implementation steps of this method are as follows:
[0194] S401: Based on the historical data of a first user on the target service platform, obtain the generation time of each interaction data generated after the first user interacts with the target service platform.
[0195] Among them, the interaction data refers to: data generated after the first user logs in to the target service platform and interacts based on the target service platform or other users in the target service platform.
[0196] Since these interaction data have a certain timeliness, recent interaction data can better reflect current interests and preferences than long-term interaction data. Therefore, in this method of introducing a time decay mechanism, it is first necessary to determine the generation time of each interaction data.
[0197] S402: Based on the obtained generation times respectively, obtain the time weights of the corresponding interaction data; among them, the time weight is negatively correlated with the duration between the generation time and the current moment.
[0198] Since data closer to the current moment can better reflect the preferences of the user, for the time weight, its value is negatively correlated with the duration between the generation time and the current moment, that is: the closer the generation time is to the current moment for the interaction data, the greater the value of the corresponding time weight.
[0199] S403: Based on the obtained time weights, perform weighted adjustment on each interaction data, and perform comprehensive feature extraction on each weighted-adjusted interaction data and other data in the historical data except for each interaction data, to obtain the attribute feature set of the first user.
[0200] After obtaining the time weights of each interaction data, based on these time weights, weighted adjustment can be performed on the interaction data, so that the interaction data closer to the current moment is more referenceable.
[0201] Then, based on each of these weighted-adjusted interaction data and other data in the historical data except for each interaction data (such as data on the age, gender, etc. of the user), perform comprehensive feature extraction to obtain the attribute feature set corresponding to a first user.
[0202] In this method, a time decay mechanism is introduced to reduce the weight of long-term interaction data in the historical data, so that the attribute feature set of the first user can better represent its current preferences, and further improve the accuracy of the allocation process in the subsequent allocation process.
[0203] In another possible implementation, a sliding window mechanism is introduced, that is, only the data within a certain time window is processed, and the expired data is automatically eliminated.
[0204] Specifically, first, based on the preset extraction time range, obtain the historical data of each first user from the target service platform; among them, the generation time of each interaction data included in each historical data is within the extraction time range; the specific value of this extraction time range can be determined according to the actual needs of the user, and this application does not limit it.
[0205] Then, based on the obtained historical data, perform feature extraction respectively, and then obtain the attribute feature set of each first user.
[0206] In this method, only the interaction data within a certain extraction time range is processed, ensuring that the extracted set of attribute features can represent the preferences of the usage object within the specified time range. Furthermore, it can make the allocation result for the first usage object more accurate and effectively reflect the test results for different time nodes. For example, the interaction data generated within the fixed selected rest days can be chosen for subsequent traffic allocation operations. In this way, the results of the AB test can more accurately reflect the feedback on the functions tested during the AB test within the rest days.
[0207] As Figure 10 shown, for the two above-mentioned time-related implementation methods, they can be implemented separately or combined and implemented in a single traffic allocation process, that is: within the corresponding window of the selected extraction time range, corresponding time weights are set for each interaction data, and corresponding feature extraction is performed on the historical data to obtain the set of attribute features.
[0208] The above introduced several possible implementation methods for attribute feature extraction. For the process of allocating the first usage object to the target test group, the embodiments of the present application also provide several possible implementation methods:
[0209] In a possible implementation method, when respectively determining the set similarity between each reference feature set and the attribute feature set of a first usage object, for each reference feature set in the respective reference feature sets of each test group, the following operations can be respectively performed:
[0210] First, based on a preset dimension threshold, dimensionality reduction processing is performed on a reference feature set to obtain a dimensionality-reduced reference feature set;
[0211] Then, based on the dimension threshold, dimensionality reduction processing is similarly performed on each associated attribute feature associated with a reference feature set in the attribute feature set of the one usage object to obtain the dimensionality-reduced associated attribute features.
[0212] Finally, based on the dimensionality-reduced reference feature set and the dimensionality-reduced associated attribute features, the set similarity between the one reference feature set and the one first usage object is obtained.
[0213] In the above process, the high-dimensional data existing in the reference feature set and each associated attribute feature is mapped to a low-dimensional space (which can be a two-dimensional grid), reducing the difficulty of determining the set similarity between the reference feature set and each associated attribute feature and improving the efficiency of obtaining the set similarity.
[0214] For the above method, the following will be described with a clustering algorithm. Exemplarily, this process can be implemented in the following manner:
[0215] First, define a two-dimensional grid structure (such as a rectangle or a hexagon). In this two-dimensional network structure, the position of each neuron is represented by coordinates (i, j), and the dimension of the weight vector of each neuron is the same as that of the input data (i.e., the associated attribute features of the first usage object).
[0216] Next, for each input data X, find the neuron closest to it (referred to as the best matching unit), that is, the neuron with the smallest Euclidean distance.
[0217] Finally, update the weights of the best matching unit and its neighboring neurons to make its weight vector approach the input data direction. Among them, the weight update formula is:
[0218] wi(t + 1) = wi(t) + α × hci × [x(t) - wi(t)] (Formula 1)
[0219] Where α is the learning rate that decays with time, hci is the neighborhood function (such as a Gaussian function, which controls the update intensity of the neurons around the best matching unit), x(t) - wi(t) is the weight update direction, which is used to push the neuron weight closer to the current input data x(t), and wi(t) is the weight of node i.
[0220] In another possible implementation, when respectively determining the set similarity between each reference feature set and the attribute feature set of a first usage object, for each reference feature set in the respective reference feature sets of each test group, the following operations can also be respectively performed:
[0221] First, obtain the associated attribute features associated with a reference feature set in the attribute feature set of a first usage object;
[0222] Secondly, respectively determine the association similarity between the associated attribute features included in each other first usage object except one first usage object in each first usage object and the associated feature attributes of one first usage object;
[0223] Finally, based on the obtained association similarities, associated attribute features, and a reference feature set, determine the set similarity between a reference feature set and a first usage object; among them, the value of the set similarity is positively correlated with the value of each association similarity.
[0224] In this way, the importance degree of the data of a first usage object is determined by the similarity degree between the associated attribute features of a first usage object and the associated attribute features of other first usage objects, that is: the higher the similarity with the associated attribute features of a first usage object, and the larger the number of other first usage objects with a higher similarity, the higher the importance degree of the data of the first usage object.
[0225] In this method, the similarity between the user object and the reference feature set of the test group is jointly judged from two aspects: the importance degree of the data and the similarity degree between the data, so that the value of the similarity is positively correlated with both the importance degree and the similarity degree, and the test group to which the first user object should be assigned can be quickly determined, improving the efficiency of the assignment process.
[0226] For the above method, its implementation process can be realized by a clustering algorithm based on the gravitational model.
[0227] Exemplarily, after obtaining the associated attribute features corresponding to each first user object and a reference feature set, for the associated attribute features corresponding to each first user object, the quality corresponding to this data point can be calculated, that is, the importance degree of this data point in the nearby area. As Figure 11 shown, combined with the above process, it can be considered that the value of each obtained associated similarity is positively correlated with the quality of this data point. In other words, the higher the associated similarity between the data point and other data points, and the more similar other data points, the greater the importance degree of this data point and the corresponding quality. Similarly, the quality of the data point corresponding to the reference feature set is calculated in the same way.
[0228] Then, through the following formula 2, combining the quality of the reference feature set, the quality of each associated attribute feature, and the distance between the two, the gravity between the two is determined.
[0229]
[0230] Among them, F represents gravity, G represents a preset gravitational constant, m i 、m j represent the quality of the data point, d(x i *x j ) represents the distance between data points (for example, Euclidean distance), and p represents the distance attenuation factor.
[0231] In this way, after calculating the gravity between each associated attribute feature and the reference feature set, the data points can be assigned to the corresponding clusters according to the magnitude of the gravity between the data points, that is: the gravity between similar data points (close distance) is large, promoting them to gather into the same cluster. That is to say, the value of the gravity is positively correlated with the similarity between two data points.
[0232] For the above two possible methods of obtaining the set similarity, they can be used alternatively in a single traffic allocation task. Among them, the method of selecting between the above two methods is as follows: Different methods can be adopted in advance to process some of the first usage objects, and based on the level of the gain effect on the AB test obtained from the results, the target processing method adopted in this traffic allocation task can be determined. Or, the two methods can also be used in combination and executed together in a single traffic allocation task. The corresponding processing method can be selected based on whether the attributes of the usage objects to be allocated meet the corresponding conditions.
[0233] It should be noted that for the method of obtaining the set similarity, other clustering algorithms can also be adopted to obtain the target test group corresponding to each first usage object, as long as it is based on the similarity between the reference attribute set and the attribute feature set. This application places no restrictions on this.
[0234] The above has respectively introduced several possible implementation methods for allocating the first usage object to the target test group. After the first usage object is allocated to the target test group, the process of allocating it to one of at least two experimental groups in the target test group can be completed through the following several possible implementation methods.
[0235] In a possible implementation method, for the first usage object allocated to the target test group, a similar allocation method to that of the second usage object described above can be adopted, and based on the identity identifier of the first usage object, it can be randomly allocated to one of at least two experimental groups.
[0236] This method can quickly complete the allocation of the experimental group for the first usage object and can ensure the randomness of the test samples in each experimental group.
[0237] In another possible implementation method, for the first usage object allocated to the target test group, a two-layer partitioning method can be adopted to complete the operation of allocating its experimental group.
[0238] Specifically, after allocating a first usage object to a target test group whose set similarity reaches the second set threshold, the following operations can continue to be performed:
[0239] First of all, it should be noted that in this target test group, in addition to at least two experimental groups, there is also a retention group. The usage objects allocated to the retention group do not participate in the relevant content of the AB test process.
[0240] Therefore, for the reserved group in the target test group, it is necessary to randomly assign a first user object based on the identity identifier of the first user object in the target service platform, and the obtained assignment result indicates whether the first user object is assigned to the reserved group.
[0241] Among them, for the process of assigning to the reserved group, the same method of randomly assigning the second user object as described above can be adopted. Briefly, through the hash operation of the identity identifier, the user object is randomly assigned to an experimental group, and the specific process is not elaborated in this application.
[0242] When a first user object is randomly assigned to the reserved group, the first user object will no longer participate in the relevant content of the A / B test.
[0243] When a first user object is not randomly assigned to the reserved group, it is necessary to assign the first user object to one of the at least two experimental groups included in the target test group.
[0244] Optionally, the above process can be implemented in a double-hash manner. For example, for each user object included in a target test group, these user objects need to be divided into a reserved group and a non-reserved group first. At this time, the above-mentioned hash modulo method can be used to randomly assign these user objects, so that some user objects are assigned to the reserved group and do not participate in the A / B test corresponding to the target test group, while the other part of the user objects are used as the user objects participating in the A / B test of the target test group.
[0245] For the user objects participating in the A / B test, based on the hash algorithm again, they are evenly split and assigned, so that these user objects can be evenly split into the corresponding experimental groups according to a proportion.
[0246] Among them, for the user objects participating in the A / B test, as Figure 12 shown, these user objects can also be grouped again according to the special attribute characteristics each user object has, so that a part of the user objects are used as special user objects and are only assigned to a certain number of experimental groups, while the other part of the user objects are used as normal user objects and are assigned to all experimental groups in an evenly split manner.
[0247] For the part of the user objects that are only assigned to a certain number of experimental groups, before the A / B test starts, the fairness of the grouping of this part of the user objects can also be ensured by means of multiple re-randomizations.
[0248] Exemplarily, different hash seeds can be utilized multiple times to perform a hash operation on the usage objects in this part, thereby obtaining corresponding bucketing results. Furthermore, from the different bucketing results, the bucketing result that best meets the requirements is selected as the allocation result for the usage objects in this part.
[0249] For the process of equal traffic splitting, it can also be based on the hash algorithm. Exemplarily, assume that the total weight of the experimental groups in the target test group is total (i.e., the total number of experimental groups). Equal traffic splitting is performed using the bucket number seq. At this time, there are total buckets in each interval, and the corresponding relationship between the number of experimental groups assigned to each bucket is randomly shuffled.
[0250] Therefore, three hash seeds need to be obtained first. These three hash seeds are respectively: the unique identifier of the experimental group (hash_seed), the equal sampling seed (equal_sampling_seed), and the interval sequence number (such as seq / total). Based on the preset Blake3 hash algorithm, a hash operation is performed on the three hash seeds to generate a random bit array h.
[0251] In this way, after generating the bit array h, classification processing can continue according to the total number of experimental groups. When the total number does not exceed the threshold (e.g., 20), it can be determined as a small-scale experiment, and the Cantor inverse expansion can be used to ensure strict full permutation randomness; while when the total number exceeds the threshold, it is determined as a large-scale experiment, and the offset modulo operation needs to be used to balance the performance and uniformity of the equal traffic splitting process.
[0252] Optionally, for the allocation process of the experimental groups, it can be based on the number of already allocated usage objects included in each experimental group.
[0253] Exemplarily, as Figure 13A shown, when there are experimental groups in which the number of already allocated usage objects included does not meet the preset quantity threshold, the first usage objects to be allocated in the target test group can be sequentially allocated to each experimental group until the number of already allocated usage objects included in each experimental group meets the preset quantity threshold.
[0254] Exemplarily, as Figure 13B shown, when the maximum difference between the numbers of already allocated usage objects included between each experimental group reaches the preset difference threshold, the first usage objects to be allocated in the target test group can be directly allocated to the experimental group with the least number of already allocated usage objects until the maximum difference does not exceed the preset difference threshold.
[0255] In this way, through the above method, it is ensured that the number of allocated users included in each experimental group within the test group is in a sufficient and uniform state, thereby ensuring the accuracy of the AB test results.
[0256] Optionally, for the allocation process of the experimental groups, it can also be based on the set similarity of the attribute feature sets of each first user within the experimental group.
[0257] Exemplarily, a first user can be allocated to one of the at least two experimental groups of the target test group, and this experimental group satisfies: the set similarity between each of the allocated users included and the attribute feature set of a first user is the largest within at least the experimental groups included in the target test group.
[0258] In this way, the set similarity between the allocated users included in each experimental group can be at a relatively high level, ensuring a relatively high degree of homogeneity among the allocated users within each experimental group, and also ensuring the accuracy of the AB test results.
[0259] Optionally, for the allocation process of the experimental groups, it can also be based on the change range of the set homogeneity degree corresponding to whether the first user is allocated among the experimental groups.
[0260] Exemplarily, a first user can be allocated to one of the at least two experimental groups of the target test group, and this experimental group satisfies: the change range of the set homogeneity degree between the attribute feature sets of the allocated users in an experimental group before and after allocating a first user to the experimental group conforms to a preset range condition.
[0261] Specifically, for each of the at least two experimental groups, the following operations can be performed respectively:
[0262] First, for the set similarity between the attribute feature sets of every two allocated users in an experimental group, perform a fusion process to obtain the corresponding first set homogeneity degree;
[0263] Secondly, take a first user as a new allocated object in an experimental group, and re-perform a fusion process on the set similarity between the attribute feature sets of every two allocated users in this experimental group to obtain the corresponding second set homogeneity degree.
[0264] In this way, through two operations, the set homogeneity degrees corresponding to before and after allocating a first user to an experimental group are obtained. Among them, the determination of the set homogeneity degree can also be determined by the vector sum or vector product of each feature vector between the attribute feature sets.
[0265] Finally, based on the homogeneity degree of the first set and the homogeneity degree of the second set, after determining a first usage object and assigning it to the one experimental group, the change range of the homogeneity degree of the set is determined.
[0266] After performing the above operations for each experimental group respectively, the change range of the homogeneity degree corresponding to each experimental group can be obtained. Therefore, based on these change ranges, the experimental group with the change range of the homogeneity degree not exceeding the range threshold can be selected as the experimental group to which a first usage object is assigned.
[0267] By allocating the first usage objects assigned to the target test group in the above manner, the difference degree between the experimental groups is ensured to be small, the homogeneity degree between the experimental groups is further improved, and thus the accuracy of the AB test result is improved.
[0268] As mentioned above, among the usage objects assigned to an experimental group, there may be first usage objects and second usage objects. Optionally, for the first usage objects and second usage objects included in the experimental group, the respective proportions they occupy can be determined by a preset proportional condition. This proportional condition can be set manually based on experience or obtained after model training. This application does not limit this.
[0269] The above introduces several possible implementation manners for traffic allocation of usage objects. After the usage objects are assigned to the corresponding experimental groups, the subsequent operations of the AB test can be performed based on these assigned usage objects.
[0270] In a possible implementation manner, after completing the allocation of the experimental group for the first usage object, the following operations can be continued:
[0271] After the first usage object is assigned to the target test group, when it is necessary to perform an AB test on the target test group, for each experimental group in the target test group, the following operations can be performed respectively:
[0272] First, based on the usage objects already assigned in an experimental group, a test version of the update function of the target service platform is tested to obtain corresponding test results; among them, there is a one-to-one correspondence between the test version of the update function and the experimental group, and the test results include: the feedback data of the usage objects already assigned for a test version.
[0273] When conducting an A / B test, for the target test group, it can be a test related to an updated function of the target service platform. In this test, it is necessary to test which test version among several test versions provided for the updated function is more popular or can obtain better feedback data. Therefore, in the target test group, at least two experimental groups can be set up so that there is a one-to-one correspondence between each experimental group and the test version of the updated function.
[0274] In this way, for each experimental group, the feedback data of each assigned user for a test version of the updated function can be obtained by using the assigned users included therein, and then used as the corresponding test result.
[0275] Secondly, when the feedback data included in the test result meets the set screening threshold, the test version corresponding to the test result can be selected and pushed to each first user included in the target test group.
[0276] In this way, because through the above steps, the degree of homogenization among the assigned users included in each experimental group is relatively high. Therefore, after screening out the test version corresponding to the updated function that meets the requirements (for example, the feedback data is relatively excellent, that is, the click-through rate or interaction rate is relatively high) through the feedback data of different experimental groups for different test versions, it can be pushed to each first user in the target test group, thereby improving the usage experience of the updated function in the overall target service platform for each first user in the target test group.
[0277] In this way, the purpose of an A / B test for the target test group can be achieved, that is: as Figure 14 shown, by conducting a small number of tests on the users in the non-reserved group in the target test group, the most suitable test version is selected and pushed to a large number of users in the target test group to maximize the usage experience of the overall users.
[0278] In a possible implementation manner, in addition to the subsequent related operations of the A / B test mentioned above, for each assigned user, each test group, and each experimental group allocated by using the traffic allocation method provided in this application, the following operations can also be continued:
[0279] Exemplarily, for the test time defined by the A / B test, the feedback data (such as clicks, views, shares, loading time, error rate, etc.) for the function to be tested during the A / B test can be obtained in different experimental groups, and then an efficient database or big data platform can be used to ensure the security and availability of the data. At the same time, the abnormal data in the feedback data is processed in a timely manner to ensure the smooth progress of the A / B test.
[0280] Thus, after the collection of feedback data is completed, statistical analysis can be performed on the feedback data of each test group to evaluate the corresponding effects of different versions of the functions to be tested. Specifically, key metrics (such as click-through rate, conversion rate, retention rate, etc.) can be calculated from the feedback data and then verified through statistical methods. Among them, the statistical methods can be any of the following methods: t-test, chi-square test, confidence interval calculation, significance judgment, multiple test correction, etc.
[0281] Next, based on these analysis results, the traffic allocation and version content can be adjusted to improve the overall business effect. Specifically, methods such as traffic adjustment, version optimization, and continuous iteration can be adopted. For example, traffic adjustment can be: if version B is significantly better than version A in a certain test group, the traffic proportion of version B can be increased; version optimization can be: according to the feedback data and data analysis, improve the version content to meet the needs of the users of a specific test group; continuous iteration can be: repeat AB testing, data analysis, and optimization to continuously improve the service quality and the user experience of the users.
[0282] The above introduces various possible implementation manners in the traffic allocation method provided in the embodiments of the present application. It should be understood that the above methods can be implemented through different combination manners. The following will introduce an example of this process through a combination manner.
[0283] As Figure 15 shown, when the traffic allocation method provided in the embodiments of the present application is adopted, for each first user among the first users who log in to the target service platform, the following operations can be performed:
[0284] First, based on the time decay mechanism, obtain the attribute feature set of the first user from the historical data; then, based on the set similarity between the attribute feature set and the reference feature sets of each test group, allocate the first user to the target test group whose set similarity reaches the second set threshold.
[0285] In the target test group, the first user is allocated twice. First, after it is determined that the first user is not allocated to the reserved group, the experimental group corresponding to the first user is allocated again.
[0286] In this process, by using the influence of the allocation of the first user to each experimental group on the homogeneity degree of the set of the experimental groups, an experimental group with the least influence is selected, and the first user is allocated to the target experimental group.
[0287] In this way, the traffic allocation process for the first usage object can be completed. After determining that the allocated objects in each experimental group meet the requirements, the test result analysis can be performed based on the feedback data of each experimental group, so as to optimize the updated function in the A / B test according to the test results.
[0288] Based on the same inventive concept, an embodiment of the present application provides a traffic allocation device, which can implement the functions corresponding to the foregoing traffic allocation method. Please refer to Figure 16 , the device includes an acquisition module 161, a processing module 162, and an allocation module 163, where:
[0289] The acquisition module 161: is used to obtain the attribute feature set of each first usage object respectively based on the historical data of each first usage object in the target service platform when each first usage object logs in to the target service platform;
[0290] The processing module 162: is used to perform the following operations for each first usage object respectively: determine the set similarity between each reference feature set and the attribute feature set of a first usage object respectively based on the reference feature sets of each test group preset; where each reference feature is: an attribute feature with a feature homogeneity degree reaching a first set threshold among the first usage objects included in a test group; each test group includes: at least two experimental groups for A / B testing;
[0291] The allocation module 163: is used to allocate a first usage object to an experimental group in a target test group whose set similarity reaches a second set threshold based on the set similarity of each obtained reference feature set.
[0292] Optionally, the processing module is further used to:
[0293] Obtain the attribute feature set of each second usage object respectively based on the historical data of each second usage object in the target service platform; where the second usage object is: other usage objects in the target service platform except each first usage object;
[0294] Perform feature clustering on each second usage object based on the set similarity between the attribute feature sets of each second usage object to obtain each test group; where in each test group, the set similarity between the attribute feature sets of any two second usage objects reaches a group threshold;
[0295] For each test group, perform the following operations respectively: obtain the reference feature set of a test group based on the feature homogeneity degree among the attribute features of the second usage objects in a test group.
[0296] Optionally, when the obtaining module is used to obtain a reference feature set of a test group based on the degree of feature homogeneity among the attribute features of each second user object in a test group, it is specifically used for:
[0297] In the attribute feature sets of each second user object included in a test group, filter out at least one candidate feature; wherein, each candidate feature is an attribute feature whose degree of difference from the attribute features of other test groups in each test group reaches a preset difference threshold;
[0298] For each candidate feature, respectively execute:
[0299] Based on a candidate feature, obtain the degree of feature homogeneity corresponding to a candidate feature based on the feature similarity between every two second user objects in a test group; wherein, the value of the degree of feature homogeneity is positively correlated with the value obtained after fusion processing of the obtained feature similarities;
[0300] When the degree of feature homogeneity corresponding to a candidate feature reaches the first set threshold, use a candidate feature as a reference feature of a test group.
[0301] Optionally, after obtaining each test group and before each first user object logs in to the target service platform, the processing module is further used for:
[0302] For each test group, respectively execute:
[0303] Respectively obtain the identity identifier of each second user object included in a test group in the target service platform;
[0304] Create at least two blank experimental groups in a test group;
[0305] Based on the obtained identity identifiers, randomly assign each second user object to one of the at least two experimental groups.
[0306] Optionally, when the obtaining module is used to obtain the attribute feature set of each first user object respectively based on the historical data of each first user object in the target service platform, it is specifically used for:
[0307] For each first user object, respectively execute:
[0308] Based on the historical data of a first user object in the target service platform, obtain the generation time of each interaction data generated after the interaction between the first user object and the target service platform;
[0309] Respectively obtain the time weights of the corresponding interaction data based on the obtained generation times; wherein, the time weight is negatively correlated with the duration between the generation time and the current moment;
[0310] Based on the obtained time weights for each, the interaction data for each is weighted and adjusted, and for the interaction data after weighted adjustment and other data in the historical data other than the interaction data for each, comprehensive feature extraction is performed to obtain an attribute feature set of a first usage object.
[0311] Optionally, when the acquisition module is used to obtain the attribute feature set of each first usage object respectively based on the historical data of each first usage object in the target service platform, it is specifically used for:
[0312] Based on a preset extraction time range, obtain the historical data of each first usage object from the target service platform; wherein, the generation time of each interaction data included in each historical data is within the extraction time range; the interaction data is: generated after the corresponding first usage object interacts with the target service platform.
[0313] For the historical data of each first usage object respectively, perform feature extraction to obtain the attribute feature set of each first usage object.
[0314] Optionally, the target test group further includes: a reserved group not used for A / B testing.
[0315] When the allocation module is used to allocate a first usage object to one of the experimental groups in a target test group where the set similarity of each obtained reference feature set reaches a second set threshold, it is specifically used for:
[0316] Based on the obtained set similarity of each reference feature set, allocate a first usage object to a target test group where the set similarity reaches a second set threshold.
[0317] For the reserved group, randomly allocate a first usage object based on the identity identifier of the first usage object in the target service platform.
[0318] When a first usage object is not randomly allocated to the reserved group, allocate the first usage object to one of the at least two experimental groups included in the target test group; wherein, one experimental group satisfies: the change range of the set homogeneity degree between the attribute feature sets of the allocated usage objects in one experimental group before and after allocating a first usage object conforms to a preset amplitude condition.
[0319] Optionally, when the allocation module is used to allocate a first usage object to one of the at least two experimental groups included in the target test group, it is specifically used for:
[0320] For each of the at least two experimental groups, respectively execute:
[0321] For the set similarity between the respective attribute feature sets of every two assigned usage objects in an experimental group, perform a fusion process to obtain the corresponding first set homogeneity degree;
[0322] Take a first usage object as a new assigned object in an experimental group, and re-perform a fusion process on the set similarity between the respective attribute feature sets of every two assigned usage objects in an experimental group to obtain the corresponding second set homogeneity degree;
[0323] Based on the first set homogeneity degree and the second set homogeneity degree, determine the change range of the set homogeneity degree after a first usage object is assigned to an experimental group;
[0324] Assign a first usage object to an experimental group where the change range of the set homogeneity degree does not exceed the range threshold.
[0325] Optionally, when the processing module is used to respectively determine the set similarity between each reference feature set and the attribute feature set of a first usage object based on the respective reference feature sets of preset test groups, it is specifically used for:
[0326] For each reference feature set in the respective reference feature sets of each test group, perform the following respectively:
[0327] Based on a preset dimension threshold, perform dimensionality reduction processing on a reference feature set to obtain a dimensionality-reduced reference feature set;
[0328] Based on the dimension threshold, perform dimensionality reduction processing on the associated attribute features associated with a reference feature set in the attribute feature set of a usage object to obtain the dimensionality-reduced associated attribute features;
[0329] Based on the dimensionality-reduced reference feature set and the dimensionality-reduced associated attribute features, obtain the set similarity between a reference feature set and a first usage object.
[0330] Optionally, when the processing module is used to respectively determine the set similarity between each reference feature set and the attribute feature set of a first usage object based on the respective reference feature sets of preset test groups, it is specifically used for:
[0331] For each reference feature set in the respective reference feature sets of each test group, perform the following respectively:
[0332] Obtain the associated attribute features associated with a reference feature set in the attribute feature set of a first usage object;
[0333] For each first usage object, determine the association similarity between the associated attribute features included in each other first usage object except one first usage object and the associated feature attributes of one first usage object;
[0334] Based on the obtained association similarities, the associated attribute features, and a reference feature set, determine the set similarity between a reference feature set and a first usage object; wherein, the value of the set similarity is positively correlated with the values of the association similarities.
[0335] Please refer to Figure 17 , which is a computer device 1700 provided by an embodiment of the present application. The computer device 1700 may be, for example, Figure 1 the client 102 or the server 101 in. The current version and historical versions of the data storage program and the application software corresponding to the data storage program may be installed on the computer device 1700. The computer device 1700 includes a processor 1780 and a memory 1720. In some embodiments, the computer device 1700 may include a display unit 1740. The display unit 1740 includes a display panel 1741 for displaying a user interaction operation interface, etc.
[0336] In a possible embodiment, the display panel 1741 may be configured in the form of a liquid crystal display (LCD) or an organic light-emitting diode (OLED), etc.
[0337] The processor 1780 is configured to read a computer program and then execute the method defined by the computer program. For example, the processor 1780 reads the data storage program or file, etc., so as to run the data storage program on the computer device 1700 and display a corresponding interface on the display unit 1740. The processor 1780 may include one or more general-purpose processors and may also include one or more DSPs (Digital Signal Processors) for performing related operations to implement the technical solutions provided by the embodiments of the present application.
[0338] The memory 1720 generally includes internal memory and external memory. The internal memory can be a random access memory (RAM), a read-only memory (ROM), a cache (CACHE), etc. The external memory can be a hard disk, an optical disc, a USB drive, a floppy disk, or a tape drive, etc. The memory 1720 is used to store computer programs and other data. The computer programs include application programs corresponding to each client, etc. The other data may include an operating system or data generated after the application programs are run. The data includes system data (such as configuration parameters of the operating system) and user data. In the embodiments of the present application, the computer programs are stored in the memory 1720, and the processor 1780 executes the computer programs in the memory 1720 to implement any one of the methods discussed in the previous figures.
[0339] As an embodiment, when each first user object logs in to the target service platform, the processor 1780 respectively obtains the attribute feature set of each first user object based on the historical data of each first user object in the target service platform. The processor 1780 stores the obtained attribute feature set in the memory 1720.
[0340] For each first user object, the processor 1780 respectively executes: based on the reference feature sets of each preset test group, respectively determining the set similarity between each reference feature set and the attribute feature set of a first user object; wherein, each reference feature is: an attribute feature with a feature homogeneity degree reaching a first set threshold among the first user objects included in a test group; each test group includes: at least two experimental groups for A / B testing.
[0341] Based on the set similarities of the obtained reference feature sets, the processor 1780 assigns a first user object to one of the experimental groups in the target test group where the set similarity reaches a second set threshold.
[0342] The above display unit 1740 is used to receive input digital information, character information, or contact touch operations / non-contact gestures, and generate signal inputs related to the user settings and function controls of the computer device 1700, etc. Specifically, in the embodiments of the present application, the display unit 1740 may include a display panel 1741. The display panel 1741 is, for example, a touch screen, which can collect touch operations of the user on or near it (such as the user using a finger, a stylus, or any suitable object or accessory to operate on the display panel 1741 or near the display panel 1741), and drive the corresponding connection device according to a preset program.
[0343] In a possible embodiment, the display panel 1741 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1780, and can receive and execute the commands sent by the processor 1780.
[0344] Among them, the display panel 1741 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the display unit 1740, in some embodiments, the computer device 1700 may further include an input unit 1730. The input unit 1730 may include an image input device 1731 and other input devices 1732. The other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), trackball, mouse, joystick, etc.
[0345] In addition to the above, the computer device 1700 may further include a power supply 1790 for powering other modules, an audio circuit 1760, a near field communication module 1770, and an RF circuit 1710. The computer device 1700 may further include one or more sensors 1750, such as an acceleration sensor, a light sensor, a pressure sensor, etc. The audio circuit 1760 specifically includes a speaker 1761 and a microphone 1762, etc. For example, the computer device 1700 can collect the user's voice through the microphone 1762 and perform corresponding operations, etc.
[0346] As an embodiment, the number of processors 1780 may be one or more. The processor 1780 and the memory 1720 may be coupled or relatively independent.
[0347] As an embodiment, Figure 17 the processor 1780 in Figure 16 may be used to implement the functions of the acquisition module 161, the processing module 162, and the distribution module 163 in
[0348] As an embodiment, Figure 17 the processor 1780 in
[0349] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by a computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the computer program is executed, it performs the steps including those of the above method embodiments. The aforementioned storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0350] Alternatively, if the above integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present invention, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. For example, it is embodied through a computer program product. The computer program product is stored in a storage medium and includes a computer program for causing a computer device to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0351] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.
Claims
1. A flow distribution method, characterized in that, Applied to a target service platform, including: When each first user object logs in to the target service platform, based on the historical data of each first user object in the target service platform, obtain the attribute feature set of each first user object respectively; Among them, for each of the first user objects, the following operations are performed respectively: Based on the reference feature sets of each preset test group, determine the set similarity between each reference feature set and the attribute feature set of a first user object respectively; among them, each reference feature is: an attribute feature with a feature homogeneity degree reaching a first set threshold among the first user objects included in a test group; each of the test groups includes: at least two experimental groups for A / B testing; Based on the set similarities of the obtained reference feature sets respectively, assign the first user object to an experimental group in a target test group whose set similarity reaches a second set threshold.
2. The method according to claim 1, wherein The test groups, and the reference feature sets of the test groups respectively, are obtained in the following manner: Based on the historical data of each second user object in the target service platform, obtain the attribute feature set of each second user object respectively; among them, the second user object is: other user objects in the target service platform except the first user objects; Based on the set similarities between the attribute feature sets of the second user objects, perform feature clustering on the second user objects to obtain the test groups; among them, in each test group, the set similarity between the attribute feature sets of any two second user objects reaches a group threshold; For each of the test groups, perform the following operations respectively: Based on the feature homogeneity degree among the attribute features of the second user objects in a test group, obtain the reference feature set of the test group.
3. The method according to claim 2, wherein The operation of obtaining the reference feature set of a test group based on the feature homogeneity degree among the attribute features of the second user objects in the test group includes: In the attribute feature sets of the second user objects included in the test group, screen out at least one candidate feature; among them, each candidate feature is: an attribute feature whose difference degree from the attribute features of other test groups in the test groups reaches a preset difference threshold; For each of the candidate features, perform the following operations respectively: Based on a candidate feature, obtain the feature homogeneity degree corresponding to the candidate feature based on the feature similarity between every two second user objects in the test group; among them, the value of the feature homogeneity degree is positively correlated with the value obtained after the fusion processing of the obtained feature similarities; When the feature homogeneity degree corresponding to a candidate feature reaches the first set threshold, use the candidate feature as a reference feature of the test group.
4. The method according to claim 2, wherein After obtaining the test groups, and before the first user objects log in to the target service platform, at least two experimental groups in each of the test groups are obtained in the following manner: For each of the test groups, perform the following operations respectively: Obtain the identity identifiers of each second user object included in a test group in the target service platform respectively; Create at least two blank experimental groups in the test group; Based on the obtained identity identifiers, randomly assign each second user object to one of the at least two experimental groups.
5. The method according to any one of claims 1 to 4, characterized in that Based on the historical data of each first user object in the target service platform respectively, obtain the attribute feature set of each first user object, including: For each first user object, perform respectively: Based on the historical data of a first user object in the target service platform, obtain the generation time of each interaction data generated after the interaction between the first user object and the target service platform; Based on the obtained generation times respectively, obtain the time weights of the corresponding interaction data; wherein, the time weight is negatively correlated with the duration between the generation time and the current moment; Based on the obtained time weights, perform weighted adjustment on the interaction data, and perform comprehensive feature extraction on the weighted-adjusted interaction data and other data in the historical data except the interaction data, to obtain the attribute feature set of the first user object.
6. The method according to any one of claims 1 to 4, characterized in that Based on the historical data of each first user object in the target service platform respectively, obtain the attribute feature set of each first user object, including: Based on a preset extraction time range, obtain the historical data of each first user object from the target service platform; wherein, the generation time of each interaction data included in each historical data is within the extraction time range; the interaction data are those generated after the corresponding first user object interacts with the target service platform; Perform feature extraction on the historical data of each first user object respectively, to obtain the attribute feature set of each first user object.
7. The method according to any one of claims 1 to 4, characterized in that The target test group further includes: a reserved group not used for AB testing; based on the set similarity of each obtained reference feature set, assigning a first user object to one of the experimental groups in the target test group whose set similarity reaches a second set threshold includes: Based on the set similarity of each obtained reference feature set, assign a first user object to a target test group whose set similarity reaches a second set threshold; For the reserved group, randomly assign the first user object based on the identity identifier of the first user object in the target service platform; When the first user object is not randomly assigned to the reserved group, assign the first user object to one of the at least two experimental groups included in the target test group; wherein, for one experimental group, the change range of the set homogeneity degree between the attribute feature sets of the assigned user objects in the experimental group before and after assigning the first user object to the experimental group meets a preset amplitude condition.
8. The method according to claim 7, wherein Assigning the one first usage object to one of at least two experimental groups included in the target test group includes: For each of the at least two experimental groups, respectively perform: Perform a fusion process on the set similarity between the attribute feature sets of every two assigned usage objects in one experimental group to obtain the corresponding first set homogeneity degree; Take the one first usage object as a new assigned object in the one experimental group, and re-perform a fusion process on the set similarity between the attribute feature sets of every two assigned usage objects in the one experimental group to obtain the corresponding second set homogeneity degree; Based on the first set homogeneity degree and the second set homogeneity degree, determine the change range of the set homogeneity degree after the one first usage object is assigned to the one experimental group; Assign the one first usage object to one of the experimental groups where the change range of the set homogeneity degree does not exceed the range threshold.
9. The method according to any one of claims 1 to 4, characterized in that The determining, based on the reference feature sets of each of the preset test groups respectively, the set similarity between each reference feature set and the attribute feature set of a first usage object includes: For each reference feature set in the reference feature sets of each of the test groups, respectively perform: Based on a preset dimension threshold, perform a dimensionality reduction process on a reference feature set to obtain a dimension-reduced reference feature set; Based on the dimension threshold, perform a dimensionality reduction process on the associated attribute features associated with the reference feature set in the attribute feature set of the one usage object to obtain the dimension-reduced associated attribute features; Based on the dimension-reduced reference feature set and the dimension-reduced associated attribute features, obtain the set similarity between the reference feature set and the first usage object.
10. The method according to any one of claims 1-4, characterized in that, The determining, based on the reference feature sets of each of the preset test groups respectively, the set similarity between each reference feature set and the attribute feature set of a first usage object includes: For each reference feature set in the reference feature sets of each of the test groups, respectively perform: Obtain the associated attribute features associated with a reference feature set in the attribute feature set of the first usage object; Respectively determine the association similarity between the associated attribute features included in each other first usage object except the first usage object among the first usage objects and the associated feature attributes of the first usage object; Based on the obtained association similarities, the associated attribute features, and the reference feature set, determine the set similarity between the reference feature set and the first usage object; wherein, the value of the set similarity is positively correlated with the values of the association similarities.
11. The method according to any one of claims 1 to 4, characterized in that, After assigning the first usage object to one of the experimental groups in the target test group where the set similarity reaches a second set threshold based on the obtained set similarities of each reference feature set, the method further includes: For each experimental group in the target test group, respectively perform: Based on each assigned user in an experimental group, test a test version of the update function for the target service platform to obtain corresponding test results; wherein, there is a one-to-one correspondence between the test version of the update function and the experimental group; the test results include: feedback data of each assigned user for the test version. When the feedback data included in the test results meets the set screening threshold, push the test version to each first user included in the target test group.
12. A flow distribution device, characterized in that, Applied to the target service platform, including: An acquisition module: used to, when each first user logs in to the target service platform, respectively obtain the attribute feature set of each first user based on the historical data of each first user in the target service platform. A processing module: used to, for each first user, respectively execute: based on the reference feature sets of each preset test group, respectively determine the set similarity between each reference feature set and the attribute feature set of a first user; wherein, each reference feature is: an attribute feature with a feature homogeneity degree reaching a first set threshold among the first users included in a test group; each test group includes: at least two experimental groups for A / B testing. An allocation module: used to, based on the set similarity of each obtained reference feature set, allocate the first user to an experimental group in the target test group where the set similarity reaches a second set threshold.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 11.
14. A computer device, characterized in that, Including: A memory, used to store the computer program; A processor, used to call the computer program stored in the memory and execute the method according to any one of claims 1 to 11 according to the obtained computer program.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is used to cause a computer to execute the method according to any one of claims 1 to 11.