Video recommendation method and system based on AB experiment, electronic equipment and storage medium

By using a pre-trained shunt model in video recommendation, the user preference score and model winning rate are calculated based on user data and experimental data, and the shunt ratio is determined to display the video recommendation results. This solves the problems of long effect recovery period, large fluctuations and influenced by various factors in the existing AB experimental methods, achieving a more stable and efficient video recommendation effect.

CN120144822APending Publication Date: 2025-06-13HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510279065.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the video recommendation, the existing AB experimental methods have problems such as long effect recovery cycle, large effect fluctuations, and are affected by time period, user population and user interest changes.

Method used

By pre-training the trained shunt model using historical user data and historical experimental data, each shunt model is obtained, and then when receiving a user request, the estimated value of the experimental indicators is estimated based on the user data and experimental data, the user preference score and model winning rate are calculated, and the shunt ratio is determined to display the video recommendation results.

Benefits of technology

The effect recovery cycle is shortened, the effect fluctuations are avoided, and the effects of different shunt models are affected by the time period, user population and user interest changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144822A_ABST
    Figure CN120144822A_ABST
Patent Text Reader

Abstract

The invention provides a video recommendation method and system based on an AB experiment, electronic equipment and a storage medium. The method comprises the steps of obtaining user data and experiment data of a user based on an online recommendation request sent by the user; loading each pre-trained shunting model, pre-estimating an estimated value of each experimental index according to the user data and the experimental data through the shunting model, and determining a user preference score of a user for a video recommendation result of the shunting model according to the estimated value of each experimental index; for each shunting model, counting the selection times of the video recommendation result of the shunting model, and determining the model winning rate of the shunting model according to the selection times; calculating a video recommendation score of each shunting model according to the user preference score and the model winning rate of each shunting model, and determining a shunting proportion among the shunting models according to the video recommendation score of each shunting model; and displaying the current video recommendation result of each shunting model to the user according to the shunting proportion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more specifically, to a video recommendation method, system, electronic device and storage medium based on AB testing. Background Art

[0002] With the continuous development of the video recommendation field, in order to improve the user experience, it is necessary to continuously conduct experimental iterations to select the best recommendation model for video recommendation. The current method for selecting a recommendation model is through AB testing, that is, evenly splitting the online traffic into two different experimental buckets, and trying to ensure that the user distribution in each bucket is uniform. Among them, different algorithm models are deployed in each bucket; after observing and statistically analyzing the experimental effects for a period of time, the algorithm model deployed on the experimental bucket with better effects is rolled out to the full scale.

[0003] Although this method can screen out a recommendation model with better effects, it requires long-term observation and statistics, the effect recovery period is relatively long, and there are certain fluctuations in the effects; moreover, different recommendation models have different focuses, resulting in the effects obtained by different recommendation models being affected by factors such as time period, user group, and changes in user preferences. Summary of the Invention

[0004] In view of this, the present invention provides a video recommendation method, system, electronic device and storage medium based on AB testing to shorten the effect recovery period, avoid fluctuations in effects, and avoid the problem that the effects of different split models are affected by factors such as time period, user group, and changes in user preferences.

[0005] The first aspect of the present application provides a video recommendation method based on AB testing, and the method includes:

[0006] When receiving an online recommendation request sent by a user, obtaining the user data and experimental data of the user based on the online recommendation request;

[0007] Loading each pre-trained split model, and using the split model to estimate the predicted value of each experimental index according to the user data and the experimental data, and determining the user preference score of the video recommendation result of the user for the split model according to the predicted values of each experimental index; wherein, the split model is trained by using historical user data and historical experimental data to train the split model to be trained;

[0008] For each split model, counting the number of times the video recommendation result of the split model is selected, and determining the model winning rate of the split model according to the number of times selected;

[0009] Calculate the video recommendation score for each of the shunt models according to the user preference score and the model winning rate of each shunt model, and determine the shunt ratio between the shunt models according to the video recommendation scores of the shunt models;

[0010] Display the current video recommendation results of each shunt model to the user according to the shunt ratio.

[0011] Optionally, loading the pre-trained shunt models, and estimating the predicted values of each experimental metric by the shunt models according to the user data and the experimental data, and determining the user preference score of the user for the video recommendation results of the shunt models includes:

[0012] Load the pre-trained shunt models, and input the user data and the experimental data into each of the shunt models; wherein, the shunt model includes a sequence encoding module, a hidden layer, a feature splicing layer and a multi-layer activation layer;

[0013] Convert each exposure video into a plurality of video feature vectors through the sequence encoding module, and generate a one-dimensional video vector according to the respective video feature vectors of each exposure video; wherein, the user data includes user exposure data and user basic attribute data, and the user exposure data includes a plurality of exposure videos sorted by exposure time;

[0014] Smooth the user basic attribute data, context data and experimental effect statistical data through the hidden layer, and splice the video vector and the smoothed user basic attribute data, context data and experimental effect statistical data through the feature splicing layer to obtain a video splicing feature;

[0015] Process the video splicing feature through the multi-layer activation layer to obtain a target video splicing feature;

[0016] Call the normalization metric function to estimate the predicted value of each experimental metric according to the target video splicing feature, and determine the user preference score of the user for the video recommendation results of the shunt model according to the predicted values of the experimental metrics.

[0017] Optionally, the converting each exposure video into a plurality of video feature vectors through the sequence encoding module, and generating a one-dimensional video vector according to the respective video feature vectors of each exposure video includes:

[0018] Convert each of the exposure videos into a plurality of video feature vectors through the sequence encoding module, and perform weighted processing on each of the video feature vectors to obtain a plurality of target video feature vectors corresponding to each exposure video;

[0019] The sequence encoding module sequentially performs pooling processing on the respective target video feature vectors of the exposure videos except the first exposure video and the activation data of its previous exposure video, and uses the activation unit to activate the obtained processing result to obtain activation data until the respective target video feature vectors of the last exposure video and the activation data of its previous exposure video are subjected to pooling processing to obtain a video vector; wherein, the activation data of the first exposure video is obtained by activating the respective target video feature vectors of the first exposure video using the activation unit.

[0020] Optionally, the calling of the normalization metric function estimates the predicted values of each experimental metric according to the target video splicing feature, and determines the user preference score of the video recommendation result of the user for the shunt model according to the predicted values of each experimental metric, including:

[0021] Call the normalization metric function to estimate the predicted value of each experimental metric according to the target video splicing feature, calculate the preference score of each experimental metric according to the predicted value of the experimental metric and its weight, and calculate the user preference score of the video recommendation result of the user for the shunt model according to the preference scores of each experimental metric.

[0022] Optionally, training the shunt model using the historical user data and historical experimental data to obtain the shunt model includes:

[0023] Obtain the historical user data, historical experimental data, and the actual predicted values of each experimental metric of the video recommendation results of the historical users for each shunt model to be trained;

[0024] Input the historical user data, historical experimental data, and the actual predicted values of each experimental metric of the video recommendation results of the historical users for each shunt model to be trained into each shunt model to be trained;

[0025] Generate a one-dimensional historical video vector through the sequence encoding module to be trained according to the user exposure data of the user data;

[0026] Smooth the historical user basic attribute data in the historical user data and the historical experimental data through the hidden layer to be trained, and splice the historical video vector and the smoothed historical user basic attribute data and historical experimental data through the feature splicing layer to be trained to obtain a historical video splicing feature;

[0027] Process the historical video splicing feature through the multi-layer activation layer to be trained to obtain a historical target video splicing feature;

[0028] Call the softmax function to estimate the predicted pre-values of each experimental metric based on the historical target video splicing features, and adjust the parameters of the to-be-trained shunt model with the predicted pre-values of each experimental metric approaching the corresponding actual pre-values as the training objective until the to-be-trained shunt model converges, obtaining the corresponding shunt model.

[0029] The second aspect of the present application provides a video recommendation system based on an AB test. The system includes:

[0030] A first acquisition unit, configured to, when receiving an online recommendation request sent by a user, acquire the user data and experimental data of the user based on the online recommendation request;

[0031] A user preference score determination unit, configured to load each pre-trained shunt model, estimate the pre-values of each experimental metric through the shunt model according to the user data and the experimental data, and determine the user preference score of the user for the video recommendation result of the shunt model according to the pre-values of each experimental metric; wherein, the shunt model is obtained by a training unit training a to-be-trained shunt model using historical user data and historical experimental data;

[0032] A model win rate determination unit, configured to, for each shunt model, count the number of times the video recommendation result of the shunt model is selected, and determine the model win rate of the shunt model according to the number of times selected;

[0033] A shunt ratio determination unit, configured to calculate the video recommendation score of each shunt model according to the user preference score and the model win rate of each shunt model, and determine the shunt ratio between each shunt model according to the video recommendation scores of each shunt model;

[0034] A video recommendation unit, configured to display the current video recommendation results of each shunt model to the user according to the shunt ratio.

[0035] Optionally, the user preference score determination unit includes:

[0036] A first input unit, configured to load each pre-trained shunt model, and input the user data and the experimental data into each shunt model; wherein, the shunt model includes a sequence encoding module, a hidden layer, a feature splicing layer, and a multi-layer activation layer;

[0037] A video vector generation unit, configured to convert each exposure video into multiple video feature vectors through the sequence encoding module, and generate a one-dimensional video vector according to the respective video feature vectors of each exposure video; wherein, the user data includes user exposure data and user basic attribute data, and the user exposure data includes multiple exposure videos sorted by exposure time.

[0038] A first feature splicing unit, configured to perform smoothing processing on the user basic attribute data, context data, and experimental effect statistical data through the hidden layer, and splice the video vector and the smoothed user basic attribute data, context data, and experimental effect statistical data through the feature splicing layer to obtain a video splicing feature.

[0039] A second processing unit, configured to process the video splicing feature through the multi-layer activation layer to obtain a target video splicing feature.

[0040] An estimation unit, configured to call a normalization metric function to estimate the predicted value of each experimental metric according to the target video splicing feature, and determine the user preference score of the video recommendation result of the user for the shunt model according to the predicted values of the respective experimental metrics.

[0041] Optionally, the video vector generation unit includes:

[0042] A weighted processing unit, configured to convert each exposure video into multiple video feature vectors through the sequence encoding module, and perform weighted processing on each video feature vector to obtain multiple target video feature vectors corresponding to each exposure video.

[0043] A video vector generation subunit, configured to sequentially perform pooling processing on the respective target video feature vectors of the exposure videos except the first exposure video and the activation data of its previous exposure video through the sequence encoding module, and activate the obtained processing result by using an activation unit to obtain activation data, until the respective target video feature vectors of the last exposure video and the activation data of its previous exposure video are subjected to pooling processing to obtain a video vector; wherein, the activation data of the first exposure video is obtained by activating the respective target video feature vectors of the first exposure video by using the activation unit.

[0044] A third aspect of the present application provides an electronic device, including: a processor and a memory, the processor and the memory are connected through a communication bus; wherein, the processor is configured to call and execute a program stored in the memory; the memory is configured to store a program, and the program is used to implement the video recommendation method based on the AB experiment provided in the first aspect of the present application.

[0045] The fourth aspect of this application provides a storage medium, in which computer-executable instructions are stored, and the computer-executable instructions are used to execute the video recommendation method based on the AB test provided in the first aspect of this application.

[0046] The embodiments of this application provide a video recommendation method, system, electronic device and storage medium based on the AB test. It is possible to pre-train each shunt model to be trained by using historical user data and historical experiment data to obtain each shunt model. When receiving an online recommendation request sent by a user, based on the online recommendation request, obtain the user data and experiment data of the user. Each shunt model estimates the predicted value of each experimental index according to the user data and experiment data, and determines the user preference score of the video recommendation result of the user for each shunt model according to the predicted values of each experimental index. At the same time, it is also possible to calculate the model winning rate of each shunt model, so as to calculate the video recommendation score of each shunt model according to the user preference score and model winning rate of each shunt model. Finally, it is possible to determine the shunt ratio between each shunt model according to the video recommendation scores of each shunt model, and display the current video recommendation results of each shunt model to the user according to the obtained shunt ratio. It can be seen that the technical solution provided in this application does not need to ensure that the user distribution of each shunt model is uniform, nor does it need to conduct long-term experimental effect observation and statistics. It can not only shorten the effect recovery period, avoid fluctuations in the effect, but also avoid the influence of factors such as time period, user population and user interest changes on the effects of different shunt models. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0048] Figure 1 It is a schematic flowchart of a video recommendation method based on the AB test provided in the embodiments of this application;

[0049] Figure 2 It is a schematic structural diagram of an AB test platform provided in the embodiments of this application;

[0050] Figure 3 It is a schematic flowchart of a training method for a shunt model provided in the embodiments of this application;

[0051] Figure 4 It is a schematic structural diagram of a shunt model to be trained provided in the embodiments of this application;

[0052] Figure 5 An example diagram for generating a historical video vector provided by an embodiment of the present application;

[0053] Figure 6 A schematic structural diagram of a video recommendation system based on an AB test provided by an embodiment of the present invention;

[0054] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] In the present application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0057] As can be seen from the above background technology, currently, it is possible to conduct an AB test, that is, evenly split the online traffic into two different experimental buckets. After observing and statistically analyzing the experimental effects for a period of time, the algorithm model deployed on the experimental bucket with better effects is rolled out in full. Among them, common AB test methods may include random splitting (randomly assigning users to the experimental group and the control group according to a random coefficient), hash-based splitting (performing a hash operation based on a certain unique identifier of the user (such as the user ID), and assigning users to different groups according to the range of the hash value), stratified sampling splitting, etc.

[0058] However, the current AB testing method faces complex and variable factors such as scenarios, user groups, and time periods, resulting in significant differences in the effects of different recommendation algorithms under different conditions. Moreover, the current AB testing method can only select the recommendation algorithm with better overall effect, unable to ensure that the adopted recommendation algorithm is the optimal algorithm under various factor conditions, and there is a risk of incomplete uniform grouping, and the statistical effect fluctuates and has deviations.

[0059] The embodiments of the present application provide a video recommendation method, system, electronic device, and storage medium based on AB testing. It can pre-train each shunt model to be trained by using historical user data and historical experiment data to obtain each shunt model. So when receiving an online recommendation request sent by a user, based on the online recommendation request, obtain the user data and experiment data of the user. Each shunt model estimates the predicted value of each experimental index according to the user data and experiment data, and determines the user preference score of the video recommendation result of the user for each shunt model according to the predicted values of each experimental index. At the same time, it can also calculate the model winning rate of each shunt model, so as to calculate the video recommendation score of each shunt model according to the user preference score and model winning rate of each shunt model. Finally, it can determine the shunt ratio between each shunt model according to the video recommendation scores of each shunt model, and display the current video recommendation results of each shunt model to the user according to the obtained shunt ratio, and can dynamically adjust the display ratio of the video recommendation results of each shunt model according to the actual situation. It can solve the problem that the current AB testing method faces complex and variable factors such as scenarios, user groups, and time periods, resulting in significant differences in the effects of different recommendation algorithms under different conditions. And it can also solve the problem that the current AB testing method can only select the recommendation algorithm with better overall effect, unable to ensure that the adopted recommendation algorithm is the optimal algorithm under various factor conditions, and there is a risk of incomplete uniform grouping, and the statistical effect fluctuates and has deviations. The technical solution provided by the present application can also achieve that different users are not shunted to different experiments in the way of user ID hashing bucket or random bucket, but display the video recommendation results of all shunt models to the same user. Therefore, the present application does not require functions such as vertical orthogonal traffic division, and only needs to implement in what proportion the video recommendation results of different shunt models are displayed to the user.

[0060] See Figure 1 , which shows a schematic flowchart of a video recommendation method based on AB testing provided by an embodiment of the present application. The process of this video recommendation method based on AB testing specifically includes the following steps:

[0061] S101: When receiving an online recommendation request sent by a user, obtain the user data and experiment data of the user based on the online recommendation request.

[0062] In the embodiments of the present application, a corresponding AB video platform can be pre-built. The AB video platform includes a data layer, a service layer, and a display layer. As Figure 2 shown, the data layer is used to store the current user data and historical user data of each user, as well as the experimental configuration data corresponding to each shunt model, the current experimental effect statistical data, and the historical experimental effect statistical data; the service layer is used to detect and receive online recommendation requests sent by users in real time, perform traffic distribution and data aggregation, and can also implement parameter settings (experimental configuration) for each shunt model to be trained in the experimental configuration data of each shunt model; the display layer is used to provide corresponding operation page configurations and provide visualization effect reports (report visualization) aggregated in various dimensions.

[0063] It should be noted that the user data may include user exposure data and user basic attribute data. Among them, the user exposure data may include multiple exposure videos sorted by exposure time, and the user basic attribute data may include user nicknames, user levels (ordinary users, VIP users, SVIP users, etc.); the historical user data may include historical user exposure data and historical user basic attribute data. The historical user exposure data may include multiple historical exposure videos sorted by exposure time, and the historical user basic attribute data may include historical user nicknames, historical user levels, etc.; the experimental effect statistical data may include statistical data such as click-through rates, conversion rates, and viewing durations of exposure videos in each experimental cross dimension (for example, the click-through rate of a certain variety show video A among group B people, the viewing duration of a certain TV drama video D during a certain time period C, etc.); the experimental configuration data includes data such as experimental parameters related to each shunt model.

[0064] It should also be noted that the multiple exposure videos are composed of the recommended videos in the video recommendation results of each shunt model.

[0065] It should also be noted that the experimental cross dimension can be video and population cross, video and time period cross, etc., which are not limited in the embodiments of this application.

[0066] In the embodiments of the present application, it can be detected in real time through the service layer whether there is an online recommendation request sent by a user. When it is detected that a user sends an online recommendation request, the service layer can receive the online recommendation request and obtain the current user data and experimental data of the user from the data layer based on the user unique identifier ID carried in the online recommendation request. Among them, the user data includes the user exposure data and user basic attribute data of the user; the experimental data includes context data and experimental effect statistical data.

[0067] It should be noted that the context data includes the video information of the previous exposure video and the next exposure video of each exposure video.

[0068] S102: Load each pre-trained shunt model, and use the shunt model to estimate the predicted values of each experimental metric based on the user data and experimental data, and determine the user preference score of the video recommendation result for the shunt model according to the predicted values of each experimental metric; among them, the shunt model is trained by using historical user data and historical experimental data to train the shunt model.

[0069] In the embodiment of the present application, for each shunt model to be trained, the shunt model to be trained can be pre-trained by using historical user data and historical experimental data to obtain the corresponding shunt model, and the obtained shunt models are stored in the AB test platform, so that after receiving an online recommendation request and obtaining the corresponding user data and experimental data, the service layer is called to load each pre-stored shunt model, and at the same time, the user data and experimental data are input into each shunt model, so that each shunt model estimates the predicted values of each experimental metric according to the input user data and experimental data, and determines the user preference score of the video recommendation result for the shunt model according to the predicted values of each experimental metric.

[0070] It should be noted that each experimental metric can be set in advance, where each experimental metric can include the click-through rate, conversion rate, viewing duration, etc. of the video recommendation result for the shunt model by the user. For example, the predicted values of each experimental metric can be that the click-through rate of the video recommendation result for the shunt model by the user is 80%, the conversion rate is 70%, and the viewing duration is two hours. Here, the embodiment of the present application does not make any limitations.

[0071] See Figure 3 , which shows a schematic flowchart of a method for training a shunt model provided by an embodiment of the present application. The method specifically includes the following steps:

[0072] S301: Obtain historical user data, historical experimental data, and the actual predicted values of each experimental metric of the historical user's video recommendation result for each shunt model to be trained;

[0073] In the process of specifically executing step S301, the service layer can be called to obtain log data from the data layer, and data cleaning is performed on the obtained log data to extract historical user data, historical experiment data, and actual estimated values of various experimental indicators of the video recommendation results of historical users for each shunt model to be trained. Finally, the historical user data, historical experiment data, and actual estimated values of various experimental indicators of the video recommendation results of historical users for each shunt model to be trained are input into each shunt model to be trained; among them, the historical user data includes historical user exposure data and historical user basic attribute data, and the historical experiment data includes historical context data, historical experiment effect statistical data, and experimental configuration data of each shunt model to be trained.

[0074] It should be noted that the historical user exposure data can include multiple historical exposure videos sorted by exposure time, and the historical user exposure data can include multiple historical exposure videos sorted by exposure time. The historical user basic attribute data can include historical user nicknames, historical user levels, etc.; the historical experiment effect statistical data can include statistical data such as click-through rate, conversion rate, and viewing duration of historical exposure videos in each experimental cross-dimension; the experimental configuration data includes experimental parameter data related to each shunt model to be trained.

[0075] S302: Input the historical user data, historical experiment data, and actual estimated values of various experimental indicators of the video recommendation results of historical users for each shunt model to be trained into each shunt model to be trained.

[0076] In the embodiment of the present application, the shunt model to be trained includes a sequence encoding module to be trained, a hidden layer to be trained, a feature splicing layer to be trained, a multi-layer activation layer to be trained, and a softmax function, as Figure 4 shown.

[0077] In the process of specifically executing step S302, for each shunt model to be trained, the historical user data, historical context data, historical experiment effect statistical data, and experimental configuration data of this shunt model to be trained can be input into this shunt model to be trained.

[0078] S303: Generate a one-dimensional historical video vector according to the user exposure data of the user data through the sequence encoding module to be trained.

[0079] In the embodiment of the present application, each historical exposure video is converted into a historical video feature vector corresponding to each experimental metric by a sequence encoding module to be trained, and the historical video feature vector corresponding to each experimental metric is weighted according to the weight corresponding to each experimental metric to obtain a historical target video feature vector corresponding to each experimental metric; starting from the second historical exposure video, the respective historical target video feature vectors of the historical exposure video are sequentially subjected to pooling processing with the historical activation data of the previous historical exposure video, and the obtained processing result is activated by the activation unit of the sequence encoding module to be trained to obtain historical activation data until the respective historical target video feature vectors of the last historical exposure video are subjected to pooling processing with the historical activation data of the previous historical exposure video to obtain a historical video vector; wherein, the historical activation data of the first historical exposure video is obtained by activating the respective historical target video feature vectors of the first historical exposure video by the activation unit of the sequence encoding module to be trained.

[0080] It should be noted that the respective experimental metrics may include click-through rate, conversion rate, and viewing duration, and the corresponding weights of the respective experimental metrics may include the weight Wa of the click-through rate, the weight Wb of the conversion rate, and the weight Wc of the viewing duration. The weight corresponding to each experimental metric can be set according to actual applications, and the embodiments of the present application do not limit this here.

[0081] For example, refer to Figure 5, The historical user-exposed videos include n historical exposure videos sorted by exposure videos. Each historical exposure video can be converted into a historical video feature vector corresponding to click-through rate, conversion rate, and viewing duration through a sequence encoding module to be trained; the historical video feature vector of the click-through rate of the historical exposure video V1 is weighted and calculated according to the preset weight Wa1 of the click-through rate to obtain the historical target video feature vector V1,a. The historical video feature vector of the conversion rate of the historical exposure video V1 is weighted and calculated according to the preset weight Wb1 of the conversion rate to obtain the historical target video feature vector V1,b. The historical video feature vector of the viewing duration of the historical exposure video V1 is weighted and calculated according to the preset weight Wc1 of the viewing duration to obtain the historical target video feature vector V1,c, that is, a three-dimensional vector (V1,a, V1,b, V1,c) is obtained;.....; the historical video feature vector of the click-through rate of the historical exposure video Vn is weighted and calculated according to the preset weight Wan of the click-through rate to obtain the historical target video feature vector Vn,a. The historical video feature vector of the conversion rate of the historical exposure video Vn is weighted and calculated according to the preset weight Wbn of the conversion rate to obtain the historical target video feature vector Vn,b. The historical video feature vector of the viewing duration of the historical exposure video Vn is weighted and calculated according to the preset weight Wcn of the viewing duration to obtain the historical target video feature vector Vn,c.

[0082] The activation unit of the sequence encoding module to be trained activates each historical target video feature vector (V1,a, V1,b, V1,c) of the historical exposure video V1 to obtain the historical activation data of the historical exposure video V1; the pooling process is performed on each historical target video feature vector (V2,a, V2,b, V2,c) of the historical exposure video V2 and the historical activation data of the historical exposure video V1, and the activation unit of the sequence encoding module to be trained is used to activate the obtained processing result to obtain the historical activation data of the historical exposure video V2; the pooling process is performed on each historical target video feature vector (V3,a, V3,b, V3,c) of the historical exposure video V3 and the historical activation data of the historical exposure video V2, and the activation unit of the sequence encoding module to be trained is used to activate the obtained processing result to obtain the historical activation data of the historical exposure video V3. This cyclic processing is performed until the pooling process is performed on each historical target video feature vector (V2,a, V2,b, V2,c) of the historical exposure video Vn and the historical activation data of the historical exposure video Vn-1, and a three-dimensional video vector is output, and the three-dimensional video vector is dimensionally reduced to obtain a one-dimensional historical video vector.

[0083] S304: Smooth the historical user basic attribute data and historical experiment data in the historical user data through the hidden layer to be trained, and splice the historical video vector with the smoothed historical user basic attribute data and historical experiment data through the feature splicing layer to be trained to obtain the historical video splicing feature.

[0084] In the specific process of executing step S304, after inputting the historical user data and historical experiment data into the shunt model to be trained, generate a one-dimensional historical video vector through the sequence encoding module to be trained according to the historical user exposure data in the historical user data. At the same time, to ensure the index confidence of the predicted estimated values of each experimental index obtained, smooth the historical user basic attribute data in the historical user data, the historical context data in the historical experiment data, the historical experiment effect statistical data, and the experimental configuration data through the hidden layer to be trained, and splice the historical video vector with the smoothed historical user basic attribute data, historical context data, historical experiment effect statistical data, and experimental configuration data through the feature splicing layer to be trained to obtain the corresponding historical video splicing feature.

[0085] S305: Process the historical video splicing feature through the multi-layer activation layer to be trained to obtain the historical target video splicing feature.

[0086] S306: Call the softmax function to estimate the predicted estimated value of each experimental index according to the historical target video splicing feature, and adjust the parameters of the shunt model to be trained with the predicted estimated values of each experimental index approaching the corresponding actual estimated values as the training target until the shunt model to be trained converges to obtain the corresponding shunt model.

[0087] In the specific process of executing step S306, after obtaining the historical target video splicing feature, the shunt model to be trained can call the corresponding softmax function to estimate the predicted estimated value of each experimental index according to the historical target video splicing feature, and construct the corresponding loss function according to the predicted estimated value of each experimental index and the actual estimated value of each experimental index, so as to use the constructed loss function to adjust the parameters of the sequence encoding module to be trained, the hidden layer to be trained, the feature splicing layer to be trained, and the multi-layer activation layer to be trained in the shunt model to be trained until the shunt model to be trained converges to obtain the corresponding shunt model. Among them, the shunt model includes a sequence encoding module, a hidden layer, a feature splicing layer, and a multi-layer activation layer.

[0088] In the embodiment of the present application, after each shunt model is pre-trained, each shunt model can be stored in the AB test platform, so that after the corresponding user data and experimental data are obtained, the service layer is called to load each pre-stored shunt model; for each shunt model, the user data and experimental data are input into the shunt model, so that the shunt model generates a one-dimensional video vector according to the user exposure data, and smooths the user basic data and experimental data in the user data, so as to splice the video vector and the smoothed user basic data and experimental data to obtain a video splicing vector, and finally determine the predicted value of each experimental index according to the video splicing vector, and determine the user preference score of the video recommendation result of the user for the shunt model according to the predicted values of each experimental index.

[0089] Optionally, the process of estimating the predicted value of each experimental index according to the user data and experimental data by the shunt model and determining the user preference score of the video recommendation result of the user for the shunt model according to the predicted values of each experimental index can be: converting each exposure video into multiple video feature vectors through a sequence encoding module, and generating a one-dimensional video vector according to each video feature vector of each exposure video; wherein, the user data includes user exposure data and user basic attribute data, and the user exposure data includes multiple exposure videos sorted by exposure time; smoothing the user basic attribute data, context data, experimental effect statistical data and experimental configuration data through a hidden layer, and splicing the video vector and the smoothed user basic attribute data, context data, experimental effect statistical data and experimental configuration data through a feature splicing layer to obtain a video splicing feature; processing the video splicing feature through a multi-layer activation layer to obtain a target video splicing feature, and estimating the predicted value of each experimental index according to the target video splicing feature, and determining the user preference score of the video recommendation result of the user for the shunt model according to the predicted values of each experimental index.

[0090] Optionally, the process of converting each exposure video into multiple video feature vectors through a sequence encoding module and generating a one-dimensional video vector based on the respective video feature vectors of each exposure video can be as follows: Each exposure video is converted into multiple video feature vectors through the sequence encoding module, and each video feature vector is weighted to obtain multiple target video feature vectors corresponding to each exposure video; Through the sequence encoding module, the respective target video feature vectors of the exposure videos except the first exposure video are sequentially subjected to pooling processing with the activation data of the previous exposure video, and the activation unit is used to activate the obtained processing result to obtain activation data until the respective target video feature vectors of the last exposure video are subjected to pooling processing with the activation data of the previous exposure video to obtain a video vector; Among them, the activation data of the first exposure video is obtained by activating the respective target video feature vectors of the first exposure video using the activation unit.

[0091] In the actual application process, each exposure video is converted into video feature vectors corresponding to each experimental index through a sequence encoding module, and the video feature vectors corresponding to each experimental index are weighted and calculated according to the weights corresponding to each experimental index to obtain target video feature vectors corresponding to each experimental index; Starting from the second exposure video, the respective target video feature vectors of the exposure videos are sequentially subjected to pooling processing with the activation data of the previous exposure video, and the activation unit of the sequence encoding module is used to activate the obtained processing result to obtain activation data. Such cyclic processing is performed until the respective target video feature vectors of the last exposure video are subjected to pooling processing with the activation data of the previous exposure video to obtain a three-dimensional video vector, and the three-dimensional video vector is dimensionally reduced to obtain a one-dimensional video vector; Among them, the activation data of the first exposure video is obtained by the sequence encoding module activating the respective target video feature vectors of the first exposure video using the corresponding activation unit.

[0092] The user basic attribute data, context data, experimental effect statistical data, and experimental configuration data are smoothed through a hidden layer, and the video vector and the smoothed user basic attribute data, context data, experimental effect statistical data, and experimental configuration data are concatenated through a feature concatenation layer to obtain a video concatenation feature; The video concatenation feature is processed through the above-mentioned multi-layer activation layer to obtain a target video concatenation feature, and the normalization index of the shunt model is called to estimate the predicted values of each experimental index according to the target video concatenation feature. Finally, the user preference score of the video recommendation result of the user for the shunt model is determined according to the predicted values of each experimental index.

[0093] Optionally, the process of calling the normalization index function to estimate the predicted value of each experimental index based on the target video splicing feature and determining the user preference score of the video recommendation result of the user for the shunt model can be as follows: Call the normalization index function to estimate the predicted value of each experimental index based on the target video splicing feature, calculate the preference score of each experimental index according to the predicted value of the experimental index and its weight, and calculate the user preference score of the video recommendation result of the user for the shunt model according to the preference scores of each experimental index.

[0094] In the actual application process, call the normalization index function to estimate the predicted value of each experimental index based on the target video splicing feature, calculate the product of the predicted value of the experimental index and its weight, obtain the preference score of each experimental index, and finally perform an addition operation on the preference scores of each experimental index to obtain the user preference score of the video recommendation result of the user for the shunt model. Among them, the calculation process of the user preference score of the video recommendation result of the user for the shunt model is shown in formula (1).

[0095] (1)

[0096] Among them, is the user preference score of the video recommendation result of the user for the shunt model A, i is the experimental index, n is the total number of each experimental index, is the weight of the experimental index i in the shunt model A, is the predicted value of the experimental index i obtained through the shunt model A. For example, if n is equal to 3, that is, including three experimental indexes of click-through rate, conversion rate, and viewing duration, then can be the weight of the click-through rate in the shunt model A, can be the weight of the conversion rate in the shunt model A, can be the weight of the viewing duration in the shunt model A.

[0097] It should be noted that for each experimental index, the weight of this experimental index in each shunt model can be the same or different.

[0098] It should also be noted that the user preference score of the shunt model can represent the preference degree of the user for the shunt model under the current conditions.

[0099] S103: For each shunt model, count the number of times the video recommendation result of the shunt model is selected, and determine the model winning rate of the shunt model according to the number of times selected.

[0100] In the process of specifically executing step S103, the corresponding total number of model interactions can be preset in advance. Furthermore, the number of times the user selects the video recommendation results of each shunt model can be counted, that is, the number of times the video recommendation results of each shunt model are selected is counted. Then, the total number of model interactions is divided by the number of times the video recommendation results of each shunt model are selected to obtain the model winning rate of each shunt model. Among them, the calculation method of the model winning rate of each shunt model is shown in formula (2).

[0101] (2)

[0102] Among them, is the model winning rate of shunt model A, is the number of times the video recommendation results of shunt model A are selected, and N is the preset total number of model interactions.

[0103] It should be noted that the total number of model interactions can include the total number of interactions under each experimental index. For example, the total number of model interactions can include the total number of clicks, the total number of conversions, and the total number of views; the number of times a shunt model is selected can include the number of times the video recommendation results of the shunt model are clicked, converted, and viewed.

[0104] It should also be noted that the model winning rate of the shunt model can represent the overall historical performance of the shunt model.

[0105] S104: Calculate the video recommendation score of each shunt model according to the user preference score and the model winning rate of each shunt model, and determine the shunt ratio between each shunt model according to the video recommendation scores of each shunt model.

[0106] In the process of specifically executing step S104, the corresponding user preference score weight and model winning rate weight can be preset in advance. After obtaining the user preference score and the model winning rate of each shunt model, for each shunt model, the product of the user preference score of the shunt model and its user preference score weight and the product of the model winning rate and its model winning rate weight can be calculated, and the sum operation is performed on the product of the user preference score and its user preference score weight and the product of the model winning rate and its model winning rate weight to obtain the video recommendation score of the shunt model. Finally, the shunt ratio between each shunt model is determined according to the video recommendation scores of each shunt model. Among them, the calculation method of the video recommendation score of the shunt model is shown in formula (3).

[0107] (3)

[0108] Among them, S is the video recommendation score, is the user preference score weight, is the model winning rate weight, is the user preference score, is the model winning rate.

[0109] S105: Display the current video recommendation results of each shunt model to the user according to the shunt ratio.

[0110] In the specific process of executing step S105, after obtaining the shunt ratio between each shunt model, the recommendation result proportion of the video recommendation results of each shunt model can be allocated according to the shunt ratio, so as to display the current video recommendation results of each shunt model to the user according to the recommendation result proportion of each shunt model. At the same time, a corresponding visualization effect report can be generated according to the recommendation result proportion of each shunt model, and the corresponding visualization effect report can be displayed through the page provided by the display layer.

[0111] The embodiment of the present application provides a video recommendation method based on the AB test. Each shunt model to be trained can be pre-trained using historical user data and historical experiment data to obtain each shunt model, so that when receiving an online recommendation request sent by a user, the user data and experiment data of the user can be obtained based on the online recommendation request. Each shunt model estimates the predicted value of each experimental index according to the user data and experiment data, and determines the user preference score of the user for the video recommendation result of each shunt model according to the predicted value of each experimental index. At the same time, the model winning rate of each shunt model can also be calculated, so as to calculate the video recommendation score of each shunt model according to the user preference score and model winning rate of each shunt model. Finally, the shunt ratio between each shunt model can be determined according to the video recommendation scores of each shunt model, so as to display the current video recommendation results of each shunt model to the user according to the obtained shunt ratio. It can be seen that the technical solution provided by the present application does not need to ensure that the user distribution of each shunt model is uniform, nor does it need to conduct long-term experimental effect observation and statistics. It can not only shorten the effect recovery period, avoid fluctuations in the effect, but also avoid the influence of factors such as time period, user population and user interest changes on the effects of different shunt models.

[0112] Based on the video recommendation method based on the AB test provided by the embodiment of the present application, correspondingly, the embodiment of the present application also provides a video recommendation system based on the AB test, as Figure 6 shown. The video recommendation system based on the AB test includes:

[0113] The first acquisition unit 61 is used to obtain the user data and experiment data of the user based on the online recommendation request when receiving the online recommendation request sent by the user;

[0114] A user preference score determination unit 62, configured to load each pre-trained shunt model, estimate the predicted value of each experimental metric according to the user data and the experimental data through the shunt model, and determine the user preference score of the user for the video recommendation result of the shunt model according to the predicted values of the respective experimental metrics; wherein, the shunt model is obtained by a training unit training the shunt model to be trained using historical user data and historical experimental data;

[0115] A model win rate determination unit 63, configured to, for each shunt model, count the number of times the video recommendation result of the shunt model is selected, and determine the model win rate of the shunt model according to the number of times selected;

[0116] A shunt ratio determination unit 64, configured to calculate the video recommendation score of each shunt model according to the user preference score and the model win rate of each shunt model, and determine the shunt ratio between the respective shunt models according to the video recommendation scores of the respective shunt models;

[0117] A video recommendation unit 65, configured to display the current video recommendation results of the respective shunt models to the user according to the shunt ratio.

[0118] For the specific principles and execution processes of each unit in the video recommendation system based on the AB test disclosed in the embodiments of the present application above, they are the same as those of the video recommendation method based on the AB test disclosed in the embodiments of the present application above. Reference can be made to the corresponding parts in the video recommendation method based on the AB test disclosed in the embodiments of the present application above, and details will not be elaborated here.

[0119] The embodiments of the present application provide a video recommendation system based on an AB test, which can pre-train each shunt model to be trained using historical user data and historical experimental data to obtain each shunt model, so that when receiving an online recommendation request sent by a user, the user data and experimental data of the user are obtained based on the online recommendation request, the predicted value of each experimental metric is estimated through each shunt model according to the user data and the experimental data, and the user preference score of the user for the video recommendation result of each shunt model is determined according to the predicted values of the respective experimental metrics. At the same time, the model win rate of each shunt model can also be calculated, so as to calculate the video recommendation score of each shunt model according to the user preference score and the model win rate of each shunt model. Finally, the shunt ratio between the respective shunt models can be determined according to the video recommendation scores of the respective shunt models, so as to display the current video recommendation results of the respective shunt models to the user according to the obtained shunt ratio. It can be seen that the technical solution provided by the present application does not need to ensure that the user distribution of each shunt model is uniform, nor does it need to conduct long-term statistical observation of experimental effects. It can not only shorten the effect recovery period, avoid fluctuations in effects, but also avoid the influence of factors such as time period, user population, and user interest changes on the effects of different shunt models.

[0120] Optionally, the user preference score determination unit includes:

[0121] A first input unit for loading each pre-trained shunt model and inputting user data and experimental data into each shunt model; wherein, the shunt model includes a sequence encoding module, a hidden layer, a feature splicing layer, and a multi-layer activation layer;

[0122] A video vector generation unit for converting each exposure video into multiple video feature vectors through the sequence encoding module and generating a one-dimensional video vector based on the respective video feature vectors of each exposure video; wherein, the user data includes user exposure data and user basic attribute data, and the user exposure data includes multiple exposure videos sorted by exposure time;

[0123] A first feature splicing unit for smoothing the user basic attribute data, context data, and experimental effect statistical data through the hidden layer, and splicing the video vector and the smoothed user basic attribute data, context data, and experimental effect statistical data through the feature splicing layer to obtain a video splicing feature;

[0124] A second processing unit for processing the video splicing feature through the multi-layer activation layer to obtain a target video splicing feature;

[0125] An estimation unit for calling a normalization index function to estimate the predicted values of each experimental index based on the target video splicing feature, and determining the user preference score of the user's video recommendation result for the shunt model according to the predicted values of each experimental index.

[0126] Optionally, the video vector generation unit includes:

[0127] A weighted processing unit for converting each exposure video into multiple video feature vectors through the sequence encoding module and performing weighted processing on each video feature vector to obtain multiple target video feature vectors corresponding to each exposure video;

[0128] A video vector generation sub-unit for sequentially performing pooling processing on the respective target video feature vectors of the exposure videos except the first exposure video and the activation data of its previous exposure video through the sequence encoding module, and activating the obtained processing result using an activation unit to obtain activation data until the respective target video feature vectors of the last exposure video and the activation data of its previous exposure video are subjected to pooling processing to obtain a video vector; wherein, the activation data of the first exposure video is obtained by activating the respective target video feature vectors of the first exposure video using an activation unit.

[0129] Optionally, the estimation unit includes:

[0130] An estimation subunit, configured to call a normalization metric function to estimate the predicted value of each experimental metric according to the target video splicing feature, calculate the preference score of each experimental metric according to the predicted value of the experimental metric and its weight, and calculate the user preference score of the video recommendation result of the user for the traffic splitting model according to the preference scores of each experimental metric.

[0131] Optionally, the training unit includes:

[0132] A second acquisition unit, configured to acquire historical user data, historical experimental data, and the actual predicted values of each experimental metric of the video recommendation results of historical users for each traffic splitting model to be trained;

[0133] A second input unit, configured to input the historical user data, historical experimental data, and the actual predicted values of each experimental metric of the video recommendation results of historical users for each traffic splitting model to be trained into each traffic splitting model to be trained;

[0134] A historical video vector generation unit, configured to generate a one-dimensional historical video vector according to the user exposure data of the user data through a sequence encoding module to be trained;

[0135] A second feature splicing unit, configured to perform smoothing processing on the historical user basic attribute data and historical experimental data in the historical user data through a hidden layer to be trained, and splice the historical video vector and the smoothed historical user basic attribute data and historical experimental data through a feature splicing layer to be trained to obtain a historical video splicing feature;

[0136] A second feature splicing unit, configured to process the historical video splicing feature through a multi-layer activation layer to be trained to obtain a historical target video splicing feature;

[0137] A training subunit, configured to call a normalization exponential function to estimate the predicted predicted value of each experimental metric according to the historical target video splicing feature, and adjust the parameters of the traffic splitting model to be trained with the predicted predicted values of each experimental metric approaching the corresponding actual predicted values as the training target until the traffic splitting model to be trained converges to obtain the corresponding traffic splitting model.

[0138] The present application also provides a storage medium, in which program instructions are stored, and when the program instructions are loaded and executed by a processor, the above-mentioned embodiments of any video recommendation method based on an AB test are implemented.

[0139] The present application also provides an electronic device, such as Figure 7As shown in the figure, the electronic device includes a processor 701 and a memory 702, and the processor and the memory are connected through a communication bus; the processor and the memory are connected through the communication bus; wherein, the processor is configured to call and execute a program stored in the memory; the memory is configured to store a program, and the program is used to implement any one of the above video recommendation methods based on the AB test.

[0140] The processor of the present application may be the CPU of the terminal, or an MCU integrated in the terminal, or a combination of the CPU and the MCU. Moreover, the processor includes a kernel, and the kernel retrieves the corresponding program from the memory, and one or more kernels may be provided.

[0141] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one memory chip.

[0142] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0143] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0144] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0145] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A video recommendation method based on AB experiment, characterized in that: The method comprises: When receiving an online recommendation request sent by a user, obtaining user data and experimental data of the user based on the online recommendation request; Loading each pre-trained traffic diversion model, and using the traffic diversion model to estimate the estimated value of each experimental indicator according to the user data and the experimental data, and determining the user preference score of the user for the video recommendation result of the traffic diversion model according to the estimated value of each experimental indicator; wherein the traffic diversion model is obtained by training the traffic diversion model to be trained using historical user data and historical experimental data; For each of the diversion models, counting the number of times the video recommendation results of the diversion model are selected, and determining the model winning rate of the diversion model according to the number of times; Calculate the video recommendation score of each diversion model according to the user preference score and model win rate of each diversion model, and determine the diversion ratio between each diversion model according to the video recommendation score of each diversion model; The current video recommendation results of each diversion model are displayed to the user according to the diversion ratio.

2. The method according to claim 1, characterized in that The loading of the pre-trained diversion models, estimating the estimated value of each experimental indicator according to the user data and the experimental data by the diversion model, and determining the user preference score of the user for the video recommendation result of the diversion model according to the estimated value of each experimental indicator, includes: Loading each pre-trained traffic diversion model, and inputting the user data and the experimental data into each of the traffic diversion models; wherein the traffic diversion model includes a sequence encoding module, a hidden layer, a feature concatenation layer, and a multi-layer activation layer; Each exposure video is converted into a plurality of video feature vectors by the sequence encoding module, and a one-dimensional video vector is generated according to each video feature vector of each exposure video; wherein the user data includes user exposure data and user basic attribute data, and the user exposure data includes a plurality of exposure videos sorted by exposure time; The user basic attribute data, context data and experimental effect statistical data are smoothed by the hidden layer, and the video vector and the smoothed user basic attribute data, context data and experimental effect statistical data are spliced ​​by the feature splicing layer to obtain a video splicing feature; Processing the video splicing features through the multi-layer activation layer to obtain target video splicing features; A normalized index function is called to estimate an estimated value of each experimental index according to the target video splicing feature, and a user preference score of the user for the video recommendation result of the diversion model is determined according to the estimated value of each experimental index.

3. The method according to claim 2, characterized in that The sequence encoding module converts each exposure video into a plurality of video feature vectors, and generates a one-dimensional video vector according to each video feature vector of each exposure video, including: Converting each of the exposure videos into a plurality of video feature vectors through the sequence encoding module, and performing weighted processing on each of the video feature vectors to obtain a plurality of target video feature vectors corresponding to each exposure video; The sequence encoding module sequentially performs pooling processing on each target video feature vector of the exposure videos except the first exposure video and the activation data of the previous exposure video, and uses the activation unit to activate the processing results to obtain activation data, until each target video feature vector of the last exposure video and the activation data of the previous exposure video are pooled to obtain a video vector; wherein the activation data of the first exposure video is obtained by activating each target video feature vector of the first exposure video using the activation unit.

4. The method according to claim 2, characterized in that: The calling of the normalized index function estimates the estimated value of each experimental index according to the target video splicing feature, and determines the user preference score of the user for the video recommendation result of the diversion model according to the estimated value of each experimental index, including: A normalized indicator function is called to estimate the estimated value of each experimental indicator according to the target video splicing feature, and a preference score of each experimental indicator is calculated according to the estimated value of the experimental indicator and its weight, and a user preference score of the user for the video recommendation result of the diversion model is calculated according to the preference score of each experimental indicator.

5. The method according to claim 1, characterized in that The method of training the trained traffic diversion model by using historical user data and historical experimental data to obtain the traffic diversion model includes: Obtain historical user data, historical experimental data, and actual estimated values ​​of various experimental indicators of historical users' video recommendation results for each diversion model to be trained; Inputting historical user data, historical experimental data, and actual estimated values ​​of various experimental indicators of video recommendation results of historical users for each diversion model to be trained into each diversion model to be trained; Generate a one-dimensional historical video vector according to the user exposure data of the user data by the sequence encoding module to be trained; The historical user basic attribute data and the historical experimental data in the historical user data are smoothed by the hidden layer to be trained, and the historical video vector and the smoothed historical user basic attribute data and the historical experimental data are spliced ​​by the feature splicing layer to be trained to obtain a historical video splicing feature; Processing the historical video splicing features through the multi-layer activation layer to be trained to obtain historical target video splicing features; A normalized exponential function is called to estimate the predicted estimated value of each experimental indicator according to the historical target video splicing features, and the parameters of the diversion model to be trained are adjusted with the predicted estimated value of each experimental indicator approaching the corresponding actual estimated value as the training goal, until the diversion model to be trained converges to obtain a corresponding diversion model.

6. A video recommendation system based on AB experiment, characterized in that: The system comprises: A first acquisition unit, configured to, when receiving an online recommendation request sent by a user, acquire user data and experimental data of the user based on the online recommendation request; A user preference score determination unit is used to load each pre-trained diversion model, and estimate the estimated value of each experimental indicator according to the user data and the experimental data through the diversion model, and determine the user preference score of the user for the video recommendation result of the diversion model according to the estimated value of each experimental indicator; wherein the diversion model is obtained by training the diversion model to be trained by the training unit using historical user data and historical experimental data; A model winning rate determining unit, used for counting the number of times the video recommendation result of the diversion model is selected for each diversion model, and determining the model winning rate of the diversion model according to the number of times the video recommendation result is selected; A diversion ratio determination unit, configured to calculate a video recommendation score for each diversion model according to a user preference score and a model win rate of each diversion model, and determine a diversion ratio between each diversion model according to the video recommendation scores of each diversion model; The video recommendation unit is used to display the current video recommendation results of each diversion model to the user according to the diversion ratio.

7. The system according to claim 6, characterized in that The user preference score determination unit comprises: A first input unit is used to load each pre-trained diversion model and input the user data and the experimental data into each diversion model; wherein the diversion model includes a sequence encoding module, a hidden layer, a feature concatenation layer and a multi-layer activation layer; A video vector generating unit, configured to convert each exposure video into a plurality of video feature vectors through the sequence encoding module, and generate a one-dimensional video vector according to each video feature vector of each exposure video; wherein the user data includes user exposure data and user basic attribute data, and the user exposure data includes a plurality of exposure videos sorted by exposure time; A first feature splicing unit is used to smooth the user basic attribute data, context data and experimental effect statistical data through the hidden layer, and to splice the video vector with the smoothed user basic attribute data, context data and experimental effect statistical data through the feature splicing layer to obtain a video splicing feature; A second processing unit, configured to process the video splicing features through the multi-layer activation layer to obtain a target video splicing feature; The estimation unit is used to call the normalized index function to estimate the estimated value of each experimental index according to the target video splicing feature, and determine the user preference score of the user for the video recommendation result of the diversion model according to the estimated value of each experimental index.

8. The system according to claim 7, characterized in that The video vector generating unit comprises: A weighted processing unit, configured to convert each of the exposure videos into a plurality of video feature vectors through the sequence encoding module, and perform weighted processing on each of the video feature vectors to obtain a plurality of target video feature vectors corresponding to each exposure video; A video vector generating subunit is used to perform pooling processing on each target video feature vector of the exposure videos except the first exposure video and the activation data of the previous exposure video in sequence through the sequence encoding module, and activate the processing results obtained by using the activation unit to obtain activation data, until each target video feature vector of the last exposure video is pooled with the activation data of the previous exposure video to obtain the video vector; wherein the activation data of the first exposure video is obtained by activating each target video feature vector of the first exposure video by using the activation unit.

9. An electronic device, characterized in that: include: A processor and a memory, wherein the processor and the memory are connected via a communication bus; wherein the processor is used to call and execute a program stored in the memory; The memory is used to store a program, and the program is used to implement the video recommendation method based on AB experiment as described in any one of claims 1 to 5.

10. A storage medium, characterized in that: The storage medium stores computer executable instructions, and the computer executable instructions are used to execute the video recommendation method based on AB experiment according to any one of claims 1 to 5.

Citation Information

Cited By

  • Algorithm information visualization interaction method and system

    CN120343362A

  • Algorithm information visualization interaction method and system

    CN120343362B