Recommendation method and system and storage medium
By offline simulation of the traffic data of the recommendation system and determining the target recommendation strategy, the problem of overloading of computing resources in non-ideal scenarios is solved, and adaptive adjustment and resource optimization are achieved.
Patent Information
- Application Number
- CN202510495018.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-25
AI Technical Summary
The existing recommendation system cannot achieve global optimization in non-ideal scenarios, resulting in the problem of overloading of computing resources.
By performing offline simulation of the traffic data in the second period, the target recommendation strategy is determined so that it meets the condition that resource consumption does not exceed the total amount within the first period, thereby avoiding overloading of computing resources.
Adaptively adjusting the recommendation strategy in different scenarios is realized, the traffic utilization and recommendation effect of the recommendation system are improved, and the computing resource overload is avoided.
Smart Images

Figure CN120372093A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of Internet technologies, and particularly to a recommendation method, system, and storage medium. Background Art
[0002] With the development of Internet technologies, there are multiple exhibition booths on the user pages of more and more Internet products. When a user accesses a page, an Internet product can, based on multi-exhibition-booth recommendation technology, recommend different content for the user at each exhibition booth to increase the exposure rate of the recommended content and provide the user with more diverse choices. A recommendation system is deployed in the Internet product, and the recommendation system can make a recommendation decision (decide what content to recommend for each exhibition booth) according to a preset recommendation evaluation function.
[0003] The above-mentioned recommendation evaluation function is modeled based on an ideal scenario. In the ideal scenario, the recommendation system can achieve a good recommendation effect by making a recommendation decision based on this recommendation evaluation function. However, in actual application, the recommendation system cannot operate in the ideal scenario. In a non-ideal scenario, the recommendation effect achieved by the recommendation system making a recommendation decision based on this recommendation evaluation function is not ideal, cannot reach the global optimum, and often causes the situation that the computing resources of the recommendation system are overloaded.
[0004] The content in the background art section is only the information known to the inventor personally, and does not represent that the above information has entered the public domain before the filing date of this disclosure, nor does it represent that it can become the prior art of this disclosure. Summary of the Invention
[0005] This specification provides a recommendation method, system, and storage medium, which can determine a corresponding target recommendation strategy for each time period to adapt to the change of the scenario, and the amount of resources required by the target recommendation strategy does not exceed the total amount of resources, which can avoid the situation of computing resource overload.
[0006] In a first aspect, this specification provides a recommendation method, which is applied to a recommendation system. The method includes: in response to receiving a recommendation request, determining a first time period to which the receiving moment of the recommendation request belongs; obtaining a target recommendation strategy adopted by the recommendation system during the first time period, where the target recommendation strategy is a strategy determined by offline simulation of traffic data during a second time period, and the recommendation evaluation index obtained by performing offline simulation on the traffic data using the target recommendation strategy meets a preset condition and the amount of resources required does not exceed the total amount of resources of the recommendation system, and the second time period is before the first time period; and determining a recommendation action corresponding to the recommendation request based on the target recommendation strategy, and executing the recommendation action to respond to the recommendation request.
[0007] In some embodiments, the target recommendation strategy is obtained in the following manner: obtaining the traffic data within the second time period; obtaining an initial recommendation strategy and determining the initial distribution information corresponding to the initial recommendation strategy; adjusting the initial distribution information based on the traffic data within the second time period to obtain updated distribution information; and determining the target recommendation strategy based on the updated distribution information, wherein the recommendation evaluation index obtained by processing the traffic data within the second time period using the target recommendation strategy is not worse than the recommendation evaluation index obtained by processing the traffic data within the second time period using the initial recommendation strategy.
[0008] In some embodiments, the adjusting the initial distribution information based on the traffic data within the second time period to obtain updated distribution information includes: performing multiple rounds of iteration on the initial distribution information based on the traffic data within the second time period until a preset end condition is reached, wherein for any positive integer i, the i-th round of iteration process includes: sampling multiple sampling recommendation strategies based on the current distribution information, where when i = 1, the current distribution information is the initial distribution information, and when i > 1, the current distribution information is the distribution information updated in the (i - 1)-th round of iteration; performing offline simulation on the traffic data within the second time period using the multiple sampling recommendation strategies respectively to obtain the recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies respectively; updating the respective recommendation evaluation indexes based on the resource consumption amounts required by the multiple sampling recommendation strategies, wherein when the resource consumption amount required by a sampling recommendation strategy is greater than the total resource amount, the updated recommendation evaluation index corresponding to this sampling recommendation strategy is less than the recommendation evaluation index before the update; and determining the updated distribution information based on the updated recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies.
[0009] In some embodiments, determining the updated distribution information based on the updated recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies includes: sorting the multiple sampling recommendation strategies based on the updated recommendation evaluation indexes corresponding to them; and determining the updated distribution information based on the multiple sampling recommendation strategies and the sorting result of the multiple sampling recommendation strategies.
[0010] In some embodiments, the determining the updated distribution information based on the multiple sampling recommendation strategies and the sorting result of the multiple sampling recommendation strategies includes: determining the top K sampling recommendation strategies with the highest sorting in the multiple sampling recommendation strategies based on the sorting result of the multiple sampling recommendation strategies, where K is an integer greater than or equal to 1; and determining the updated distribution information according to the distribution information of the K sampling recommendation strategies.
[0011] In some embodiments, determining the updated distribution information based on the multiple sampling recommendation strategies and the sorting results of the multiple sampling recommendation strategies includes: determining the weights corresponding to the multiple sampling recommendation strategies respectively through a polytope mapping algorithm based on the multiple sampling recommendation strategies and the sorting results of the multiple sampling recommendation strategies, where the weights corresponding to the top K sampling recommendation strategies in the sorting are greater than the weights corresponding to the sampling recommendation strategies ranked later, and K is an integer greater than or equal to 1; and determining the updated distribution information based on the multiple sampling recommendation strategies, the weights corresponding to the multiple sampling recommendation strategies respectively, and K.
[0012] In some embodiments, for any one of the multiple sampling recommendation strategies, the following formula is used for the update process of the recommendation evaluation index corresponding to the sampling recommendation strategy:
[0013] Value′(θ) = Value(θ) - λmin{C - Cost(θ), 0}
[0014] where Value(θ) is the recommendation evaluation index corresponding to the sampling recommendation strategy, Value′(θ) is the updated recommendation evaluation index, λ is a constant coefficient, C is the total amount of resources of the recommendation system, and Cost(θ) is the amount of resources consumed by the sampling recommendation strategy.
[0015] In some embodiments, using the multiple sampling recommendation strategies to perform offline simulation on the traffic data respectively to obtain the recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies respectively includes: traversing the multiple sampling recommendation strategies, and performing the following steps for the current sampling recommendation strategy during the traversal: inputting the traffic data into a pre-trained simulation system, and controlling the simulation system to process the traffic data using the current sampling recommendation strategy to obtain the recommendation evaluation index corresponding to the current sampling recommendation strategy.
[0016] In some embodiments, the preset end condition includes at least one of the following: the current iteration number reaches a preset number; or the difference between the updated distribution information and the current distribution information is less than or equal to a preset difference.
[0017] In some embodiments, the initial recommendation strategy is a strategy determined based on the traffic data in the third time period. The recommendation evaluation index obtained by performing recommendation simulation on the traffic data using the initial recommendation strategy meets a preset condition and the amount of resources consumed does not exceed the total amount of resources of the recommendation system. The third time period is before the second time period.
[0018] In some embodiments, the initial recommendation strategy is a recommendation strategy obtained by random initialization.
[0019] In some embodiments, the target recommendation strategy is a strategy determined by the recommendation system through offline simulation of traffic data in the second period during the fourth period. The fourth period is before the first period and after the second period. Determining the target recommendation strategy based on the updated distribution information includes: determining a candidate recommendation strategy based on the updated distribution information, where the recommendation evaluation index obtained by processing the traffic data in the second period using the candidate recommendation strategy is not worse than the recommendation evaluation index obtained by processing the traffic data in the second period using the initial recommendation strategy; determining at least part of the traffic in the fourth period as target traffic data, and obtaining the true recommendation evaluation index obtained by the recommendation system processing the target traffic data in the fourth period; performing offline simulation on the target traffic data using the candidate recommendation strategy to obtain the predicted recommendation evaluation index corresponding to the candidate recommendation strategy; and in the case where the predicted recommendation evaluation index is greater than or equal to the true recommendation evaluation index, determining the candidate recommendation strategy as the target recommendation strategy, or in the case where the predicted recommendation evaluation index is less than the true recommendation evaluation index, using the recommendation strategy adopted by the recommendation system in the fourth period as the target recommendation strategy.
[0020] In some embodiments, the recommendation system includes a recall module and a fine ranking module, and the target recommendation strategy includes at least one of the following: the recall channels used in the recall module and the queue lengths of the recall channels; or the metric prediction model used in the fine ranking module.
[0021] In some embodiments, the recommendation request is used to request information recommendation for multiple exhibition booths, and the target recommendation strategy includes sub-recommendation strategies corresponding to the multiple exhibition booths respectively.
[0022] In some embodiments, determining the recommendation action corresponding to the recommendation request based on the target recommendation strategy includes: for each target booth among the multiple exhibition booths: obtaining the recommendation prediction data corresponding to the target booth, where the recommendation prediction data at least includes: the response time corresponding to the target booth and the historical recommendation evaluation index corresponding to the target booth; and determining the recommendation action corresponding to the recommendation request based on the recommendation prediction data corresponding to the target booth and the sub-recommendation strategy corresponding to the target booth.
[0023] In a second aspect, this specification also provides a recommendation system, including: at least one storage medium storing at least one instruction set for performing data processing related to content recommendation; and at least one processor communicatively connected to the at least one storage medium. When the recommendation system runs, the at least one processor reads the at least one instruction set and implements the method according to any one of the first aspects according to the instructions of the at least one instruction set.
[0024] In a third aspect, this specification also provides a computer-readable non-volatile storage medium, where at least one instruction set is stored in the computer-readable non-volatile storage medium, and when the at least one instruction set is executed by at least one processor, the method according to any one of the first aspects is implemented.
[0025] Other functions of the recommendation method, system, and storage medium provided in this specification will be partially listed in the following description. The creative aspects of the recommendation method, system, and storage medium provided in this specification can be fully explained by practicing or using the methods, devices, and combinations described in the detailed examples below. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0027] Figure 1 FIG. shows a schematic diagram of a recommendation scenario provided according to an embodiment of this specification;
[0028] Figure 2 FIG. shows a hardware structure diagram of a computing system provided according to an embodiment of this specification;
[0029] Figure 3 FIG. shows a flowchart of a recommendation method provided according to an embodiment of this specification;
[0030] Figure 4 FIG. shows a timing diagram when a recommendation method provided according to an embodiment of this specification is applied;
[0031] Figure 5 FIG. shows a schematic diagram of determining a target recommendation strategy through a cross-entropy evolution algorithm; and
[0032] Figure 6 FIG. shows a schematic diagram of responding to a recommendation request according to a target recommendation strategy provided according to an embodiment of this specification. Detailed Implementation Modes
[0033] The following description provides specific application scenarios and requirements of this specification, aiming to enable those skilled in the art to manufacture and use the content in this specification. For those skilled in the art, various local modifications to the disclosed embodiments are obvious, and without departing from the spirit and scope of this specification, the general principles defined here can be applied to other embodiments and applications. Therefore, this specification is not limited to the disclosed embodiments, but has the broadest scope consistent with the claims.
[0034] The terms used herein are for the purpose of describing specific example embodiments only and are not restrictive. For example, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" used herein may also include the plural forms. When used in this specification, the terms "comprises", "comprising" and / or "containing" mean that the associated integers, steps, operations, elements and / or components exist, but do not exclude the existence of one or more other features, integers, steps, operations, elements, components and / or groups, or the addition of other features, integers, steps, operations, elements, components and / or groups in the system / method.
[0035] In view of the following description, these features and other features of this specification, as well as the operations and functions of the relevant elements of the structure, and the combination and manufacturing economy of the components can be significantly improved. Referring to the accompanying drawings, all of these form a part of this specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0036] The flowcharts used in this specification illustrate the operations implemented by the system according to some embodiments in this specification. It should be clearly understood that the operations in the flowchart may not be implemented in sequence. On the contrary, the operations may be implemented in reverse order or simultaneously. In addition, one or more other operations may be added to the flowchart. One or more operations may be removed from the flowchart.
[0037] The application scenarios of this specification are introduced below.
[0038] The technical solution provided in this specification is applicable to the scenario of recommending information to users. In this scenario, multiple exhibition booths can be set in the application program or target page corresponding to the recommendation system, and the exhibition booths are used to display the relevant content of the products to be promoted. When a user enters the application program or accesses the target page, the recommendation system can determine the recommendation actions corresponding to each exhibition booth and execute the recommendation actions to display the relevant content of the products to be promoted corresponding to the recommendation actions in each exhibition booth. Among them, the products to be promoted can provide products or services in any form, including but not limited to providing credit services, providing trading services, and providing information services, etc. The relevant content of the products to be promoted is used to display the specific content of the services provided by the products to be promoted.
[0039] This specification provides a recommendation method, which can be executed by a recommendation system. First, when the recommendation system receives a recommendation request, it can respond to the received recommendation request and determine the first time period to which the reception moment of the recommendation request belongs. Then, the recommendation system can determine the recommendation action corresponding to the recommendation request according to the target recommendation strategy adopted in the first time period and execute the recommendation action to respond to the recommendation request. Among them, the target recommendation strategy adopted in the first time period is a strategy determined by the recommendation system through offline simulation of the traffic data in the second time period. The recommendation evaluation index obtained by the recommendation system using the target recommendation strategy for offline simulation of the traffic data meets the preset conditions and the required resource consumption does not exceed the total resources of the recommendation system. The second time period is before the first time period.
[0040] In this specification, the recommendation evaluation index is an evaluation index related to the expected recommendation effect and is used to evaluate whether the current recommendation achieves the expected recommendation effect. The expected recommendation effect is related to the requirements of the actual application scenario, and this specification does not limit the expected recommendation effect. For example, the expected recommendation effect may include, but is not limited to, one or more of the following: high conversion rate, high revenue, high number of clicks, etc.
[0041] In the recommendation method provided in this specification, the traffic conditions that the recommendation system needs to process in adjacent time periods are basically the same. Therefore, the recommendation system can perform offline simulation on the traffic data in the second time period, update the recommendation strategy, and obtain the target recommendation strategy applied in the first time period. The recommendation system can make recommendations for the recommendation requests received in the first time period based on the target recommendation strategy. In the process of offline determining the target recommendation strategy, the recommendation system takes the resource consumption not exceeding the total amount of resources as a hard constraint, so that the target recommendation strategy determined offline will not cause a situation of computational resource overload during execution. Moreover, since the recommendation system uses different target recommendation strategies in different time periods, when the recommendation system determines the recommended content to be displayed at each booth, it can adaptively adjust the recommendation strategy according to the scenario where the recommendation system is located. The target recommendation strategy is more suitable for the current scenario, and the recommendation evaluation index for making recommendations according to the target recommendation strategy is better. Thus, the utilization rate of traffic by the recommendation system is improved, and the recommendation effect is further enhanced.
[0042] Figure 1 FIG. shows a schematic diagram of a recommendation scenario provided according to an embodiment of this specification. As Figure 1 shown, the scenario 100 may include a recommendation system 11 and a client 12. Figure 1 The shown scenario 100 may be a recommendation scenario of a certain Internet product. For example, the Internet product may include a client and a server. Among them, Figure 1 the client 12 therein may correspond to the client of the Internet product, Figure 1 the recommendation system 11 therein may correspond to the server of the Internet product, or correspond to a subsystem in the server of the Internet product.
[0043] Referring to Figure 1 , the application program may be the target application program corresponding to the Internet product in the client 12, and the target page may refer to the page related to the Internet product displayed to the user by the client 12 through an application with a page access function. The application program or the target page may include K booths, and each booth is used to display the relevant content of a promoted product.
[0044] Referring to Figure 1 , the recommendation system 11 may include a policy update module. The policy update module may perform offline simulation based on the traffic data received in the previous time period, and determine the target recommendation strategy for the next time period.
[0045] In some embodiments, referring to Figure 1, the recommendation system 11 may further include an online recommendation module. The online recommendation module may be directly or indirectly communicatively connected to the client 12. For example, when the client 12 detects that a target user requests to access a target page or launch a target application, it may send a recommendation request to the recommendation system 11 based on the booths in the target page or target application. The online recommendation module in the recommendation system 11 receives the recommendation request and, in response to the recommendation request, determines the recommendation actions for each booth corresponding to the recommendation request according to the target recommendation strategy. Further, the online recommendation module in the recommendation system 11 may execute the recommendation actions corresponding to each booth and obtain the recommended content to be displayed corresponding to each booth. Then, the recommendation system 11 may send the recommended content to be displayed corresponding to each booth to the client 12 so that the client 12 can render and display the recommended content corresponding to each booth based on the above-mentioned recommended content.
[0046] In some embodiments, the recommendation method provided in this specification may be executed on the recommendation system 11. For example, the recommendation method provided in this specification may be executed by the online recommendation module in the recommendation system 11. At this time, the recommendation system 11 may store the data or instructions for executing the recommendation method described in this specification and may execute or be used to execute the data or instructions. In some embodiments, the recommendation system 11 may include a hardware device with data information processing capabilities and the necessary programs for driving the hardware device to work.
[0047] The recommendation system 11 may correspond to a single computing device or a computing cluster composed of multiple computing devices.
[0048] In some embodiments, one or more applications (APPs) may be installed on the client 12. The APP can provide the ability to receive user operations and an interface. The APPs include, but are not limited to: financial APP programs, web browser APP programs, search APP programs, chat APP programs, shopping APP programs, video APP programs, financial management APP programs, instant messaging tools, email clients, social platform software, and so on.
[0049] In some embodiments, a target application (APP) corresponding to an Internet product may be installed on the client 12. The target APP may send a recommendation request to the recommendation system 11 when starting up. The target APP may also receive the recommended content to be displayed corresponding to each booth from the recommendation system 11 and display the relevant content of the corresponding products placed in each booth on the client 12.
[0050] In some embodiments, the client 12 can receive a user's operation to access a target page through applications such as a web browser-like APP or a shopping APP. When the client 12 responds to the user's operation of accessing the target page, it can send a recommendation request to the recommendation system 11 based on the booths in the target page. Then, the client 12 can also receive the recommended content to be displayed corresponding to each booth from the recommendation system 11 and display the relevant content of the corresponding placement products in each booth on the client 12.
[0051] It should be understood that Figure 1 the number of clients 12 in
[0052] Figure 2 shows a hardware structure diagram of a computing system provided according to an embodiment of the present specification. The computing system 200 can be used as Figure 1 the recommendation system 11 in
[0053] As Figure 2 shown, the computing system 200 can include at least one storage medium 230 and at least one processor 220. In some embodiments, the computing system 200 can also include a communication port 250 and an internal communication bus 210. The computing system 200 can also include I / O components 260.
[0054] The internal communication bus 210 can connect different system components. For example, the internal communication bus 210 can connect the storage medium 230, the processor 220, the communication port 250, and the I / O components 260, etc.
[0055] The I / O components 260 support input / output between the computing system 200 and other components.
[0056] The communication port 250 is used for data communication between the computing system 200 and the outside world. For example, the communication port 250 can be used for data communication between the computing system 200 and a network. The communication port 250 can be a wired communication port or a wireless communication port.
[0057] The storage medium 230 can include a data storage device. The data storage device can be a non-transitory storage medium or a transitory storage medium. For example, the data storage device can include one or more of a magnetic disk 232, a read-only storage medium (ROM) 234, or a random access storage medium (RAM) 235. The storage medium 230 also includes at least one instruction set stored in the data storage device. The instruction set can include computer program code, and the computer program code can include programs, routines, objects, components, data structures, processes, modules, and so on.
[0058] At least one processor 220 may be communicatively coupled to at least one storage medium 230. When the computing system 200 is running, the at least one processor 220 reads the at least one instruction set and executes the recommendation method provided in this specification according to the instructions of the at least one instruction set. The processor 220 may execute the steps included in the recommendation method. The processor 220 may be in the form of one or more processors. In some embodiments, the processor 220 may include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application specific integrated circuit (ASIC), an application specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physics processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, etc., or any combination thereof.
[0059] For illustrative purposes only, only one processor 220 is shown in the computing system 200 in the drawings. However, it should be noted that the computing system 200 in this specification may also include multiple processors. Therefore, the operations and / or method steps disclosed in this specification may be executed by one processor or jointly executed by multiple processors. For example, if it is described in this specification that the processor 220 of the computing system 200 executes step A and step B, it should be understood that step A and step B may also be executed jointly or separately by two different processors 220 (e.g., the first processor executes step A, the second processor executes step B, or the first and second processors jointly execute steps A and B).
[0060] Figure 3 A flowchart of a recommendation method provided according to an embodiment of this specification is shown. As before, the recommendation system 11 may execute this recommendation method.
[0061] As Figure 3 shown, the recommendation method may include:
[0062] S310: In response to receiving a recommendation request, determine a first time period to which the reception time of the recommendation request belongs.
[0063] In some embodiments, the recommendation request may be a recommendation request sent by a client to a recommendation system, and the recommendation request is used to request information recommendation for multiple booths. The trigger of the recommendation request may include various situations. For example, the recommendation request may be sent when the client responds to a preset operation performed by the target user. When the recommendation system receives a recommendation request from the client, it means that information recommendation needs to be performed at at least one corresponding booth. As an example, the preset operation includes but is not limited to: payment operation, adding to the shopping cart operation, submitting an order operation, and confirming receipt of goods operation. This specification does not limit this.
[0064] In some embodiments, the recommendation system may divide a day into multiple time periods. Since the traffic conditions in adjacent time periods are basically the same. Therefore, within the current time period, the recommendation system can perform offline simulation on the traffic data of the previous time period to update the target recommendation strategy, so as to use the updated target recommendation strategy for recommendation when receiving a recommendation request in the next time period. The recommendation system may use the time period corresponding to the receiving moment of the recommendation request as the first time period.
[0065] In some embodiments, the duration of each time period may be 5 minutes, 10 minutes, 15 minutes, 20 minutes, etc. The duration of each time period can be determined according to the requirements in actual applications. This specification does not limit this.
[0066] Figure 4 Shows a timing diagram when applying a recommendation method provided according to an embodiment of this specification.
[0067] In some embodiments, referring to Figure 4 , the recommendation system includes an online recommendation module and a policy update module. During the t time period, the online recommendation module can process the received traffic data according to the target recommendation strategy of the t time period, and in the policy update module, according to the traffic data received during the t - 1 time period, iteratively update the target recommendation strategy of the t time period to obtain the target recommendation strategy of the t + 1 time period. Among them, the target recommendation strategy of the t time period is obtained by the policy update module iteratively updating the target recommendation strategy of the t - 1 time period according to the traffic data received during the t - 2 time period during the t - 1 time period.
[0068] Among them, when the receiving moment of the recommendation system for the recommendation request belongs to the t time period, it is determined that the t time period is the first time period, and the t - 2 time period is the second time period. Similarly, when the receiving moment of the recommendation system for the recommendation request belongs to the t + 1 time period, the t + 1 time period is the first time period, and the t - 1 time period is the second time period.
[0069] It should be noted that Figure 4The scenario shown is applicable to the case where the recommendation system has been running for at least two time periods after startup. In the first two time periods after the recommendation system starts up, recommendations can be made in the following ways:
[0070] In the first time period after the recommendation system starts up, since there is no previous time period, in this time period, the recommendation strategy used by the online recommendation module can be a recommendation strategy obtained through random initialization. And due to the lack of traffic data from the previous time period, the strategy update module cannot perform offline simulation in this time period. In the second time period after the recommendation system starts up, the online recommendation module still uses the recommendation strategy obtained through random initialization to make recommendations, and in the strategy update module, offline simulation is performed based on the traffic data from the first time period and the recommendation strategy obtained through random initialization to obtain the target recommendation strategy for the third time period.
[0071] S320: Obtain the target recommendation strategy adopted by the recommendation system in the first time period. The target recommendation strategy is a strategy determined by performing offline simulation on the traffic data in the second time period. The recommendation evaluation index obtained by performing offline simulation on the traffic data using the target recommendation strategy meets the preset conditions, and the amount of resources required does not exceed the total resources of the recommendation system. The second time period is before the first time period.
[0072] In some embodiments, referring to Figure 4 , the target recommendation strategy (the target recommendation strategy for the t-th time period) adopted by the recommendation system in the first time period (for example, the t-th time period) can be a strategy determined by performing offline simulation on the traffic data in the second time period (the t-2-th time period) in the t-1-th time period and updating the target recommendation strategy for the t-1-th time period.
[0073] Among them, the recommendation requests received in the first time period can be used to request information recommendations at multiple booths. In this case, the target recommendation strategy can include sub-recommendation strategies corresponding to each of the multiple booths. After the recommendation system executes the recommendation strategy, the recommendation action corresponding to the booth can be determined, that is, the recommended content to be displayed. As an example, assuming that the recall module and the fine-tuning module are used when the recommendation system makes recommendations, the recommendation strategy can represent the recall channels used in the recall module, the queue lengths of the recall channels, or the metric prediction models used in the fine-tuning module. A recommendation strategy can represent a permutation and combination of the recall channels, queue lengths, and metric prediction models.
[0074] In some embodiments, the target recommendation strategy is obtained in the following manner: obtaining traffic data within a second time period; obtaining an initial recommendation strategy and determining the initial distribution information corresponding to the initial recommendation strategy; adjusting the initial distribution information based on the traffic data within the second time period to obtain updated distribution information; and determining the target recommendation strategy based on the updated distribution information, wherein the recommendation evaluation index obtained by processing the traffic data within the second time period using the target recommendation strategy is not worse than the recommendation evaluation index obtained by processing the traffic data within the second time period using the initial recommendation strategy.
[0075] In some embodiments, offline simulation refers to the recommendation system simulating the behavior of processing historical traffic data in an offline environment based on the target recommendation strategy of the current time period or a pre-set recommendation strategy and predicting the corresponding recommendation evaluation index. The recommendation system can determine the target recommendation strategy to be used in the next time period according to the results of the offline simulation. Refer to Figure 4 , after receiving a recommendation request within the second time period, the online recommendation module of the recommendation system can record and save the received recommendation request as the traffic data within the second time period. In the first time period, the policy update module of the recommendation system can use the traffic data within the second time period recorded by the online recommendation module as historical traffic data for offline simulation to determine the target recommendation strategy to be used in the next time period.
[0076] In some embodiments, the recommendation system can perform multiple rounds of iteration on the initial distribution information based on the traffic data until a preset end condition is reached to obtain updated distribution information. For example, the recommendation system can perform multiple rounds of iteration on the initial distribution information based on the traffic data through the Cross Entropy Method (CEM), and the distribution information after the last round of iteration is the updated distribution information. The recommendation system can determine the target recommendation strategy based on the updated distribution information.
[0077] Figure 5 FIG. shows a schematic diagram of determining a target recommendation strategy through the cross entropy evolution algorithm according to an embodiment of the present specification.
[0078] In some embodiments, the initial recommendation strategy can be a recommendation strategy obtained by random initialization. As an example, in the first time period after the recommendation system is started, since there is no previous time period, within this time period, a recommendation strategy obtained by random initialization can be used. For example, the recommendation system can set the initial distribution information corresponding to the recommendation strategy to a standard normal distribution, that is, μ 0 in which all the means are 0, and σ 0 in which all the variances are 1; alternatively, the recommendation system can also set arbitrary parameters for the initial distribution information corresponding to the recommendation strategy, that is, μ 0All the means in it are randomly set, σ 0 and all the variances in it are also randomly set.
[0079] In some embodiments, the initial recommendation strategy is a strategy determined based on the traffic data in the third time period. The recommendation evaluation index obtained by performing recommendation simulation on the traffic data using the initial recommendation strategy meets the preset conditions, and the required resource consumption does not exceed the total resources of the recommendation system. The third time period is before the second time period.
[0080] As an example, referring to Figure 4 and Figure 5 , assume that the first time period is the t time period, and the target recommendation strategy in the t time period is determined by the recommendation system through offline simulation based on the initial recommendation strategy and the traffic data in the second time period during the t - 1 time period. When the recommendation system performs offline simulation during the t - 1 time period, the initial recommendation strategy is the target recommendation strategy in the t - 1 time period, and the traffic data in the second time period is the traffic data in the t - 2 time period. Among them, the initial recommendation strategy is determined based on the traffic data in the third time period (i.e., the t - 3 time period) during the t - 2 time period.
[0081] In some embodiments, the recommendation system can obtain the corresponding initial distribution information according to the initial recommendation strategy. The initial distribution information is used to represent the probability distribution of different recommendation actions being selected. As an example, the target recommendation strategy in the t - 1 time period may include sub - recommendation strategies corresponding to multiple booths, and each sub - recommendation strategy represents the recommendation action used for the corresponding booth. Obtaining the corresponding initial distribution information for the initial recommendation strategy may include the distribution information corresponding to each sub - recommendation strategy. For example, the distribution information corresponding to each sub - recommendation strategy can be represented by a normal distribution, that is, each sub - recommendation strategy follows a normal distribution, and the normal distribution followed by the sub - recommendation strategy can represent the probability of different recommendation actions being selected in the sub - recommendation strategy. Among them, the mean of the normal distribution can represent the expected value when selecting a recommendation action (i.e., the most likely recommended action to be selected), and the variance of the normal distribution can represent the degree of dispersion of multiple candidate recommendation actions.
[0082] In some embodiments, assume that the target recommendation strategy in the t - 1 time period includes N all sub - recommendation strategies corresponding to each booth, and N all is an integer greater than or equal to 1. Then, the mean (μ 0 ) and variance (σ 0 ) corresponding to the normal distribution followed by the initial distribution information can be represented by the following formula:
[0083]
[0084] Among them, represents the mean of the normal distribution followed by the sub - recommendation strategy corresponding to the 1st booth, Denotes the variance of the normal distribution followed by the sub-recommendation strategy corresponding to the 1st booth; Denotes the Nth all mean of the normal distribution followed by the sub-recommendation strategy corresponding to the booth, Denotes the Nth all variance of the normal distribution followed by the sub-recommendation strategy corresponding to the booth.
[0085] In some embodiments, when the recommendation system adjusts the initial distribution information based on the traffic data in the second time period, it may perform multiple rounds of iteration on the initial distribution information based on the traffic data in the second time period until a preset end condition is reached. Refer to Figure 5 , for any positive integer i, the i-th round of iteration process includes:
[0086] The recommendation system samples multiple sampling recommendation strategies based on the current distribution information. When i = 1, the current distribution information is the initial distribution information. When i > 1, the current distribution information is the distribution information updated in the (i - 1)-th round of iteration. The recommendation system performs offline simulation on the traffic data using multiple sampling recommendation strategies respectively, and obtains the recommendation evaluation indicators corresponding to the multiple sampling recommendation strategies respectively. The recommendation system updates the respective recommendation evaluation indicators based on the resource consumption amounts required by the multiple sampling recommendation strategies. Among them, when the resource consumption amount required by a sampling recommendation strategy is greater than the total resource amount, the updated recommendation evaluation indicator corresponding to this sampling recommendation strategy is less than the recommendation evaluation indicator before the update. And, the recommendation system determines the updated distribution information based on the updated recommendation evaluation indicators corresponding to the multiple sampling recommendation strategies.
[0087] In some embodiments, in the i-th round of iteration, the recommendation system may sample based on the current distribution information to obtain N sample sampling recommendation strategies, where N sample is an integer greater than or equal to 1. The N sample sampling recommendation strategies can be represented by the following formula:
[0088]
[0089] where, θ1 represents the first sampling recommendation strategy, denotes the Nth sample sampling recommendation strategy, and each sampling recommendation strategy follows the parameters N(μ i-1 , σ i-1 ) of the normal distribution corresponding to the current distribution information. As an example, in the 1st round of iteration, i = 1, then the current distribution information N(μ i-1 , σ i-1 ) is the above-mentioned initial distribution information N(μ 0 , σ 0 ).
[0090] In some embodiments, when sampling the current distribution information, it includes, but is not limited to, translating the current distribution information by a random distance to obtain the sampled distribution information; or scaling the current distribution information by a random multiple to obtain the sampled distribution information, and generating a sampled recommendation strategy according to the sampled distribution information. Then, the recommendation system obtains multiple sampled recommendation strategies according to multiple sampled distribution information.
[0091] In some embodiments, the recommendation system can traverse multiple sampled recommendation strategies and perform the following steps for the current sampled recommendation strategy during the traversal: input the traffic data into a pre-trained simulation system, and control the simulation system to sample and process the traffic data using the current sampled recommendation strategy to obtain the recommendation evaluation index corresponding to the current sampled recommendation strategy.
[0092] Among them, the pre-trained simulation system can, based on the traffic data, simulate the execution of the recommendation actions indicated by the recommendation strategy and predict the recommendation evaluation index after the recommendation strategy executes the recommendation actions. As an example, the simulation system can be a prediction model trained based on historical traffic, historical recommendation strategies, and historical recommendation evaluation indexes. The type and training method of the simulation system are not limited in this specification.
[0093] In some embodiments, referring to Figure 5 , the recommendation system can, through the simulation system, sequentially simulate the use of each sampled recommendation strategy to process the traffic data within the second time period, and predict and obtain the recommendation evaluation index corresponding to each sampled recommendation strategy.
[0094] In some embodiments, in order to ensure that there will be no problem of resource overload when the recommendation system makes recommendations based on the recommendation strategy, the total amount of resources of the recommendation system can be used as a constraint condition, and the recommendation evaluation indexes of each are updated based on the resource consumption amounts required by multiple sampled recommendation strategies.
[0095] In some embodiments, for any one of the multiple sampled recommendation strategies, the update process of the recommendation evaluation index corresponding to the sampled recommendation strategy adopts the following formula:
[0096] Value′(θ)=Value(θ)-λmin{C-Cost(θ),0}
[0097] Among them, Value(θ) is the recommendation evaluation index corresponding to the sampled recommendation strategy, Value′(θ) is the updated recommendation evaluation index, λ is a constant coefficient, C is the total amount of resources of the recommendation system, and Cost(θ) is the resource consumption amount required by the sampled recommendation strategy.
[0098] In some embodiments, when the amount of resources Cost(θ) consumed by the sampled recommendation strategy is less than or equal to the total amount of resources C of the recommendation system, C - Cost(θ) is a value greater than or equal to 0, and min{C - Cost(θ), 0} takes the value 0. In this case, Value′(θ) = Value(θ), that is, the updated recommendation evaluation index is the same as the recommendation evaluation index corresponding to the sampled recommendation strategy.
[0099] In some embodiments, when the amount of resources Cost(θ) consumed by the sampled recommendation strategy is greater than the total amount of resources C of the recommendation system, C - Cost(θ) is a value less than 0, and min{C - Cost(θ), 0} takes the value C - Cost(λ). And θ can be set to a maximum value. In this case, Value′(θ) is equal to Value(θ) minus a maximum value, that is, the updated recommendation evaluation index of the current sampled recommendation strategy will be much smaller than the recommendation evaluation index corresponding to the sampled recommendation strategy. That is to say, when the sampled recommendation strategy violates the constraint condition (that is, the amount of resources Cost(θ) consumed by the sampled recommendation strategy is greater than the total amount of resources C of the recommendation system), the updated recommendation evaluation index of this sampled recommendation strategy will be much smaller than that of other sampled recommendation strategies that do not violate the constraint condition.
[0100] In some embodiments, the recommendation system can sort multiple sampled recommendation strategies based on the updated recommendation evaluation indexes corresponding to them. For example, the recommendation system can sort the updated recommendation evaluation indexes corresponding to multiple sampled recommendation strategies in descending order. In the sorting result obtained by the recommendation system, the higher the ranking of the sampled recommendation strategy, the better the corresponding recommendation evaluation index. Among them, since the updated recommendation evaluation indexes of the sampled recommendation strategies that violate the constraint condition are extremely small, these sampled recommendation strategies that violate the constraint condition will be at the end in the sorting result.
[0101] In some embodiments, the recommendation system can ignore the sampled recommendation strategies at the end in the sorting result, or ignore the sampled recommendation strategies whose recommendation evaluation indexes are less than a preset threshold, so as to reduce the impact of the recommendation strategies whose consumed resources exceed the total amount of resources of the recommendation system on subsequent processing.
[0102] In this embodiment, when the amount of resources consumed by the sampled recommendation strategy is greater than the total amount of resources of the recommendation system, the recommendation system updates the recommendation evaluation index corresponding to the sampled recommendation strategy to an extremely small value, so that this sampled recommendation strategy is ignored in subsequent processing. It avoids the recommendation system selecting a recommendation strategy whose consumed resources exceed the total amount of resources of the recommendation system, so that the target recommendation strategy determined offline will not cause a situation of computational resource overload during execution.
[0103] In some embodiments, the recommendation system may determine updated distribution information based on multiple sampled recommendation strategies and the sorting results of the multiple sampled recommendation strategies. As an example, the recommendation system may generate updated distribution information based on the sorting results of the multiple sampled recommendation strategies through the top-k method. For example, the recommendation system may generate updated distribution information based on the sorting results of the multiple sampled recommendation strategies through the hard top-k method. Among them, the hard top-k method means that when the recommendation system generates updated distribution information, only the top k sampled recommendation strategies in the sorting results are used.
[0104] Alternatively, the recommendation system may generate updated distribution information based on the sorting results of the multiple sampled recommendation strategies using the soft top-k method. Among them, the soft top-k method may be the differential top-k method, that is, when the recommendation system generates updated distribution information, not only the top k sampled recommendation strategies in the sorting results are used, but also at least one other sampled recommendation strategy is considered.
[0105] In some embodiments, the recommendation system may generate updated distribution information through the hard top-k method based on the sorting results of the multiple sampled recommendation strategies, which may be implemented through the following steps: The recommendation system may determine the top K sampled recommendation strategies with the highest sorting in the multiple sampled recommendation strategies based on the sorting results of the multiple sampled recommendation strategies, where K is an integer greater than or equal to 1. Then, the recommendation system may determine the updated distribution information according to the distribution information of the K sampled recommendation strategies.
[0106] As an example, assume that K = 3 and the number of sampled recommendation strategies is 5. The 5 sampled recommendation strategies may be represented by the following vector:
[0107] [X1,X2,X3,X4,X5]
[0108] Among them, the recommendation system sorts the 5 sampled recommendation strategies in descending order according to the updated recommendation evaluation index corresponding to each sampled recommendation strategy. The vector corresponding to the 5 sorted sampled recommendation strategies is:
[0109] [X1,X3,X4,X2,X5]
[0110] K = 3, which means that the recommendation system uses the distribution information of the first 3 sampled recommendation strategies (i.e., X1, X3, X4) to determine the updated distribution information. The recommendation system may generate the updated distribution information through the following formula:
[0111] [X1,X2,X3,X4,X5]*[1,0,1,1,0] T / 3
[0112] This formula represents multiplying the vector [X1, X2, X3, X4, X5] by the transpose of the vector [1, 0, 1, 1, 0] and then dividing by 3, i.e.:
[0113]
[0114] In this case, when the recommendation system uses the hard top-k method to generate the updated distribution information, the updated distribution information is the average of the distribution information of K sampled recommendation strategies.
[0115] In some embodiments, the recommendation system can generate the updated distribution information through the differential top-k method based on the ranking results of multiple sampled recommendation strategies, which can be achieved through the following steps: The recommendation system can determine the weights corresponding to each of the multiple sampled recommendation strategies through the polytope mapping algorithm based on the multiple sampled recommendation strategies and the ranking results of the multiple sampled recommendation strategies. Among them, the weights corresponding to the top K sampled recommendation strategies are greater than the weights corresponding to the sampled recommendation strategies ranked later, and K is an integer greater than or equal to 1. Then, the recommendation system can determine the updated distribution information based on the multiple sampled recommendation strategies, the weights corresponding to each of the multiple sampled recommendation strategies, and K.
[0116] As an example, assume K = 3 and the number of sampled recommendation strategies is 5. The 5 sampled recommendation strategies can be represented by the following vectors:
[0117] [X1, X2, X3, X4, X5]
[0118] Among them, the 5 sampled recommendation strategies are sorted in descending order according to the updated recommendation evaluation index corresponding to each sampled recommendation strategy.
[0119] The vectors corresponding to the 5 sorted sampled recommendation strategies are:
[0120] [X1, X3, X4, X2, X5]
[0121] Then, the recommendation system can determine the weights corresponding to each sampled recommendation strategy through the polytope mapping algorithm based on the vectors corresponding to the 5 sorted sampled recommendation strategies and the value of K.
[0122] As an example, the weight corresponding to each sampled recommendation strategy can be calculated by the following formula:
[0123]
[0124] Among them, this formula represents the objective function of the Limited Multi-Label Projection Layer in the polytope mapping algorithm. In this formula, y *denotes the weight corresponding to the sampling recommendation strategy, x represents the input vector (i.e., the vector corresponding to the 5 sorted sampling recommendation strategies), y represents a sparse probability distribution that satisfies the constraints, and the purpose of this formula is to map the input vector x to the sparse probability distribution y that satisfies the constraints. In this formula, denotes the variable to be minimized subsequently, is the constraint condition, that is, the sum of the weights (y * ) corresponding to all sampling recommendation strategies must be equal to K. There are two variables in this formula, among which, is a linear term. In order to minimize , it is necessary to maximize the inner product of y and x, that is, when the value of x is larger (i.e., the updated recommendation evaluation index corresponding to the sampling recommendation strategy is larger and it is more forward in the sorting result), a higher weight is assigned to the sampling recommendation strategy corresponding to x. -H b (y) is a binary entropy regularization term, which is used to encourage the values corresponding to y to be neither close to 0 nor close to 1, but distributed in the middle region between 0 and 1 (i.e., y satisfies the sparse probability distribution of the constraints), so as to prevent the occurrence of gradient sparsity.
[0125] As an example, the weights corresponding to the vector [X1, X2, X3, X4, X5] can be [0.9, 0.4, 0.7, 0.6, 0.4], then the recommendation system can generate the updated distribution information through the following formula:
[0126] [X1, X2, X3, X4, X5] * [0.9, 0.4, 0.7, 0.6, 0.4] T / 3
[0127] This formula means multiplying the vector [X1, X2, X3, X4, X5] by the transpose of the vector [0.9, 0.4, 0.7, 0.6, 0.4] and then dividing by 3, that is:
[0128]
[0129] In this embodiment, the recommendation system determines a corresponding weight for each sampling recommendation strategy through the polytope mapping algorithm. The updated distribution information determined according to each sampling recommendation strategy and the corresponding weight not only considers the characteristics of the sampling recommendation strategies that are forward in the sorting result, but also takes into account the characteristics of other sampling recommendation strategies in the sorting result. The updated distribution information determined by the recommendation system takes into account the characteristics of each sampling recommendation strategy, can represent the characteristics of multiple sampling recommendation strategies more objectively and accurately, and the target recommendation strategy determined according to the updated distribution information can adapt to more scenarios and has better effects.
[0130] In some embodiments, after obtaining the updated distribution information, the recommendation system may determine whether a preset end condition is met according to the updated distribution information or the number of iterations. When the preset end condition is met, the recommendation system may determine the target recommendation strategy according to the updated distribution information. When the preset end condition is not met, the recommendation system may use the updated distribution information as the current distribution information and perform the next iteration.
[0131] In some embodiments, the preset end condition may be: the current number of iterations reaches a preset number, or the difference between the updated distribution information and the current distribution information is less than or equal to a preset difference. As an example, the preset number may be 100,000 times. Then, when the number of iterations is 100,000, the recommendation system determines that the preset end condition is met and ends the iteration. Or, when the number of iterations does not reach 100,000 times, if the difference between the updated distribution information and the current distribution information is less than or equal to the preset difference, the recommendation system may also determine that the preset end condition is met and end the iteration.
[0132] As an example, the difference between the updated distribution information and the current distribution information may be a difference value, Euclidean distance, etc. The preset difference may be determined according to the actually used difference and the actual accuracy requirement, and this specification does not limit this.
[0133] In some embodiments, the target recommendation strategy is a strategy determined by the recommendation system through offline simulation of the traffic data in the second time period during the fourth time period. The fourth time period is before the first time period and after the second time period. The recommendation system determines the target recommendation strategy based on the updated distribution information, including: the recommendation system determines a candidate recommendation strategy based on the updated distribution information, and the recommendation evaluation index obtained by processing the traffic data in the second time period by the candidate recommendation strategy is not worse than the recommendation evaluation index obtained by processing the traffic data in the second time period by using the initial recommendation strategy. Determining at least part of the traffic in the fourth time period as the target traffic data, the recommendation system may obtain the true recommendation evaluation index obtained by the recommendation system processing the target traffic data in the fourth time period. Then, the recommendation system may perform offline simulation on the target traffic data by using the candidate recommendation strategy to obtain the predicted recommendation evaluation index corresponding to the candidate recommendation strategy. Finally, when the predicted recommendation evaluation index is greater than or equal to the true recommendation evaluation index, the recommendation system determines the candidate recommendation strategy as the target recommendation strategy. Or, when the predicted recommendation evaluation index is less than the true recommendation evaluation index, the recommendation system uses the recommendation strategy adopted in the fourth time period as the target recommendation strategy.
[0134] In some embodiments, as an example, the fourth time period may be the duration required for the policy update module to perform offline simulation on the traffic data. Refer to Figure 4, assuming that the first time period is the t+1 period and the second time period is the t-1 period, then the fourth time period can be within the t period. For example, if the durations of the t-1 period, the t period, and the t+1 period are 15 minutes, and the duration required for the policy update module to perform offline simulation on the traffic data within the second time period is 10 minutes, then the fourth time period can be the first 10 minutes within the t period.
[0135] Continue to refer to Figure 4 , the online recommendation module will continuously receive recommendation requests within the 15 minutes corresponding to the t period, and the policy update module will complete the offline simulation of the traffic data within the second time period and obtain candidate recommendation policies within the first 10 minutes corresponding to the t period.
[0136] In this case, the recommendation system can use all the recommendation requests (or some of them) received within the first 10 minutes of the t period as the target traffic, and obtain the true recommendation evaluation indicators obtained by the online recommendation module processing the target traffic data. As an example, refer to Figure 4 , the target traffic can be all the recommendation requests received during the fourth time period. After the online recommendation module processes the received recommendation requests according to the target recommendation policy of the t period during the fourth time period, it can obtain the true sub-recommendation evaluation indicators corresponding to each recommendation request in the target traffic data. The true recommendation evaluation indicator obtained by the recommendation system processing the target traffic data during the fourth time period is the sum of the true sub-recommendation evaluation indicators corresponding to each recommendation request in the target traffic data.
[0137] In some embodiments, the method by which the recommendation system can perform offline simulation on the target traffic data using the candidate recommendation policy to obtain the predicted recommendation evaluation indicator corresponding to the candidate recommendation policy is similar to the method of obtaining the recommendation evaluation indicators corresponding to the sampled recommendation policies through the offline simulation system, and will not be elaborated here.
[0138] When the predicted recommendation evaluation indicator of the recommendation system is greater than or equal to the true recommendation evaluation indicator, the candidate recommendation policy is determined as the target recommendation policy. Alternatively, when the predicted recommendation evaluation indicator of the recommendation system is less than the true recommendation evaluation indicator, the recommendation policy adopted by the recommendation system during the fourth time period is used as the target recommendation policy.
[0139] In this embodiment, the recommendation system can verify candidate recommendation strategies. When the recommendation evaluation metrics obtained by offline simulation of the target traffic data based on the candidate recommendation strategies are better than the real recommendation evaluation metrics obtained by the online recommendation module when processing the target traffic data, the recommendation strategy adopted by the recommendation system in the fourth period is determined as the target recommendation strategy. Otherwise, the recommendation system can continue to use the recommendation strategy adopted by the online recommendation module as the target recommendation strategy. In this way, the recommendation system can always maintain the target recommendation strategy that is more suitable for the current scenario, and the recommendation evaluation metrics for recommendation according to the target recommendation strategy are better, thereby improving the utilization rate of traffic by the recommendation system and further improving the recommendation effect.
[0140] S330: Determine the recommendation action corresponding to the recommendation request based on the target recommendation strategy, and execute the recommendation action to respond to the recommendation request.
[0141] In some embodiments, the recommendation request is used to request information recommendation for multiple exhibition booths. The target recommendation strategy includes the recall channel, queue length, and metric prediction model corresponding to each exhibition booth. The recommendation system may include a recall module and a fine ranking module. The recommendation system can configure the recall channel used in the recall module corresponding to each exhibition booth, the queue length of the recall channel, and the metric prediction model used in the fine ranking module through the target recommendation strategy.
[0142] Among them, one recall channel can correspond to one recommended product, and different recommended products include their respective corresponding recommended content. Each recall channel can also correspond to multiple queue lengths, and the queue length is used to represent the number of recommended content recalled in the recall channel. The metric prediction model represents the model used for fine ranking of the recommended content in the recall channel. Different metric prediction models can estimate the evaluation metrics when displaying the recommended content in the exhibition booth through different algorithms or logics, and then determine the recommended content to be displayed in the exhibition booth.
[0143] In some embodiments, the recommendation system can, for each target exhibition booth among multiple exhibition booths: obtain the recommendation prediction data corresponding to the target exhibition booth, where the recommendation prediction data at least includes: the response time corresponding to the target exhibition booth, and the historical recommendation evaluation metrics corresponding to the target exhibition booth. Then, the recommendation system can determine the recommendation action corresponding to the recommendation request based on the recommendation prediction data corresponding to the target exhibition booth and the sub-recommendation strategy corresponding to the target exhibition booth.
[0144] Figure 6 Shows a schematic diagram of responding to a recommendation request according to a target recommendation strategy provided by an embodiment of the present specification.
[0145] Reference Figure 6,The recommendation system can determine the recommendation action corresponding to the recommendation request through the online recommendation module, and execute the recommendation action to respond to the recommendation request.
[0146] In some embodiments, reference Figure 6 , the recommendation request is used to request information recommendation for booth 1, booth 2, booth 3 and booth 4, then booth 1, booth 2, booth 3 and booth 4 can all be used as target booths. After receiving the recommendation request, the recommendation system can obtain the recommendation prediction data corresponding to booth 1, the recommendation prediction data corresponding to booth 2, the recommendation prediction data corresponding to booth 3 and the recommendation prediction data corresponding to booth 4 respectively.
[0147] In some embodiments, the recommendation prediction data corresponding to each booth may include the historical recommendation evaluation index corresponding to the target booth obtained according to the historical data stored in the recommendation system, and the response time determined according to the operation status of the recommendation system, wherein the response time refers to the time interval between the recommendation system receiving the recommendation request for the booth and executing the recommendation action corresponding to the booth. The recommendation prediction data may also include the user feature data corresponding to the client that initiates the recommendation request, the exposure rate of the booth, the click rate of the booth, etc.
[0148] In some embodiments, taking booth 1 as the target booth as an example, the sub-recommendation strategy corresponding to booth 1 includes a recall channel corresponding to booth 1, a queue length corresponding to booth 1, and an indicator estimation model corresponding to booth 1.
[0149] The recommendation system can input the recall channel corresponding to booth 1 into the recall module. The recall module can recall the corresponding recommended content based on the recall channel corresponding to booth 1. The number of recalled recommended content can be determined according to the queue length corresponding to booth 1. For example, when the queue length corresponding to booth 1 is 5, the recommendation system can recall 5 corresponding recommended contents based on the recall channel corresponding to booth 1 through the recall module. Then, the recommendation system can estimate the evaluation index of the recommended content in the recall channel based on the recommendation prediction data and the index estimation model corresponding to booth 1 through the fine ranking module, and determine the recommended action corresponding to booth 1 according to the estimated evaluation index. Among them, the recommended action corresponding to booth 1 may include the recommended content displayed in booth 1 and the display order and display duration of the recommended content.
[0150] In summary, in the recommendation method and recommendation system provided in this specification, the recommendation system can obtain a target recommendation strategy applied within the first time period by performing offline simulation on the traffic data in the second time period and updating the recommendation strategy. The recommendation system can make recommendations for recommendation requests received within the first time period based on the target recommendation strategy. During the process of offline determining the target recommendation strategy, the recommendation system uses the amount of resources required not exceeding the total amount of resources as a hard constraint, so that the target recommendation strategy determined offline will not cause a situation of computational resource overload during execution. Moreover, since the recommendation system uses different target recommendation strategies in different time periods, when determining the recommended content to be displayed at each booth, the recommendation system can adaptively adjust the recommendation strategy according to the scenario where the recommendation system is located. The target recommendation strategy is more suitable for the current scenario, and the recommendation evaluation index for making recommendations based on the target recommendation strategy is better. As a result, the utilization rate of traffic by the recommendation system is improved, and the recommendation effect is further enhanced.
[0151] On the other hand, this specification provides a computer-readable non-transitory storage medium storing at least one instruction set for content recommendation. When the at least one instruction set is executed by a processor, the at least one instruction set directs the processor to perform the steps of the recommendation method described in this specification. In some possible implementations, aspects of this specification can also be implemented in the form of a program product, which includes program code. When the program product runs on a computing system 200, the program code is used to cause the computing system 200 to perform the steps of the recommendation method described in this specification. The program product for implementing the above method can be a portable compact disc read-only memory (CD-ROM) including program code and can run on the computing system 200. However, the program product of this specification is not limited to this. In this specification, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system. The program product can be any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. The computer-readable storage medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above. The program code for performing the operations of this specification can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the computing system 200, partially on the computing system 200, executed as an independent software package, partially on the computing system 200 and partially on a remote computing device, or entirely on a remote computing device.
[0152] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require a particular order or a sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0153] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and is not necessarily limiting. Although not explicitly stated herein, those skilled in the art will understand that this specification is intended to encompass various reasonable changes, improvements, and modifications to the embodiments. These changes, improvements, and modifications are intended to be proposed by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0154] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean that the specific features, structures, or characteristics described in connection with that embodiment may be included in at least one embodiment of this specification. Thus, it should be emphasized and understood that two or more references to "an embodiment" or "one embodiment" or "alternative embodiments" in various parts of this specification do not necessarily all refer to the same embodiment. Additionally, the specific features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0155] It should be understood that in the foregoing description of the embodiments of this specification, for the purpose of helping to understand a feature, and for the purpose of simplifying this specification, this specification combines various features in a single embodiment, figure, or its description. However, this does not mean that the combination of these features is necessary, and those skilled in the art may well mark out some of the devices as separate embodiments for understanding when reading this specification. That is to say, the embodiments in this specification may also be understood as the integration of multiple sub - embodiments. And the content of each sub - embodiment is also valid when it has fewer features than all the features of a single foregoing disclosed embodiment.
[0156] Each patent, patent application, publication of patent application, and other materials cited herein, such as articles, books, specifications, publications, documents, items, etc., except for those that are inconsistent with or conflict with this document, or those that have a restrictive effect on the broadest scope of the claims, may be incorporated herein by reference and used for all purposes now or hereafter associated with this document. In addition, in the event of any inconsistency or conflict between the description, definition, and / or use of relevant terms in any material and the description, definition, and / or use of relevant terms in this document, the terms in this document shall prevail.
[0157] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A recommendation method, applied to a recommendation system, the method comprising: Responding to receiving a recommendation request, determining a first time period to which the receiving moment of the recommendation request belongs; Obtaining a target recommendation strategy adopted by the recommendation system within the first time period, where the target recommendation strategy is a strategy determined by offline simulation of traffic data within a second time period, and the recommendation evaluation index obtained by performing offline simulation on the traffic data using the target recommendation strategy meets a preset condition and the required resource consumption does not exceed the total resources of the recommendation system, and the second time period is before the first time period; And Based on the target recommendation strategy, determining a recommendation action corresponding to the recommendation request, and executing the recommendation action to respond to the recommendation request.
2. The method according to claim 1, wherein, The target recommendation strategy is obtained by the following method: Obtaining the traffic data within the second time period; Obtaining an initial recommendation strategy and determining initial distribution information corresponding to the initial recommendation strategy; Adjusting the initial distribution information based on the traffic data within the second time period to obtain updated distribution information; And Based on the updated distribution information, determining the target recommendation strategy, where the recommendation evaluation index obtained by processing the traffic data within the second time period using the target recommendation strategy is not worse than the recommendation evaluation index obtained by processing the traffic data within the second time period using the initial recommendation strategy.
3. The method according to claim 2, wherein The adjusting the initial distribution information based on the traffic data within the second time period to obtain updated distribution information includes: Performing multiple rounds of iteration on the initial distribution information based on the traffic data within the second time period until a preset end condition is reached. For any positive integer i, the i-th round of iteration process includes: Sampling multiple sampling recommendation strategies based on the current distribution information. When i = 1, the current distribution information is the initial distribution information. When i > 1, the current distribution information is the distribution information updated in the (i - 1)-th round of iteration; Performing offline simulation on the traffic data within the second time period using the multiple sampling recommendation strategies respectively to obtain recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies respectively; Updating the respective recommendation evaluation indexes based on the resource consumption required by the multiple sampling recommendation strategies. When the resource consumption required by a sampling recommendation strategy is greater than the total resources, the updated recommendation evaluation index corresponding to the sampling recommendation strategy is less than the recommendation evaluation index before the update; and Based on the updated recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies, determining the updated distribution information.
4. The method according to claim 3, wherein Based on the updated recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies, determining the updated distribution information includes: Sorting the multiple sampling recommendation strategies based on the updated recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies; and Based on the multiple sampling recommendation strategies and the sorting result of the multiple sampling recommendation strategies, determining the updated distribution information.
5. The method according to claim 4, wherein, Determining the updated distribution information based on the multiple sampling recommendation strategies and the sorting results of the multiple sampling recommendation strategies includes: Based on the sorting results of the multiple sampling recommendation strategies, determining the top K sampling recommendation strategies with the highest sorting among the multiple sampling recommendation strategies, where K is an integer greater than or equal to 1; and Determining the updated distribution information according to the distribution information of the K sampling recommendation strategies.
6. The method according to claim 4, wherein, Determining the updated distribution information based on the multiple sampling recommendation strategies and the sorting results of the multiple sampling recommendation strategies includes: Based on the multiple sampling recommendation strategies and the sorting results of the multiple sampling recommendation strategies, determining the weights corresponding to the multiple sampling recommendation strategies respectively through a polytope mapping algorithm, where the weights corresponding to the top K sampling recommendation strategies with the highest sorting are greater than the weights corresponding to the sampling recommendation strategies with lower sorting, and K is an integer greater than or equal to 1; and Determining the updated distribution information based on the multiple sampling recommendation strategies, the weights corresponding to the multiple sampling recommendation strategies respectively, and the K.
7. The method according to claim 3, wherein For any one of the multiple sampling recommendation strategies, the following formula is adopted for the update process of the recommendation evaluation index corresponding to the sampling recommendation strategy: Value′(θ)=Value(θ)-λmin{C-Cost(θ),0} Wherein, Value(θ) is the recommendation evaluation index corresponding to the sampling recommendation strategy, Value′(θ) is the updated recommendation evaluation index, λ is a constant coefficient, C is the total amount of resources of the recommendation system, and Cost(θ) is the amount of resources consumed by the sampling recommendation strategy.
8. The method according to claim 3, wherein The offline simulation of the traffic data by using the multiple sampling recommendation strategies respectively to obtain the recommendation evaluation indexes corresponding to the multiple sampling recommendation strategies respectively includes: Traversing the multiple sampling recommendation strategies, and performing the following steps for the current sampling recommendation strategy in the traversal: Inputting the traffic data into a pre-trained simulation system, and controlling the simulation system to process the traffic data by using the current sampling recommendation strategy to obtain the recommendation evaluation index corresponding to the current sampling recommendation strategy.
9. The method according to claim 3, wherein The preset end conditions include at least one of the following: The current iteration number reaches the preset number; or The difference between the updated distribution information and the current distribution information is less than or equal to the preset difference.
10. The method according to claim 2, wherein, The initial recommendation strategy is a strategy determined based on the traffic data in the third time period. The recommendation evaluation index obtained by performing recommendation simulation on the traffic data by using the initial recommendation strategy meets the preset conditions and the amount of resources consumed does not exceed the total amount of resources of the recommendation system. The third time period is before the second time period.
11. The method according to claim 2, wherein, The initial recommendation strategy is a recommendation strategy obtained by random initialization.
12. The method according to claim 2, wherein, The target recommendation strategy is a strategy determined by the recommendation system through offline simulation of the traffic data in the second time period in the fourth time period. The fourth time period is before the first time period and after the second time period. Determining the target recommendation strategy based on the updated distribution information includes: Determine a candidate recommendation strategy based on the updated distribution information, where the recommendation evaluation index obtained by processing the traffic data in the second period using the candidate recommendation strategy is not worse than the recommendation evaluation index obtained by processing the traffic data in the second period using the initial recommendation strategy; Determine at least part of the traffic in the fourth period as target traffic data, and obtain the true recommendation evaluation index obtained by the recommendation system processing the target traffic data in the fourth period; Perform offline simulation on the target traffic data using the candidate recommendation strategy to obtain the predicted recommendation evaluation index corresponding to the candidate recommendation strategy; and In the case where the predicted recommendation evaluation index is greater than or equal to the true recommendation evaluation index, determine the candidate recommendation strategy as the target recommendation strategy, or, in the case where the predicted recommendation evaluation index is less than the true recommendation evaluation index, use the recommendation strategy adopted by the recommendation system in the fourth period as the target recommendation strategy.
13. The method according to claim 1, wherein The recommendation system includes a recall module and a fine ranking module, and the target recommendation strategy includes at least one of the following: The recall channels used in the recall module and the queue lengths of the recall channels; or The metric prediction model used in the fine ranking module.
14. The method according to claim 1, wherein The recommendation request is used to request information recommendation for multiple exhibition booths, and the target recommendation strategy includes sub-recommendation strategies corresponding to each of the multiple exhibition booths.
15. The method according to claim 14, wherein, Determining the recommendation action corresponding to the recommendation request based on the target recommendation strategy includes: For each target exhibition booth among the multiple exhibition booths: Obtain the recommendation prediction data corresponding to the target exhibition booth, where the recommendation prediction data at least includes: the response time corresponding to the target exhibition booth and the historical recommendation evaluation index corresponding to the target exhibition booth; and Based on the recommendation prediction data corresponding to the target exhibition booth and the sub-recommendation strategy corresponding to the target exhibition booth, determine the recommendation action corresponding to the recommendation request.
16. A recommendation system, comprising: At least one storage medium storing at least one instruction set for performing data processing related to content recommendation; And At least one processor communicatively connected to the at least one storage medium, where when the recommendation system runs, the at least one processor reads the at least one instruction set and implements the method according to any one of claims 1-15 based on the instructions of the at least one instruction set.
17. A computer-readable non-volatile storage medium, wherein, At least one instruction set is stored in the computer-readable non-volatile storage medium, and when the at least one instruction set is executed by at least one processor, the method according to any one of claims 1-15 is implemented.