Information recommendation method, device and electronic device based on artificial intelligence

By obtaining and processing the expected items and uncertain items of the candidate recommendation information set, determining the diversity characteristics and recommendation index, the problem of insufficient information recommendation accuracy in the existing technology is solved, and higher recommendation accuracy and resource conservation are achieved.

CN113590929BActive Publication Date: 2025-10-14TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110120525.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-28
Publication Date
2025-10-14
Estimated Expiration
2041-01-28

AI Technical Summary

Technical Problem

In the existing technology, information recommendation based on click-through rate is difficult to effectively characterize the positive impact of cold start information on user behavior and the diversity of user interests, which affects the accuracy of information recommendation.

Method used

By obtaining the expected items and uncertain items of the information features of the candidate recommendation information set, aggregation processing is performed to obtain the upper confidence bound features, the diversity features and recommendation index are determined, and finally the candidate recommendation information set with the highest recommendation index is selected for recommendation.

Benefits of technology

It improves the accuracy of information recommendations, ensures wide information coverage, deeply mines information that users are interested in, avoids invalid recommendations, and saves server computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113590929B_ABST
    Figure CN113590929B_ABST
Patent Text Reader

Abstract

The application provides an information recommendation method and device based on artificial intelligence, an electronic device and a computer readable storage medium. The method comprises: obtaining a plurality of candidate recommendation information sets, determining an expected item and an uncertain item of information characteristics of each candidate recommendation information set; performing aggregation processing on the expected item and the uncertain item of each candidate recommendation information set to obtain an upper confidence limit feature of each candidate recommendation information set; determining a diversity feature corresponding to each candidate recommendation information set; determining a recommendation index corresponding to each candidate recommendation information set according to the upper confidence limit feature of each candidate recommendation information set and a constraint violation feature; and taking the candidate recommendation information set with the highest recommendation index as a to-be-recommended information set to perform a recommendation operation on the to-be-recommended information set. Through the application, the recommendation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to artificial intelligence technology, and in particular to an information recommendation method and device based on artificial intelligence, an electronic device, and a computer readable storage medium. BACKGROUND

[0002] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0003] Information recommendation is an important application of artificial intelligence. In related technologies, in order to improve the recommendation rate, the click rate and the like is predicted, and recommendation is performed based on the predicted click rate. However, the applicant has found in the implementation of the embodiments of the present application that it is difficult to effectively depict the positive influence of cold-start information on user behavior and depict user diversity interest based on only the click rate for recommendation, thereby affecting the accuracy of information recommendation. SUMMARY

[0004] The embodiments of the present application provide an information recommendation method and device based on artificial intelligence, an electronic device, and a computer readable storage medium, which can improve the recommendation accuracy.

[0005] The technical solutions of the embodiments of the present application are as follows:

[0006] The embodiments of the present application provide an information recommendation method based on artificial intelligence, comprising:

[0007] Obtaining a plurality of candidate recommendation information sets, determining the expected item and the uncertain item of the information feature of each candidate recommendation information set;

[0008] Aggregating the expected item and the uncertain item of each candidate recommendation information set to obtain the upper confidence limit feature of each candidate recommendation information set;

[0009] Determining the diversity feature corresponding to each candidate recommendation information set;

[0010] Determining the recommendation index corresponding to each candidate recommendation information set according to the upper confidence limit feature and the constraint violation feature of each candidate recommendation information set;

[0011] Taking the candidate recommendation information set with the highest recommendation index as the to-be-recommended information set to perform a recommendation operation on the to-be-recommended information set.

[0012] The embodiments of the present application provide an information recommendation device based on artificial intelligence, comprising:

[0013] An acquisition module, configured to acquire a plurality of candidate recommendation information sets and determine expected items and uncertain items of information features of each candidate recommendation information set;

[0014] an aggregation module, configured to aggregate the expected items and uncertain items of each candidate recommendation information set to obtain an upper confidence bound feature of each candidate recommendation information set;

[0015] A diversity module, configured to determine a diversity feature corresponding to each of the candidate recommendation information sets;

[0016] An index module, configured to determine a recommendation index corresponding to each candidate recommendation information set based on an upper confidence bound feature and a constraint violation feature of each candidate recommendation information set;

[0017] The recommendation module is configured to take the candidate recommendation information set with the highest recommendation index as the information set to be recommended, so as to perform a recommendation operation on the information set to be recommended.

[0018] In the above scheme, the acquisition module is also used to: perform at least one of the following processing to obtain multiple candidate recommendation information sets: obtain multiple candidate recommendation information sets based on a linear estimation function; obtain multiple candidate recommendation information sets based on a quadratic estimation function; obtain multiple candidate recommendation information sets through an action evaluation framework; obtain multiple candidate recommendation information sets by combining a soft attention mechanism and a hard attention mechanism; obtain multiple candidate recommendation information sets through a Bernoulli distribution.

[0019] In the above scheme, the acquisition module is also used to: perform mapping processing on the i-th column vector of the L column vectors of the unit matrix to obtain the mapping processing result corresponding to the i-th column vector; wherein, the L column vectors correspond one-to-one to L information; L is an integer greater than or equal to 2, and the value range of i satisfies 1≤i≤L; using the mapping processing result of the column vector of the corresponding information as the weight, weighted summation processing is performed on the action data of the L information to obtain a linear estimation function; wherein, the action data represents whether the corresponding information is selected or not; determine the action data of the L information while satisfying the following conditions: when the action data of the L information are substituted into the linear estimation function, the value of the linear estimation function is the maximizing convergence value; the action data of the L information represent that at least one of the selected information from the L information satisfies the diversity constraint; and the at least one selected information from the L information is composed into the candidate recommendation information set.

[0020] In the foregoing solution, the obtaining module is further configured to: perform mapping processing on an i-th column vector in L column vectors of the unit matrix to obtain a mapping processing result corresponding to the i-th column vector, and take the mapping processing result corresponding to the i-th column vector as a matrix element; perform summation processing on the i-th column vector and a j-th column vector in the L column vectors of the unit matrix, perform mapping processing on a summation processing result, and obtain a mapping processing result corresponding to the i-th column vector and the j-th column vector; wherein L is an integer greater than or equal to 2, the values of i and j satisfy 1≤i,j≤L, and i and j are different; perform average processing on the mapping processing result corresponding to the i-th column vector and the mapping processing result corresponding to the j-th column vector, and perform subtraction processing on the mapping processing result corresponding to the i-th column vector and the j-th column vector and the average processing result, to obtain a matrix element; construct a matrix according to the matrix element; perform multiplication processing on a transpose of an action data matrix corresponding to the L information, the matrix, and the action data matrix, to obtain a quadratic estimation function; wherein the action data matrix includes action data corresponding to the L information, and the action data represents corresponding selected or unselected; determine action data of the L information that simultaneously satisfies the following conditions: when the action data of the L information is substituted into the quadratic estimation function, a value of the quadratic estimation function is a maximized convergent value; and the action data of the L information represents that at least one selected information in the L information satisfies a diversity constraint; and group at least one selected information in the L information to form the candidate recommended information set.

[0021] In the foregoing solution, the obtaining module is further configured to: generate an action matrix with L column vectors through an action network in an action evaluation framework, and determine a candidate recommended information set corresponding to the action matrix; wherein column identifiers of the L column vectors correspond to L information one by one, L is an integer greater than or equal to 2, and a value of the column vector represents action data corresponding to the information; and perform the following processing on the action matrix any number of times: perform transposition processing on any two different column vectors in the L column vectors in the action matrix to obtain a new action matrix, and determine a candidate recommended information set corresponding to the new action matrix.

[0022] In the scheme, the obtaining module is further configured to: generate action data corresponding to each of the information through an action network in the action evaluation framework; sort the L information in descending order according to the action data of each of the information; update the action data of a plurality of information ranked in the front in the L information to one, and update the action data of other information to zero; wherein the other information is information other than the plurality of information ranked in the front in the L information; convert the updated action data of each of the information into a column vector corresponding to the information to obtain an action matrix with the L column vectors.

[0023] In the scheme, the obtaining module is further configured to: initialize the evaluation network and the action network of the action evaluation framework; perform K times of iteration processing on the action evaluation framework, and perform the following processing in each iteration processing: perform T rounds of update processing on the action network and the evaluation network of the action evaluation framework according to the trade-off coefficient of the expectation item and the diversity feature, and update the trade-off coefficient according to the result of the Tth round of update processing; wherein T and K are integers greater than or equal to 2; determine the action network obtained in the Kth iteration processing as the action network for generating the action matrix with the L column vectors.

[0024] In the scheme, the obtaining module is further configured to: perform T rounds of iteration processing on the action evaluation framework, and perform the following processing in each round of iteration processing: predict a candidate recommended information set sample through the action network, and obtain an expectation item and a diversity feature corresponding to the candidate recommended information set sample; determine a value function value corresponding to the candidate recommended information set sample through the evaluation network, and determine a comprehensive value corresponding to the candidate recommended information set sample according to the expectation item, the diversity feature, the trade-off coefficient, and the value function value; obtain an error between the comprehensive value and the value function value, and update parameters of the evaluation network according to a gradient item corresponding to the error; determine a punitive value function value corresponding to the candidate recommended information set sample according to the expectation item, the diversity feature, and the trade-off coefficient, and update parameters of the action network according to a gradient item corresponding to the punitive value function.

[0025] In the above scheme, the acquisition module is also used to: obtain local observation data corresponding to each information in L information, and encode the local observation data into observation features; determine at least one interactive information in the L information that has an interactive relationship with the i-th information based on the hard attention mechanism and combined with the observation features of each information; determine the interaction weight between each interactive information and the i-th information based on the soft attention mechanism, and determine the interaction features of all the interactive information corresponding to the i-th information based on the interaction weight; determine the policy prediction value corresponding to the i-th information through the policy network based on the observation features and interaction features of the i-th information; wherein L is an integer greater than or equal to 2, i is an integer whose value increases from 1, and the value range of i satisfies 1≤i≤L; obtain the candidate recommendation information set based on the policy prediction value of each information in the L information.

[0026] In the above scheme, the acquisition module is also used to: merge the observation features of the i-th information with the observation features of each other information different from the i-th information to obtain the merged features corresponding to each of the other information; map each of the merged features through a bidirectional long-short-term memory artificial neural network, and perform maximum likelihood processing on the mapping results to obtain the hard attention value corresponding to each of the other information; and determine the other information whose hard attention value is greater than the hard attention threshold as the interactive information among the L information that has an interactive relationship with the i-th information.

[0027] In the above scheme, the acquisition module is also used to: perform the following processing for each of the interactive information: obtain the i-th embedding feature of the i-th information, and linearly map the i-th embedding feature according to the query parameter of the soft attention mechanism to obtain the query feature corresponding to the i-th information; obtain the interactive embedding feature of the interactive information, and linearly map the interactive embedding feature according to the key parameter of the soft attention mechanism to obtain the key feature corresponding to the interactive information; determine a soft attention value that is exponentially positively correlated with the key feature, the query feature and the hard attention value as the interaction weight corresponding to the interactive information; and weight the observation feature of each of the interactive information according to the interaction weight corresponding to the interactive information to obtain the interactive features of all the interactive information for the i-th information.

[0028] In the above scheme, the acquisition module is further configured to perform any one of the following processing: acquire a plurality of information from the L information, wherein the corresponding strategy prediction value of the information is greater than a strategy prediction threshold value, and sample K sampling information from the plurality of information to constitute the candidate recommendation information set; sort the L information in descending order according to the information strategy prediction value, and acquire K information ranked in the front to constitute the candidate recommendation information set; wherein K is the number of recommendation information in the candidate recommendation information set.

[0029] In the above scheme, the acquisition module is further configured to: acquire a training sample set, wherein the training sample set includes N candidate recommendation information set samples corresponding to N rounds of historical recommendations, N is an integer greater than or equal to 2; divide the N rounds of historical recommendations to obtain a plurality of historical recommendation periods, wherein each historical recommendation period includes M rounds of historical recommendations, M is an integer greater than 1 and less than N; initialize an objective function, wherein the objective function is used to represent the maximization of the penalty value function value in the M rounds of historical recommendations, and the objective function includes a Bernoulli distribution corresponding to the qth historical recommendation period and a Bernoulli distribution corresponding to the q-1th historical recommendation period, q is an integer greater than or equal to 2; in each historical recommendation period, the following processing is performed: acquiring a Bernoulli distribution corresponding to the historical recommendation period, and generating a candidate recommendation information set corresponding to each round of historical recommendation according to the Bernoulli distribution; determine the penalty value function value corresponding to each candidate recommendation information set sample, and substitute it into the objective function to perform gradient descent processing of the objective function for the Bernoulli distribution corresponding to the qth historical recommendation period, to obtain the Bernoulli distribution corresponding to the q+1th historical recommendation period; generate a candidate recommendation information set based on the Bernoulli distribution of the last historical recommendation period.

[0030] In the above scheme, the acquisition module is further configured to: generate a new candidate recommendation information set according to a teacher-student mechanism and in combination with the acquired plurality of candidate recommendation information sets; or generate a new candidate recommendation information set according to a beta distribution sampling mechanism and in combination with the acquired plurality of candidate recommendation information sets.

[0031] In the above scheme, the acquisition module is also used to: obtain the expected items and diversity characteristics of each historical candidate recommendation information set to determine the penalty value function value corresponding to each of the historical candidate recommendation information sets, and determine the historical candidate recommendation information set with the highest corresponding penalty value function value as the teacher set, and determine each candidate recommendation information set as the student set; for any student set, perform at least one of the following processing: map the any student set and the teacher set according to the operator to obtain a new candidate recommendation information set, or map the any student set and another student set different from the any student set according to the operator to obtain a new candidate recommendation information set.

[0032] In the above scheme, the acquisition module is further used to: perform the following processing on each candidate recommendation information set: perform perturbation processing on the action data of each recommendation information in the candidate recommendation information set to obtain a perturbation value of each action data in the candidate recommendation information set; perform perturbation processing on the action data of other information to obtain a perturbation value of each other information, wherein the other information is information other than the recommendation information in L information, and L is an integer greater than or equal to 2; based on the perturbation value corresponding to each recommendation information, obtain the beta distribution corresponding to the recommendation information, and based on the perturbation value corresponding to each other information, obtain the beta distribution corresponding to the other information; sample from the beta distribution corresponding to the recommendation information to obtain sampled action data corresponding to each recommendation information, and sample from the beta distribution corresponding to the other information to obtain sampled action data corresponding to each other information; based on the sampled action data corresponding to each recommendation information and the sampled action data corresponding to each other information, perform mixed descending sorting on the other information and the recommended information, and obtain the top K pieces of information to form a new candidate recommendation information set; wherein K is the number of recommended information in the candidate recommendation information set.

[0033] In the above scheme, the acquisition module is also used to: forward propagate the information features of each candidate recommendation information set in the belief neural network to obtain the expected items corresponding to each candidate recommendation information set; obtain the gradient function of the belief neural network, and substitute the information features of each candidate recommendation information set into the gradient function to obtain the uncertain items corresponding to each candidate recommendation information set.

[0034] In the foregoing scheme, the diversity module is further configured to: perform multiple recommendation information extraction processes on each candidate recommendation information set to obtain multiple recommendation information subsets; wherein two recommendation information are extracted in each recommendation information extraction process, and each recommendation information subset includes the two recommendation information extracted in the corresponding recommendation information extraction process; obtain a total number of the recommendation information subsets and a number of the recommendation information subsets that do not satisfy the diversity constraint, determine a ratio between the number of the recommendation information subsets that do not satisfy the diversity constraint and the total number, and determine a diversity feature corresponding to the ratio.

[0035] An electronic device is provided in an embodiment of the application, and the electronic device comprises:

[0036] a memory configured to store executable instructions;

[0037] a processor configured to execute the executable instructions stored in the memory to implement the information recommendation method based on artificial intelligence provided in the embodiments of the application.

[0038] A computer readable storage medium is provided in an embodiment of the application, and the computer readable storage medium stores executable instructions, and the executable instructions are configured to be executed by a processor to implement the information recommendation method based on artificial intelligence provided in the embodiments of the application.

[0039] The embodiments of the application have the following beneficial effects:

[0040] Based on the information features of the candidate recommendation information set, the expected item and the uncertain item for predicting the recommendation revenue are described for the candidate recommendation information set, the contribution of the information features to the user behavior prediction is considered, and the information coverage range of the candidate recommendation information set is ensured to be wide through the diversity feature, so as to deeply mine the information interested by the user, ensure the information recommendation accuracy of subsequent information recommendation, effectively avoid invalid recommendation, and further save the computing resources related to the recommendation logic in the server. BRIEF DESCRIPTION OF DRAWINGS

[0041] Figures 1A-1B FIG. 1 is a structural schematic diagram of an information recommendation system based on artificial intelligence provided in an embodiment of the application;

[0042] Figure 2 FIG. 2 is a structural schematic diagram of an electronic device provided in an embodiment of the application;

[0043] Figures 3A-3D FIG. 3 is a flow schematic diagram of an information recommendation method based on artificial intelligence provided in an embodiment of the application;

[0044] Figure 4 FIG. 4 is an architecture schematic diagram of an information recommendation system based on artificial intelligence provided in an embodiment of the application;

[0045] Figures 5A-5B FIG. 1 is a schematic diagram of a model of an information recommendation system based on artificial intelligence according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be described in further detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without making creative labor fall within the scope of protection of the present application.

[0047] In the following description, "some embodiments" are related to a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0048] The related data collection and processing in the embodiments of the present application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0050] Before the embodiments of the present application are described in further detail, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations.

[0051] 1) expected item, the expected behavior feature is a historical average return value, for example, in a recommendation system, the historical click rate (historical average return value) of any information can be predicted to a certain extent.

[0052] 2) uncertain item, the uncertainty feature is an upper bound value representing the uncertainty of the historical average return value, for example, in a recommendation system, some information has a small number of exposures, so its historical click rate cannot accurately predict the real click rate of the information, and thus the expected behavior feature is modified by the uncertainty feature.

[0053] 3) upper confidence bound feature, the upper confidence bound feature is a value obtained by weighted calculation based on the expected item and the uncertain item, and is used to predict the positive return after performing the recommendation operation.

[0054] In the related art, a recommendation system is usually packaged as a slot machine problem to decide an optimal solution as the decision result of each round of recommendation. However, the applicant found during implementation of the present application that the reward feedback in the recommendation system mainly reflects reward feedback and constraint feedback, and then explained the recommendation decision problem of the recommendation system as an optimization problem with complex constraints and sparse nonlinear feedback. Currently, there is no related technical solution to solve the optimization problem with complex constraints and sparse nonlinear feedback. This problem aims to select K (K is an integer greater than or equal to 2) information from L (L is an integer greater than or equal to 2) information with unknown rewards in each round to recommend to the user, so as to maximize the reward of T (T is an integer greater than or equal to 2) rounds of interaction. In the related art, it is assumed that the feedback of each information will be received after each decision, rather than only the total feedback (sparse feedback) of the selected information. The total feedback (sparse feedback) is the sum of the feedback of each selected information. In the related art, a small number of constraints such as base constraints (the number of selected information in each round is a fixed value K) and knapsack constraints are applied to the selection of information, rather than complex constraints such as diversity constraints.

[0055] The embodiments of the present application provide a recommendation method and device based on artificial intelligence, electronic equipment and computer readable storage medium, which can consider the contribution of information features to user behavior prediction, and ensure wide information coverage of the candidate recommendation information set through diversity features, thereby improving the recommendation accuracy. The following describes an exemplary application of the electronic equipment provided by the embodiments of the present application. The electronic equipment provided by the embodiments of the present application can be implemented as a server. The following describes an exemplary application when the device is implemented as a server.

[0056] The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, and basic cloud computing services such as big data and artificial intelligence platforms. The server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.

[0057] Artificial intelligence cloud service, also commonly referred to as AIaaS (AI as a Service). It is a mainstream service mode of the current artificial intelligence platform. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service mode is similar to opening an AI theme mall: all developers can access and use one or more artificial intelligence services provided by the platform through API interfaces. Some experienced developers can also use the AI framework and AI infrastructure provided by the platform to deploy and operate their own exclusive cloud artificial intelligence services. In the recommendation method based on artificial intelligence provided in the embodiments of the present application, the AI framework and AI infrastructure provided by the artificial intelligence cloud service can be used to deploy and operate the recommendation system.

[0058] Referring to Figure 1A , Figure 1A is a schematic diagram of the architecture of the recommendation system based on artificial intelligence provided in the embodiments of the present application. The recommendation system can be used to support various information recommendation scenarios and search scenarios. The search scenario is a special recommendation scenario, i.e., a scenario in which recommendations are made in response to a search query input by a user. The recommendation system includes application scenarios such as recommending news, recommending products, and recommending videos. Depending on the application scenario, the information can be news, video articles, text, or information related to products (such as real products such as clothes and virtual products such as game props). During the use of the client by the user, in response to a recommendation request of the terminal 400, the server 200 can obtain a plurality of candidate recommendation information sets from the database 500, and obtain information features (including user dimensions, environmental dimensions, and information itself dimensions) of each candidate recommendation information set based on user behaviors, environmental data, and attribute data of the candidate recommendation information set. The server 200 determines a corresponding recommendation index based on the information features of each candidate recommendation information set. The server determines a candidate recommendation information set with the highest recommendation index as a to-be-recommended information set based on the corresponding recommendation index, and recommends the information in the to-be-recommended information set to the terminal 400.

[0059] The specific architecture of the recommendation system will be described below. Referring to Figure 1B , based on Figure 1A , Figure 1BFig. 1 is a schematic diagram of an architecture of a recommendation system based on artificial intelligence provided by an embodiment of the present application. In the recommendation system, a terminal 400 is connected to a server 200 through a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two. The server 200 can be abstracted as a server cluster, including a master server 200-1 and multiple slave servers 200-2, …, 200-7. The master server 200-1 estimates a non-linear feedback function h(.) to provide reward feedback for the slave servers. Meanwhile, the master server 200-1 constructs an evaluator of the diversity constraint violation degree of a candidate recommendation information set in combination with the diversity constraint violation degree to provide constraint feedback for the slave servers. The reward feedback included in the profit feedback for the slave servers by the master server is the actual value of the estimated non-linear feedback function h(.), so that the slave servers 200-2, …, 200-7 generate multiple candidate recommendation information sets based on the received profit feedback and send them to the master server 200-1. After receiving the candidate recommendation information sets satisfying the diversity constraint, the master server 200-1 decides the candidate recommendation information set with the highest recommendation index (in combination with the confidence feature and the diversity feature) to make a decision.

[0060] In some embodiments, the server 200 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud service, cloud database, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and big data and artificial intelligence platform. The terminal 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the embodiments of the present application.

[0061] Referring to Figure 2 , Figure 2 Fig. 2 is a schematic diagram of a structure of a server 200 applying a recommendation method based on artificial intelligence provided by an embodiment of the present application. Figure 2 The server 200 shown in Fig. 2 includes at least one processor 210, a memory 250, and at least one network interface 220. The various components in the server 200 are coupled together by a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between the components. The bus system 240 includes not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all kinds of buses are marked as the bus system 240 in Figure 2 .

[0062] The processor 210 can be an integrated circuit chip that has a processing capability of signals, such as a general purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc., wherein the general purpose processor can be a microprocessor or any conventional processor.

[0063] The memory 250 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical drives, etc. The memory 250 optionally includes one or more storage devices remotely located from the processor 210.

[0064] The memory 250 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0065] In some embodiments, the memory 250 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are exemplarily illustrated below.

[0066] The operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks.

[0067] The network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including Bluetooth, wireless fidelity (WiFi), and universal serial bus (USB), etc.

[0068] In some embodiments, the artificial intelligence-based recommendation device provided by the embodiments of the present application can be realized in a software manner, Figure 2 An artificial intelligence-based recommendation device 255 stored in the memory 250 is shown, which includes a plurality of modules, the modules can be software in the form of programs and plug-ins, including the following software modules: an acquisition module 2551, an aggregation module 2552, a diversity module 2553, an index module 2554, and a recommendation module 2555, these modules are logical, and thus can be combined or further split according to the functions implemented, the functions of each module will be described below.

[0069] The artificial intelligence-based recommendation method provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server provided in the embodiment of the present application.

[0070] See also Figure 3A , Figure 3A This is a flowchart of the artificial intelligence-based recommendation method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.

[0071] In step 101, a plurality of candidate recommendation information sets are obtained, and expected items and uncertain items of information features of each candidate recommendation information set are determined.

[0072] In some embodiments, see Figure 3B , Figure 3B This is a flowchart of the artificial intelligence-based recommendation method provided in the embodiment of the present application, which will be combined with Figure 3B The steps shown in FIG. 1 are used to illustrate that obtaining multiple candidate recommendation information sets in step 101 can be achieved by performing at least one of the following steps 1011-1015, that is, performing at least one of the following steps 1011-1015 to obtain multiple candidate recommendation information sets.

[0073] In step 1011 , multiple candidate recommendation information sets are obtained according to a linear estimation function.

[0074] In some embodiments, the above-mentioned acquisition of multiple candidate recommendation information sets based on the linear estimation function can be achieved through the following technical solutions: mapping processing is performed on the i-th column vector of the L column vectors of the unit matrix to obtain the mapping processing result corresponding to the i-th column vector; wherein, the L column vectors correspond one-to-one to the L information; L is an integer greater than or equal to 2, and the value range of i satisfies 1≤i≤L; using the mapping processing result of the column vector of the corresponding information as the weight, weighted summation processing is performed on the action data of the L information to obtain a linear estimation function; wherein, the action data represents whether the corresponding information is selected or not; action data that satisfies the following conditions while determining the L information: when the action data of the L information are substituted into the linear estimation function, the value of the linear estimation function is the maximizing convergence value; the action data of the L information represent that at least one selected information among the L information satisfies the diversity constraint; at least one selected information among the L information is composed into a candidate recommendation information set.

[0075] As an example, the nonlinear feedback function h' of the belief neural network is linearly estimated, and the identity matrix I L*L The L column vectors (L column vectors correspond to L information one by one) are input into the belief neural network to obtain L outputs {b i} iL =1 The process of processing the L column vectors by the confidence neural network is a mapping process, and the mapping result corresponding to the i-th column vector is b i i b i is the weight corresponding to the i-th information (the i-th information corresponds to the i-th column vector), and is calculated according to {b i} =1 The following linear integer programming problem is constructed, wherein x satisfies a diversity constraint, x represents action data corresponding to each information, for example, x = 1 represents being selected, and x = 0 represents not being selected, and see formulas (1)-(3):

[0076]

[0077] Wherein, L is the number of information, K is the number of information in the candidate recommended information set, x i is the action data of the i-th information, and the action data of 1 represents that the i-th information is selected into the candidate recommended information set, the above problem is solved by using an optimizer Gurobi, and the candidate recommended information set is obtained, the result obtained is K information selected from the L information, the action data corresponding to the K information is 1, and the diversity constraint is satisfied so that formula (1) is maximized and converges, and formula (1) is a linear estimation function.

[0078] In step 1012, a plurality of candidate recommended information sets are obtained according to the quadratic estimation function.

[0079] ​In some embodiments, the above-mentioned acquisition of multiple candidate recommendation information sets based on the quadratic estimation function can be implemented by the following technical solutions: mapping processing is performed on the i-th column vector among the L column vectors of the unit matrix to obtain the mapping processing result corresponding to the i-th column vector, and the mapping processing result corresponding to the i-th column vector is used as the matrix element; summing the i-th column vector and the j-th column vector among the L column vectors of the unit matrix, mapping the summation result to obtain the mapping processing result corresponding to the i-th column vector and the j-th column vector; wherein L is an integer greater than or equal to 2, the value range of i and j satisfies 1≤i, j≤L, and the values ​​of i and j are different; averaging the mapping processing result corresponding to the i-th column vector and the mapping processing result corresponding to the j-th column vector, and performing the mapping processing on the i-th column vector. The mapping processing results of the i-th column vector and the j-th column vector are subtracted from the average processing results to obtain matrix elements; a matrix is ​​constructed according to the matrix elements; the transpose of the action data matrix corresponding to L information and the matrix are multiplied with the action data matrix to obtain a quadratic estimation function; wherein the action data matrix includes action data corresponding to the L information one by one, and the action data represents whether the corresponding information is selected or not; the action data of the L information that meets the following conditions is determined: when the action data of the L information are substituted into the quadratic estimation function, the value of the quadratic estimation function is the value that maximizes convergence; the action data of the L information represent that at least one information selected from the L information meets the diversity constraint; at least one information selected from the L information is composed of a candidate recommendation information set.

[0080] As an example, the nonlinear feedback function h' of the belief neural network is estimated quadratically, and the identity matrix I L*L The L column vectors (L column vectors correspond to L information one by one) are input into the belief neural network to obtain L outputs That is, the mapping processing result corresponding to each column vector, and the mapping processing result of the corresponding column vector is used as the matrix element Q ii , for i∈[L], the matrix element Q ii =b i , assuming h′≈x T Qx,Q∈R L*L , Q=Q T , e i is the identity matrix I L*L The i-th column vector of the unit matrix is ​​summed with the i-th column vector and the j-th column vector of the L column vectors of the unit matrix, and the summation result is mapped, that is, Determine as the input of the belief neural network, and perform mapping processing through the belief neural network to obtain the output {O ij} i,j∈[L],i≠j As the mapping processing result corresponding to the i-th column vector and the j-th column vector, the mapping processing result b corresponding to the i-th column vectori a mapping processing result b corresponding to the jth column vector j performing average processing on the mapping processing results b corresponding to the ith column vector and the jth column vector ij performing subtraction processing on the average processing result to obtain a matrix element Q ij , that is, for i, j e [L], i≠j, Q ij = o ij -(b i +b j ) / 2; since , for i, j e [L], i≠j, let Q ij = o ij -(b i +b j ) / 2, after obtaining the matrix Q, the following quadratic integer programming problem is established, in which x satisfies the diversity constraint, see formulas (4)-(6):

[0081] max x x T Qx (4);

[0082]

[0083] where L is the number of information, K is the number of information in the candidate recommended information set, x i is the action data of the ith information, and the action data is 1, indicating that the ith information is selected into the candidate recommended information set. The above problem is solved by using the optimizer Gurobi to obtain the candidate recommended information set. The result obtained is K information selected from L information, and the action data of the K information is 1. The diversity constraint is satisfied so that formula (4) converges to maximum. Formula (4) is a quadratic estimation function.

[0084] In step 1013, a plurality of candidate recommended information sets are obtained through the action evaluation framework.

[0085] In some embodiments, the above obtaining a plurality of candidate recommended information sets through the action evaluation framework can be implemented by the following technical solution: generating an action matrix with L column vectors through the action network in the action evaluation framework, and determining a candidate recommended information set corresponding to the action matrix; wherein the column identifiers of the L column vectors one-to-one correspond to L information, L is an integer greater than or equal to 2, and the value of the column vector represents the action data of the corresponding information; performing the following processing on the action matrix any number of times: transposing any two different column vectors in the L column vectors in the action matrix to obtain a new action matrix, and determining a candidate recommended information set corresponding to the new action matrix.

[0086] As an example, the action evaluation framework is composed of an action network and an evaluation network, the action network makes a decision to obtain a candidate action data set according to the state feature and the information feature, the candidate action data set can be represented by an action matrix with L column vectors, the column identifiers of the L column vectors correspond to the L information one by one, L is an integer greater than or equal to 2, the value of the column vector represents the action data of the corresponding information, the action data represents the selected information in the candidate recommended information set, randomly transposing some two different components (any two different column vectors) of the action matrix to obtain a new action matrix, the values of any two different column vectors represent the action data of any two different information, for example, the action data corresponding to information i is 1, and the action data corresponding to information j is 0, after the transposition processing, the action data corresponding to information i is 0, and the action data corresponding to information j is 1, repeat the above random transposition operation multiple times to obtain multiple new action matrices, and multiple new candidate action data sets correspond to multiple new candidate recommended information sets.

[0087] As an example, after obtaining the multiple new candidate recommended information sets and the candidate recommended information set decided by the action network, the evaluation network can score these candidate recommended information sets to obtain the value function value corresponding to each candidate recommended information set, and select multiple candidate recommended information sets whose value function values exceed the value function threshold or are ranked in the top in descending order as the candidate recommended information set in step 101.

[0088] In some embodiments, the above-mentioned unit matrix with L column vectors generated by the action network in the action evaluation framework can be realized by the following technical solutions: generating action data corresponding to each information by the action network in the action evaluation framework; descendingly sorting the L information according to the action data of each information; updating the action data of the multiple information ranked in the top in the L information to one, and updating the action data of the other information to zero; wherein the other information is the information in the L information except the multiple information ranked in the top; converting the updated action data of each information into a column vector corresponding to the information to obtain a unit matrix with L column vectors.

[0089] As an example, the candidate action data set (corresponding to the candidate recommended information set) obtained by the action network according to the state feature and the information feature can be an original action data set not belonging to the action space, for example, the action data in the candidate action data set is an arbitrary real number value, rather than a pre-set value (for example, 0 and 1) of the action space, so that multiple candidate action data sets most similar to the original action data set are searched in the action space, and the action space is

[0090] A1 is 0, indicating that the first information is not selected; A1 is 1, indicating that the first information is selected. t , for example, PA t There are L column vectors in the equation, the value of the first column vector is 0.9, the value of the second column vector is 0.95, the value of the third column vector is 0.4, ..., the value of the Lth column vector is 0.75, for PA t The components are sorted in descending order, and the components (action data) ranked in the top K are set to 1 (the action data of the top multiple pieces of information in the L pieces of information are updated to one, and one is used as the action data of the corresponding information), and the other components are set to 0 (the action data of the other pieces of information are updated to zero, and zero is used as the action data of the corresponding information). K is the number of information in the candidate recommendation information set. The action data of each piece of information after the update is converted into a column vector of the corresponding information to obtain an action matrix with L column vectors.

[0091] In some embodiments, before the above-mentioned generation of a unit matrix with L column vectors through the action network in the action evaluation framework, it can be achieved through the following technical solutions: initialize the evaluation network and the action network of the action evaluation framework; perform K iterative processing on the action evaluation framework, and perform the following processing during each iterative processing: perform T rounds of update processing on the action network and the evaluation network of the action evaluation framework according to the trade-off coefficient between the expected items and the diversity characteristics, and update the trade-off coefficient according to the result of the T-th round of update processing; wherein T and K are both integers greater than or equal to 2; determine the action network obtained by the K-th iterative processing as the action network used to generate the action matrix with L column vectors.

[0092] As an example, before generating a unit matrix with L column vectors through the action network in the action evaluation framework, it is necessary to train the action evaluation framework and update the trade-off coefficient between the expectation item (which can be understood as reward feedback) and the diversity feature (which can be understood as constraint feedback) when training the action evaluation framework. The trade-off coefficient between reward feedback and constraint feedback is very important. It will affect the evaluation network's action value estimation accuracy for different candidate action data sets and the training stability. Therefore, it is of great significance to design a trade-off coefficient that is adaptively adjusted as the training progresses. In the tth round state s t Next take action a t Will get reward feedback r(s t ,a t ) and constraint feedback c(s t ,a t ), let the constraint function C(s t )=F(c(s t ,a t ),…,c(s N ,a N), N is the total number of recommended rounds, F function is self-defined according to different scenarios, μ is the distribution subject to the initial state, and the reward feedback is shown in formula (7):

[0093]

[0094] Wherein, S is the state space, π is the sampling basis of the candidate action data set, and the following problem is solved by reward constraint policy optimization, see formula (8):

[0095]

[0096] Wherein, γ t is the parameter of the tth round, r t is the reward feedback of the tth round, and μ(s) is the state feature, is the estimated reward feedback output by the evaluation network, and C(s) is the predicted constraint feedback of the candidate recommended information set recommended in each round.

[0097] The above formula (8) problem is solved by taking the Lagrange relaxation method, that is, the above formula (8) problem is converted into the following optimization problem, see formula (9):

[0098]

[0099] The optimization problem described in formula (9) is to solve θ to maximize Then fix θ to solve λ to minimize The process of solving θ is to update the network of the action network, and solving λ and solving θ are not in the same time dimension, so the double time dimension method is used to solve the optimization problem in formula (9). On the level of fast time dimension, the parameters of the action evaluation framework are always updated to maximize the benefit J R On the level of slow time dimension, the Lagrange multiplier is also slowly updated to maximize J C The final goal of the action evaluation framework is to find a saddle point (θ * (λ * ), λ * ), the variable weighting parameter λ and two evaluation networks are introduced in the reward constraint policy optimization, one evaluation network is responsible for fitting the return about the actual reward (estimated reward feedback), and the other is responsible for fitting the return of the actual constraint (estimated constraint feedback). Then the two are weighted by λ to obtain the action value function value, see formula (10):

[0100]

[0101] Wherein, r(s, a) is the estimated reward feedback output by the evaluation network for each round of recommendation, and c(s, a) is the estimated constraint feedback output by the evaluation network for each round of recommendation, is the value function value obtained by the evaluation network for multiple rounds of recommendation, is the reward feedback obtained by the evaluation network for multiple rounds of recommendation, is the constraint feedback obtained by the evaluation network for multiple rounds of recommendation.

[0102] The evaluation network, the action network and λ are updated in turn, and the learning rates (lr) of the three satisfy the following relationship: lr(λ) < lr(action network) < lr(evaluation network). The training process includes two time dimensions, i.e. two kinds of cycles. The large cycle is updated with the iteration number as the time dimension (updating λ), and the small cycle is updated with the recommendation round as the time dimension (updating the action network and the evaluation network). For each iteration, multiple rounds of recommendations are performed, i.e. the action network and the evaluation network are updated multiple times, and then λ is updated once. The action evaluation framework is processed for K times of iteration, and T rounds of recommendation are performed in each iteration process. The parameters of the action network and the evaluation network are updated in each round of recommendation. After T rounds of recommendation are completed, it is equivalent to completing an iteration. After completing an iteration, the weighting coefficient is updated. See formula (11):

[0103]

[0104] where λ k+1 is the parameter of the evaluation network updated after each iteration, λ k is the weighting coefficient before updating, Γ λ is a projection operator, Γ λ is set to be an operator that limits λ to the interval [0, λ max ]; is set to be the average constraint violation rate of the candidate recommendation information set corresponding to the π θ distribution in the last T rounds; α is set to be the upper bound of the constraint violation rate of the candidate recommendation information set, which needs to be determined according to the specific situation.

[0105] In some embodiments, the T-round updating process of the action network and the evaluation network of the action evaluation framework according to the trade-off coefficient of the expected item and the diversity feature can be implemented through the following technical scheme: T-round iteration is performed on the action evaluation framework, and the following processing is performed in each round of iteration: the candidate recommended information set sample is predicted through the action network, and the expected item and the diversity feature corresponding to the candidate recommended information set sample are obtained; the value function value corresponding to the candidate recommended information set sample is determined through the evaluation network, and the comprehensive value of the corresponding candidate recommended information set sample is determined according to the expected item, the diversity feature, the trade-off coefficient and the value function value; the error between the comprehensive value and the value function value is obtained, and the parameters of the evaluation network are updated according to the gradient term of the corresponding error; the penalty value function value of the corresponding candidate recommended information set sample is determined according to the expected item, the diversity feature and the trade-off coefficient, and the parameters of the action network are updated according to the gradient term of the corresponding penalty value function.

[0106] As an example, first input the actual constraint feedback c, the estimated feedback constraint C, the threshold α, the parameters of the evaluation network, the action network and the learning rate of λ, initialize the parameters θ of the action network, the parameters v of the evaluation network, the Lagrange multiplier λ, first according to the iteration number K to perform loop calculation, in each iteration process, T-round recommendation is performed, in the recommendation process, the candidate recommended information set (candidate action data set) is {a t}, the state feature after the recommendation is s t+1 , the actual constraint feedback is c t , the comprehensive value is determined according to the actual reward feedback r t , the actual constraint feedback c t and the value function value output by the evaluation network , see formula (12):

[0107]

[0108] Among them, is the comprehensive value, r t is the actual reward feedback, c t is the actual constraint feedback, γ is a parameter, is the value function value output by the action evaluation framework corresponding to the state feature s t .

[0109] Based on the determined comprehensive value, the parameters of the evaluation network are updated, and the parameters of the action network are updated, see formulas (13) and (14):

[0110]

[0111] Among them, v k+1is the parameter of the evaluation network updated after each recommendation, v k is the parameter of the evaluation network before updating, θ k is the parameter of the action network before updating, θ k+1 is the parameter of the action network updated after each recommendation. Γ θ is a projection operator, Γ θ is set to the identity operator.

[0112] In step 1014, a plurality of candidate recommendation information sets are obtained by combining the soft attention mechanism and the hard attention mechanism.

[0113] In some embodiments, the above-mentioned combination of the soft attention mechanism and the hard attention mechanism to obtain a plurality of candidate recommendation information sets can be implemented by the following technical solutions: obtaining local observation data corresponding to each information in L information, and encoding the local observation data into observation features; according to the hard attention mechanism, and combining the observation features of each information, at least one interaction information existing interaction relationship between the L information and the i information is determined; according to the soft attention mechanism, the interaction weight between each interaction information and the i information is determined, and the interaction feature corresponding to the i information of all interaction information is determined according to the interaction weight; according to the observation feature and the interaction feature of the i information, the strategy prediction value corresponding to the i information is determined through the strategy network; wherein L is an integer greater than or equal to 2, i is an integer starting from 1 and increasing, and the value range of i satisfies 1≤i≤L; according to the strategy prediction value of each information in the L information, a candidate recommendation information set is obtained.

[0114] As an example, L information can be set as L agents, where L is an integer greater than or equal to 2. In a multi-agent environment, there are interactions between a large number of agents. When implementing the embodiments of the present application, the applicant found that each agent does not need to interact with all agents all the time during the decision-making process, but only needs to interact with neighboring agents. Through the artificial intelligence-based information recommendation method provided by the embodiments of the present application, the interaction relationship between each two agents is modeled, that is, whether there is interaction between the two agents. If interaction exists, the importance of the interaction on the agent strategy is determined. The multi-agent system is modeled as a graph network, that is, a fully connected topology graph. Each node in the graph network represents an agent, that is, each information represents an agent. A node in the graph network and the edges between nodes represent the interaction relationship between two agents, that is, the interaction relationship between two information. Two attention mechanisms are used to infer the interaction mechanism between agents. First, the irrelevant interaction edges are determined by the hard attention mechanism. According to the hard attention mechanism and combined with the observation characteristics of each information, at least one interaction information that has an interaction relationship with the i-th information among the L information is determined, where i is an integer starting from 1 and the value range of i satisfies 1≤i≤L. Then, the soft attention mechanism is used to judge the importance weight of the interaction edge retained by the hard attention mechanism, and the interaction weight between each interaction information and the i-th information is determined according to the soft attention mechanism. The interaction characteristics of all interaction information corresponding to the i-th information are determined according to the interaction weight. That is, the interaction information of each information (agent) and the interaction weight corresponding to each interaction information are obtained through the hard attention mechanism and the soft attention mechanism.

[0115] As an example, the hard attention mechanism is implemented through observation features, which are obtained by encoding local observation data. For the i-th information (agent), its local observation data Encoded into observation features by the multi-layer perceptron The local observation data consists of the ratio of the information selected from the past to the current round, the average ratio of the information and the information that is not allowed to be concentrated to be selected at the same time, and the average benefit and standard deviation of the overall action in the round in which the information is selected. Based on the observation characteristics and interaction characteristics of the i-th information, the policy prediction value corresponding to the i-th information is determined through the policy network; based on the policy prediction value of each information in the L information, the candidate recommendation information set is obtained. A simplified graph is obtained through a two-stage graph attention network, in which each information is only connected to the information that needs to interact. The observation characteristics of the interactive information output by the soft attention mechanism are weighted to obtain the interaction feature x i Finally, the policy gradient algorithm is used to reinforce learning to obtain the strategy of each agent, a i =π(h i , x i) is the action data of the i-th information, π is the fully connected layer for the final action decision, where h i ,x i They represent the observation characteristics of the i-th information and the interaction characteristics of other information on the i-th information respectively.

[0116] In some embodiments, the above-mentioned determination of at least one interactive information among the L information that has an interactive relationship with the i-th information based on the hard attention mechanism and in combination with the observation characteristics of each information can be achieved through the following technical solutions: merging the observation characteristics of the i-th information with the observation characteristics of each other information different from the i-th information to obtain a merged feature corresponding to each other information; mapping each merged feature through a bidirectional long-short-term memory artificial neural network, and performing maximum likelihood processing on the mapping results to obtain a hard attention value corresponding to each other information; determining other information with a hard attention value greater than the hard attention threshold as the interactive information among the L information that has an interactive relationship with the i-th information.

[0117] As an example, a bidirectional long short-term memory artificial neural network is first used to implement a hard attention mechanism to determine whether there is an interactive relationship between agents. The observation feature of the i-th information is merged with the observation feature of each other information different from the i-th information to obtain the merged feature corresponding to each other information. For example, for the i-th information and the j-th information, the observation features of the i-th information and the j-th information are merged to obtain the merged feature of the i-th information corresponding to the j-th information (h i ,h j ), merge the features (h i ,h j ) Input the bidirectional long short-term memory artificial neural network to obtain the mapping processing result h i,j =f(Bi-LSTM(h i ,h j )), where f is a fully connected layer, and then the gumbel-softmax function (maximum likelihood processing) is used for maximum likelihood processing to obtain the hard attention value corresponding to each other information Get the real value between 0 and 1 of the edge between the i-th information and the j-th information. If If the value of the jth information is greater than the hard attention threshold, then the jth information is the interactive information with the ith information. All the interactive information with the ith information is obtained from the L information to obtain the subgraph G of the ith information. i .

[0118] In some embodiments, the above-mentioned determination of the interaction weight between each interaction information and the i-th information based on the soft attention mechanism, and determination of the interaction features of all interaction information for the i-th information based on the interaction weight can be achieved through the following technical solution: performing the following processing for each interaction information: obtaining the i-th embedding feature of the i-th information, and linearly mapping the i-th embedding feature according to the query parameter of the soft attention mechanism to obtain the query feature corresponding to the i-th information; obtaining the interaction embedding feature of the interaction information, and linearly mapping the interaction embedding feature according to the key parameter of the soft attention mechanism to obtain the key feature of the corresponding interaction information; determining a soft attention value that is exponentially positively correlated with the key feature, the query feature and the hard attention value as the interaction weight of the corresponding interaction information; weighting the observation feature of each interaction information according to the interaction weight of the corresponding interaction information to obtain the interaction feature of all interaction information for the i-th information.

[0119] As an example, we use the soft attention mechanism to learn the subgraph G i The interaction weight of each edge in G i The interaction weight of the edge between the i-th information and the j-th information is where e i and e j are the embedding features of the i-th information and the j-th information respectively, W k and W q are key linear mapping and query linear mapping respectively, W k e j Convert to a key vector, W q e i Convert to a query vector.

[0120] In some embodiments, the above-mentioned acquisition of a candidate recommendation information set based on the policy prediction value of each information in the L information can be achieved through the following technical solutions: performing any one of the following processing: obtaining multiple information whose corresponding policy prediction values ​​are greater than the policy prediction threshold from the L information, and sampling K sampled information from the multiple information to form a candidate recommendation information set; sorting the L information in descending order according to the policy prediction value of each information, and obtaining the top K information to form a candidate recommendation information set; wherein K is the number of recommended information in the candidate recommendation information set.

[0121] As an example, we use the policy gradient algorithm to reinforce learning to get the recommendation strategy for each information, a i =π(h i , x i ) is the action data of the i-th information, π is the fully connected layer for the final action decision, where h i ,x iThey represent the observation features of the i-th information and the interaction features of other information on the i-th information respectively. The fully connected layer that finally makes action decisions outputs a real number in the interval [0, 1] for each information as the strategy prediction value. To obtain the final decision of each round, three methods can be adopted: for the i-th information, if the output of the graph attention network on the i-th information is greater than 0.5, then the i-th information is selected; the real values ​​output by the graph attention network for all information are sorted in descending order, and the information corresponding to the top K real values ​​in the sorting is selected; it can also be assumed that a is output for information i i , calculate the sampling probability, see formula (15):

[0122] U~uniform(0,1),b i =a i -log(-logU) (15);

[0123] Calculate b by formula (15) i , b i Obey corresponds to a i Gumbel distribution, choose b i The information corresponding to the top K values ​​in the sorting is i The order of the top K values ​​in the graph is distinguished, and the ordered probabilities can be converted into unordered probabilities. That is, each new ranking is permuted K times, and the K different ordered probabilities are averaged to obtain the unordered probabilities. Since calculating K probabilities is computationally intensive, M permutations for each K piece of information can be randomly generated and the corresponding M probabilities are averaged. The probability of the final decision in each round is obtained by sampling the output of the graph attention network.

[0124] The graph attention network can be updated by minimizing the loss function of each round of recommendation. The loss function of the t-th round recommendation is -r t +λc t , where r t is the nonlinear feedback function based on the master server, h' is the reward feedback calculated by the graph attention network for the tth round of decision, c t is the constraint violation rate (constraint feedback) of the t-th round of decision making.

[0125] In step 1015 , multiple candidate recommendation information sets are obtained through Bernoulli distribution.

[0126] In some embodiments, the above-mentioned acquisition of multiple candidate recommendation information sets through Bernoulli distribution can be implemented through the following technical solutions: obtaining a training sample set, wherein the training sample set includes N candidate recommendation information set samples corresponding one-to-one to N rounds of historical recommendations, where N is an integer greater than or equal to 2; dividing the N rounds of historical recommendations to obtain multiple historical recommendation cycles, wherein each historical recommendation cycle includes M rounds of historical recommendations, where M is an integer greater than 1 and less than N; initializing the objective function, wherein the objective function is used to characterize the maximization of the penalty value function value in the M rounds of historical recommendations, and the objective function includes the Bernoulli distribution corresponding to the qth historical recommendation cycle. and the Bernoulli distribution corresponding to the q-1th historical recommendation cycle, where q is an integer greater than or equal to 2; in each historical recommendation cycle, perform the following processing: obtain the Bernoulli distribution corresponding to the historical recommendation cycle, and generate a sample of the candidate recommendation information set corresponding to each round of historical recommendation according to the Bernoulli distribution; determine the penalty value function value corresponding to each candidate recommendation information set sample, and substitute it into the objective function to perform gradient descent processing of the objective function for the Bernoulli distribution corresponding to the qth historical recommendation cycle, and obtain the Bernoulli distribution corresponding to the q+1th historical recommendation cycle; generate a candidate recommendation information set based on the Bernoulli distribution of the last historical recommendation cycle.

[0127] As an example, for any i∈[L],A t,i Whether it is 1 indicates whether the information i in round t is selected, A t,i Subject to the mean μ i Bernoulli distribution, for any Let the initial μ i is 0.5, Q∈{1,2,...,int(T / N)}, (int(T / N) is the lower integer of T / N, N is the number of recommended rounds of an epoch, and T is a positive integer), t∈{QN+1,...,min{(Q+1)N,T}}, on each component i∈[L], from the mean value μ i A is obtained by sampling from the Bernoulli distribution t,i , for t=min{(Q+1)N,T}, according to h'(.)-λc(.) is {A QN+1 ,...,A min{(Q+1)N,T}}, where h'(.) is the reward income from the master server, c(.) is the diversity constraint violation rate, and the vectors with the top ρ percent scores are averaged to obtain the new (μ1,μ2,...,μ L ) vector, so that (μ1,μ2,...,μ L) generates as high a score as possible, we need to have a more accurate understanding of the layout of the candidate recommendation information set, which requires that the length N of the training round should be as large as possible. However, a large value of N will lead to (μ1,μ2,...,μ L ) is updated slowly and converges slowly, which makes the cross entropy method unable to play an advantage in the online decision environment. Therefore, the number of training batches N is divided into multiple time segments of length M. At the critical point of each small segment, the proximal strategy is used to optimize the pair (μ1,μ2,...,μ L ) performs multi-step gradient descent processing. For j∈{1,2,...,int(N / M)}, when t=min{QN+(j+1)n,T}, use formula (16) as the objective function and adopt the descent gradient algorithm to update u new,i , see formula (16):

[0128]

[0129] Among them, formula (16) represents the value of the penalty value function in maximizing M rounds of historical recommendations, u old is the parameter followed by sampling in the M rounds closest to the time t = min{QN+(j+1)n,T}, u new is the parameter to be sampled in the next M rounds of recommendation starting from min{QN+(j+1)n+1,T} rounds. If A t,i =1, then P(A t,i |u i )=u i If A t,i =0, then P(A t,i |u i )=1-u i , update J every M rounds (a historical recommendation cycle) t , and J t About U new Do multiple steps of gradient descent to achieve timely parameter updates required for online decision making, and update u based on the last historical recommendation cycle. new,i , and according to the updated u new,i Construct Bernoulli distribution to sample and obtain candidate recommendation information set. When sampling, when u new,1 When it is 0.6, a random number can be generated for the first information. When the random number is not greater than 0.6, the action data representing the first information is 1. When the random number is greater than 0.6, the action data representing the first information is 0.

[0130] In some embodiments, see Figure 3C , Figure 3C This is a flowchart of the artificial intelligence-based recommendation method provided in the embodiment of the present application, which will be combined with Figure 3CThe illustrated steps are described, and the expected item and the uncertain item of the information feature of each candidate recommendation information set are determined in step 101, which can be implemented by steps 1016-1017.

[0131] In step 1016, the information feature of each candidate recommendation information set is forward propagated in the confidence neural network to obtain the expected item corresponding to each candidate recommendation information set.

[0132] In step 1017, the gradient function of the confidence neural network is obtained, and the information feature of each candidate recommendation information set is substituted into the gradient function to obtain the uncertain item corresponding to each candidate recommendation information set.

[0133] As an example, the confidence neural network calculates the upper confidence bound feature of the corresponding candidate action data set for each candidate action data set x (corresponding information feature), see formula (17) and formula (18):

[0134]

[0135] U x =h′(x;θ)+γVar (18);

[0136] Wherein, Var is the uncertain item corresponding to the candidate action data set x, h'(x; θ) is the expected item corresponding to the candidate action data set x, θ represents the parameters of the confidence neural network, Z -1 And γ are parameters. Based on the uncertain item and the expected item, the upper confidence bound feature U x Corresponding to the candidate action data set x is obtained, and the upper confidence bound feature (including reward feedback) and the diversity feature (constraint feedback used to represent the constraint violation degree) are aggregated by using the trade-off coefficient λ of the upper confidence bound feature U x And the diversity feature C(x) to obtain the recommendation index, and the candidate action data set with the highest recommendation index is taken as the final decision. The candidate action data set refers to the action set, which is usually identified by an L-dimensional vector. Each dimension of the vector is used to represent whether the information of the corresponding dimension is selected.

[0137] In some embodiments, the nonlinear feedback function is estimated according to the following update process to obtain the estimated nonlinear feedback function h'(x; θ). First, the confidence neural network is initialized, and for 1≤l W {i,j} ~N(0,4 / m), for L1, let W l =(W T -W T ), W {i}~N(0, 2 / m), in the tth round of recommendation, the candidate action data set is obtained, and the belief neural network calculates the upper confidence bound feature of the corresponding candidate action data set for each candidate action data set x (corresponding information feature), see formula (19) and formula (20):

[0138]

[0139] U t,x =h′(x;θ t-1 )+γ t-1 Var t (20);

[0140] Among them, Var t is the uncertainty term corresponding to the candidate action data set x, h′(x; θ t-1 ) is the expected term corresponding to the candidate action data set x, θ t-1 The nonlinear feedback function that characterizes the confidence neural network is obtained based on the previous t-1 rounds of recommendations. and γ t-1 It is the parameter obtained based on the previous t-1 rounds of recommendation, and the confidence bound feature U on the corresponding candidate action data set x is obtained based on the uncertainty and expectation items. t,x , and then determine the diversity feature C(x) of the candidate action data set x, and use the trade-off coefficient λ between the upper confidence bound feature (including reward feedback) and the diversity feature (constraint feedback used to characterize the degree of constraint violation) to calculate the upper confidence bound feature U t,x Aggregate with the diversity feature C(x) to obtain the recommendation index, and take the candidate action data set with the highest recommendation index as the final decision for the t-th round recommendation. The candidate action data set refers to the action set. When updating the parameters, for the parameter Z t , refer to formula (21) for update:

[0141] Z t =Z t-1 +g(x t θ t-1 )g(x t θ t-1 ) T / m (21);

[0142] Among them, Z t-1 is the parameter updated based on the previous t-1 rounds of recommendations, Z t is the parameter updated based on the recommendation in the tth round, g(x t θ t-1 ) is x t Substitute the nonlinear feedback function h′(x; θ t-1 ) to obtain the gradient.

[0143] In some embodiments, when updating parameters, for θ t , refer to formula (22) for update:

[0144]

[0145] Among them, starting from θ=θ0, the loss function L(θ) is gradient-dropped in J steps, θ t is obtained from the last iteration, i.e. θ (0) =θ0, Setting γ in the confidence neural network t It is 0.1 and can be adjusted according to different situations. The master server follows the above process of the belief neural network to continuously update the nonlinear feedback function h'(.). In each round of recommendation, the master server will collect the candidate recommendation information set provided by each slave server. The candidate recommendation information set can be represented by the candidate action data set. Assuming that the action data is 1, it represents that the information is selected, and the action data is 0, it represents that it is not selected, then there can be the following two candidate action data sets: (1, 1, 0) is used to represent the candidate recommendation information set including the first information and the second information, and (0, 1, 1) is used to represent the candidate recommendation information set including the second information and the third information.

[0146] In step 102, the expected items and uncertain items of each candidate recommendation information set are aggregated to obtain the upper confidence bound feature of each candidate recommendation information set.

[0147] As an example, see equation (23) and equation (24):

[0148]

[0149] U x =h′(x;θ)+γVar (24);

[0150] Among them, Var is the uncertainty term corresponding to the candidate action data set x, h′(x;θ) is the expectation term corresponding to the candidate action data set x, θ represents the parameters of the belief neural network, Z -1 , γ and m are parameters, Based on the uncertainty and expectation terms, the upper confidence bound feature U corresponding to the candidate action data set x is obtained x The expected term is used to characterize the expected benefit of the corresponding candidate recommendation information set, and the uncertainty term characterizes the uncertainty of the expected benefit. The upper confidence bound feature actually takes the uncertainty into account.

[0151] In step 103, the diversity feature corresponding to each candidate recommendation information set is determined.

[0152] In some embodiments, see Figure 3D, Figure 3D is a flowchart of a recommendation method based on artificial intelligence provided by an embodiment of the present application, which will be described in combination with Figure 3D The steps shown will be described. In step 103, the diversity feature corresponding to each candidate recommendation information set can be implemented through steps 1031-1032.

[0153] In step 1031, multiple recommendation information extraction processes are performed on each candidate recommendation information set, and multiple recommendation information subsets are obtained correspondingly.

[0154] As an example, two recommendation information are extracted in each recommendation information extraction process, and each recommendation information subset includes two recommendation information extracted in the corresponding recommendation information extraction process.

[0155] In step 1032, the total number of recommendation information subsets and the number of recommendation information subsets that do not satisfy the diversity constraint are obtained, the ratio between the number of recommendation information subsets that do not satisfy the diversity constraint and the total number is determined, and the diversity feature corresponding to the ratio is determined.

[0156] As an example, multiple recommendation information extraction processes are performed on each candidate recommendation information set, and multiple recommendation information subsets are obtained, two recommendation information are extracted in each recommendation information extraction process, and each recommendation information subset includes two recommendation information extracted in the corresponding recommendation information extraction process. For example, 10 information are included in the candidate recommendation information set, 10 out of 2 calculation is performed, and 45 recommendation information subsets can be obtained after multiple extraction. The total number of recommendation information subsets and the number of recommendation information subsets that do not satisfy the diversity constraint are obtained, the diversity constraint requires that the feature distance between the two information in the recommendation information subset be greater than the feature distance threshold, and it is assumed that the feature distance between the two information in 20 recommendation information subsets is greater than the feature distance threshold. The number of recommendation information subsets that do not satisfy the diversity constraint is 20, the ratio between the number of recommendation information subsets that do not satisfy the diversity constraint and the total number is determined, and the diversity feature corresponding to the ratio is determined.

[0157] In step 104, the recommendation index corresponding to each candidate recommendation information set is determined according to the upper confidence bound feature and the constraint violation feature of each candidate recommendation information set.

[0158] As an example, the upper confidence bound feature (including reward feedback) and the diversity feature (constraint feedback for representing the degree of constraint violation) are aggregated by using the trade-off coefficient λ of the upper confidence bound feature and the diversity feature to obtain the recommendation index.

[0159] In step 105, the candidate recommendation information set with the highest recommendation index is taken as the to-be-recommended information set to perform the recommendation operation on the to-be-recommended information set.

[0160] In some embodiments, a new candidate recommendation information set can be generated based on the teacher-student mechanism and in combination with the multiple candidate recommendation information sets obtained in step 101; or a new candidate recommendation information set can be generated based on the Beta distribution sampling mechanism and in combination with the multiple candidate recommendation information sets obtained in step 101.

[0161] In some embodiments, the above-mentioned generation of a new candidate recommendation information set based on the teacher-student mechanism and in combination with the obtained multiple candidate recommendation information sets can be achieved through the following technical solutions: obtaining the expected items and diversity characteristics of each historical candidate recommendation information set to determine the penalty value function value corresponding to each historical candidate recommendation information set, and determining the historical candidate recommendation information set with the highest corresponding penalty value function value as the teacher set, and determining each candidate recommendation information set as the student set; for any student set, performing at least one of the following processing: mapping any student set and the teacher set according to the operator to obtain a new candidate recommendation information set, or mapping any student set and another student set different from any student set according to the operator to obtain a new candidate recommendation information set.

[0162] As an example, the expected item and the diversity feature of each historical candidate recommendation information set are obtained to determine the corresponding penalty value function value (h'(.)-λc(.)) of each historical candidate recommendation information set, the historical candidate recommendation information set is the candidate recommendation information set determined to be the recommended information set before step 101 is performed, the historical candidate recommendation information set with the highest corresponding penalty value function value is determined as the teacher set, each candidate recommendation information set obtained through steps 1011-1015 is determined as the student set, interaction is performed between the teacher set and the student set, between the student set and the student set, a new candidate recommendation information set is generated, the teacher candidate recommendation information set is T, the student set is S, for any one student set, any one student set and the teacher set can be mapped according to the operator to obtain a new candidate recommendation information set, any one student set is selected from the S set as a candidate recommendation information set A of the student set, B=A+rand*(T-A) (mapping processing), rand is a random number in the interval [0, 1], the candidate recommendation information set A exists corresponding candidate action set A, for example {1, 0, 0, 0, 1}, indicating that the first information and the fifth information are selected, the candidate recommendation information set T (teacher set) exists corresponding candidate action set T, for example {0, 1, 1, 0, 0}, indicating that the second information and the third information are selected, the candidate action set B obtained after mapping is {0.8, 1.1, 1.5, 0, 0}.7}, sort the components of the mapping results, set the components in the top K positions to 1, and set the other components to 0, and obtain a new candidate recommendation information set B, which includes the second information and the third information. Repeat the above operation several times to obtain multiple new candidate recommendation information sets, or map any student set and another student set different from any student set according to the operator to obtain a new candidate recommendation information set. Select any two candidate recommendation information sets A and B as student sets from the S set, which correspond to the candidate action set A and the candidate action set B respectively. Determine the penalty value function value h'(A)-λc(A) for the candidate recommendation information set A, and determine the penalty value function value h'(B)-λc(B) for the candidate recommendation information set B. If h'(A)-λc(A) <h'(B)-λc(B),C=A+rand*(B-A),否则C=A+rand*(A-B),h'(A)-λc(A)<h'(B)-λc(B),通过算子C=A+rand*(B-A)进行映射,候选推荐信息集合A存在对应的候选动作集合A,例如{1,0,0,0,1},表征第一个信息和第五个信息被选择,候选推荐信息集合B存在对应的候选动作集合B,例如{0,1,1,0,0},表征第二个信息和第三个信息被选择,经过映射得到的候选动作集合C为{0.8,1.1,1.5,0,0.7},将映射处理结果的分量进行排序,并将排在前K位的分量设为1,其他分量设为0,得到新的候选推荐信息集合C,新的候选推荐信息集合中包括第二个信息和第三个信息。.

[0163] In some embodiments, the above-mentioned sampling mechanism according to the beta distribution, in combination with the obtained multiple candidate recommendation information sets, to generate a new candidate recommendation information set can be implemented through the following technical solutions: for each candidate recommendation information set, the following processing is performed: perturbation processing is performed on the action data of each recommendation information of the candidate recommendation information set to obtain a perturbation value of each action data of the candidate recommendation information set; perturbation processing is performed on the action data of other information to obtain a perturbation value of each other information, wherein the other information is information other than the recommendation information in the L information, and L is an integer greater than or equal to 2; based on the perturbation value corresponding to each recommendation information, a beta distribution corresponding to the recommendation information is obtained, and based on the perturbation value corresponding to each other information, a beta distribution corresponding to the other information is obtained; sampling is performed from the beta distribution corresponding to the recommendation information to obtain sampled action data corresponding to each recommendation information, and sampling is performed from the beta distribution corresponding to the other information to obtain sampled action data corresponding to each other information; based on the sampled action data corresponding to each recommendation information and the sampled action data corresponding to each other information, the other information and the recommendation information are mixed and sorted in descending order, and the top K information is obtained to form a new candidate recommendation information set; wherein K is the number of recommendation information in the candidate recommendation information set.

[0164] As an example, perturbation processing is performed on the action data of each recommendation information of the candidate recommendation information set to obtain a perturbation value of each action data of the candidate recommendation information set (for the recommended information), perturbation processing is performed on the action data of other information to obtain a perturbation value of each other information (for the information not recommended in the L information), wherein the other information is information other than the recommendation information in the L information, and L is an integer greater than or equal to 2, a candidate action set A corresponding to the candidate recommendation information set is obtained (for the action data of the L information), the candidate action set A can be represented by a 0-1 integer value vector A, the component corresponding to 0 (action data) represents that the corresponding information is not selected, and the component corresponding to 1 (action data) represents that the corresponding information is selected, each component (each action data) of the integer value vector A is perturbed and the perturbation value is taken as the parameter of the beta distribution, based on the perturbation value corresponding to each recommendation information, a beta distribution corresponding to the recommendation information is obtained, and based on the perturbation value corresponding to each other information, a beta distribution corresponding to the other information is obtained, for example, a [0, 1] interval real value vector B ∈ [0, 1] L For i ∈ [L], if A i = 1, let B i = 1-τ; otherwise, let B i = τ, to (B i , 1-B i) is a parameter of the Beta distribution, sampling from the Beta distribution corresponding to the recommended information to obtain the sampled action data corresponding to each recommended information, and sampling from the Beta distribution corresponding to other information to obtain the sampled action data corresponding to each other information, and randomly sampling a real value C from the Beta distribution corresponding to each information i , based on C1, C2,..., C L , a vector C is formed, based on the sampled action data (real value) corresponding to each recommended information, the sampled action data (real value) corresponding to each other information, the other information and the recommended information are mixedly sorted in descending order, and the top K information is obtained to form a new candidate recommended information set; wherein K is the number of recommended information in the candidate recommended information set, the components (real values) in C are sorted, the top K components are set to 1 (representing action data as 1, selected), and the other components are set to 0 (representing action data as 0, not selected), to obtain a new candidate recommended information set C.

[0165] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0166] Taking a recommendation system as an example, the recommendation system selects K information (K is an integer greater than or equal to 2) from L information (L is an integer greater than or equal to 2) to recommend to the terminal of the user in each round of interaction with the user, receives the total feedback of the K information from the user after performing the recommendation operation, and the total revenue obtained after the user is interacted with for multiple rounds in the evaluation system of the recommendation system is expected to be as large as possible, which can be decomposed into the total feedback received after each round of recommendation operation being as large as possible, and the diversity of each information participating in each round of recommendation being as large as possible, which represents that the Euclidean distance between the features of each information participating in each round of recommendation is greater than the Euclidean distance threshold.

[0167] In some embodiments, the nature and constraint form of the recommendation system optimization problem are first converted, and then the master-slave server architecture is used to estimate the optimal solution of the nonlinear feedback function and the diversity sampling.

[0168] In some embodiments, in the recommendation system, selecting K information from L information in each round of recommendation is a combination tiger machine problem, and in each round of selection process, the action data of each round of selection for the L information is represented by a binary vector A t =(A t,1 ,A t,2 ,...,A t,L )∈{0,1} L , wherein A t,L is the action data of the Lth information in the tth round of recommendation, and the action data is 1 or 0, which are used to represent selection and non-selection respectively, and the total reward of each round of recommendation is At a nonlinear feedback function of A t ) plus a noise term that is Gaussian distributed, i.e., h(A t ) + ε t , where the noise term ε t is not necessary, h(A t ) is a nonlinear feedback function of A t , by representing the action data as a binary vector, the multi-armed bandit problem is converted into a contextual multi-armed bandit problem, for which a contextual upper confidence bound algorithm can be used to estimate the nonlinear feedback function h(·) while balancing the exploration of unknown data and exploitation of known data.

[0169] In some embodiments, a complex and non-differentiable diversity constraint in a recommendation system is transformed into a differentiable linear constraint, the diversity constraint requires that the distance between the feature vectors of each two information in each round of recommendation is greater than a Euclidean distance threshold, which involves a multi-dimensional operation and a non-differentiable process of taking absolute value and then truncation, in order to make the diversity constraint differentiable, according to the diversity constraint that the distance between the feature vectors of each two information is greater than the Euclidean distance threshold, an unallowable information set is constructed for each information, if there exist information i and information j that are simultaneously in each other's unallowable information set (i.e., the distance between the feature vectors of information i and information j is less than or equal to the Euclidean distance threshold), for information i and information j that are in each other's unallowable information set, the sum of Ai and Aj is greater than 1, which means that information i and information j are simultaneously recommended, which violates the diversity constraint, i.e., for any A t = (A t,1 , A t,2 ,..., A t,L ), the sum of A t,i i and A t,j j is less than or equal to 1 as much as possible, which is a simple linear constraint that is differentiable.

[0170] In some embodiments, referring to Figure 4 , Figure 4is a master-slave server architecture diagram of an information recommendation method based on artificial intelligence provided by the embodiment of the application. The master server is mainly used for estimating a nonlinear feedback function h(.), providing a surrogate reward (Surrogate Reward) as a benefit feedback for the slave server, wherein the benefit feedback comprises a reward feedback h(.) and a constraint feedback. The reward feedback is an actual value of the estimated nonlinear feedback function, and an evaluator (Sample Evaluato r) of a candidate recommendation information set is constructed by combining the reward feedback with a diversity constraint violation degree. The connection line on the left side of the master server and the slave server is unidirectional, the arrow representing the surrogate reward is downward, which is used to represent that the master server sends the reward feedback to the slave server, the arrow representing the evaluator (Sample Evaluato r) is downward, which is used to represent that the master server sends the constraint feedback to the slave server. The connection line on the right side of the master server and the slave server is unidirectional, the slave server sends a candidate recommendation information set satisfying the diversity constraint to the master server, so that the master server selects a candidate recommendation information set with the highest recommendation index from the multiple candidate recommendation information sets for decision-making.

[0171] In some embodiments, the slave server receives the benefit feedback provided by the master server, and generates a candidate recommendation information set for the master server to make a decision. The slave server comprises six servers that are complementary to each other: a solution strategy sampler (hereinafter referred to as a first slave server) composed of an optimization solver Gurobi; a Wolpertinger sampler (hereinafter referred to as a second slave server), which combines the primal-dual method and the Wolpertinger strategy for solving reinforcement learning methods with a large-scale discrete action space; a graph attention network sampler (hereinafter referred to as a third slave server): each information is regarded as an agent and is taken as a node of the graph attention network, and the connection between different information in the nonlinear feedback function is inferred by means of the forward propagation of the graph attention network, so as to make a comprehensive decision for each information; an improved cross-entropy method-enhanced deep learning evolution sampler (hereinafter referred to as a fourth slave server): which combines the evolutionary strategy cross-entropy algorithm and the classic reinforcement learning algorithm (proximal policy optimization), and weighs the diversity constraint and the nonlinear feedback function, and selects a candidate recommendation information set with better comprehensive performance; a random sampler (hereinafter referred to as a fifth slave server): randomly selects multiple vectors satisfying the cardinality constraint from the set {0,1} L ; a teacher-student sampler (hereinafter referred to as a sixth slave server): taking the historical best candidate recommendation information set as a teacher, and the candidate recommendation information sets provided by other slave servers as students, the teacher-student and student-student interact to generate a new candidate recommendation information set.

[0172] In some embodiments, the master server is configured to implement a contextual upper confidence bound algorithm to solve a contextual bandit problem with nonlinear feedback, a linear upper confidence bound algorithm is mainly used to estimate a feedback function of a linear feedback bandit problem, and the contextual upper confidence bound algorithm is based on the linear upper confidence bound algorithm and combines a deep neural network to estimate a nonlinear feedback function. In the information recommendation method based on artificial intelligence provided by the embodiments of the present application, the contextual upper confidence bound algorithm is used to solve a combined bandit problem with complex constraints, and the Wolpertinger strategy can be used to solve a reinforcement learning problem with a large-scale discrete action space. The Wolpertinger strategy generates a set of original action data that may not be feasible through an action evaluation framework when the action space is huge and complex constraints are required, and then searches for a solution with the maximum action value function in the set of Y closest actionable actions from the original action data set. The effect is stable and is not prone to local optimization. The operations optimization solver Gurobi can solve a linear integer programming problem with many linear constraints. The nonlinear feedback function is linearized or approximated to the second order, so that when the nonlinear feedback function is a quadratic function or a more complex function, the operations optimization solver Gurobi can generate an optimal solution.

[0173] The information recommendation method based on artificial intelligence provided by the embodiments of the present application can solve an online combined bandit problem with a sparse and unknown form feedback function and many constraints on a candidate recommendation information set. The method can accurately estimate the feedback function while considering sampling and optimization, intelligently search a huge action space, and balance the constraints and the real feedback in the reward term to approach the real optimal solution and recommend to the user, thereby improving the recommendation accuracy.

[0174] The specific problem applied in the recommendation system is as follows: K information is selected from L information with unknown returns in each round to maximize the total return of T rounds of interaction. The individual return of each information in each round follows a sub-Gaussian distribution with different expectations and variances. The total return of the selected K information is a nonlinear function of the individual return of each information plus a sub-Gaussian noise. After the agent makes a selection action in each round, it will only receive the total return of the K information as sparse feedback, and cannot obtain the individual return of each information. In addition, the information selected by the agent in each round satisfies the diversity constraint, that is, the Euclidean distance between the feature vectors of each two selected information should be greater than a certain threshold. The evaluation index of the above problem consists of two parts: the total return R T of T rounds of selection (the sum of the return feedback U t,x of each round of recommendation); and the cumulative constraint violation rate C T of T rounds. When there are M diversity constraints, the number of violated diversity constraints when making a decision in the tth round is n t , then The comprehensive evaluation index of multi-round recommendation is R T -λCT , the value of λ is determined by the specific situation.

[0175] In some embodiments, the frequency of interaction between the master server and the slave servers is as follows: the belief neural network in the master server runs once in each round; the first slave server, the fifth slave server, and the sixth slave server run once in each round; when the number of rounds that have been run is greater than the round threshold (for example, the round threshold is 4 times the number of information, i.e., 4L), the Wolpertinger sampler, the graph attention network sampler, and the improved cross entropy-deep reinforcement learning evolution sampler can be run 20 times every 5 rounds, and the running interval and number of times can be flexibly adjusted according to the specific situation and training conditions.

[0176] In some embodiments, the master server is mainly composed of a belief neural network, which is a nonlinearization of the belief algorithm on a linear context. It uses an L1 layer perceptron h'(.) to estimate the nonlinear feedback function h(.). Where m is the parameter of the belief neural network, for example, 4, 8, 16, σ(x) = max{x,0}, W1∈R m*L ,W l ∈R m*m , W l is the perceptron corresponding to the lth layer in the L1 layer perceptron, 2≤l≤L1-1,W L1 ∈R m*1 ,θ=[vec(W1) T ,...,vec(W L1 ) T )] T ∈R p ,p=m+mL+m 2 (L1-1), in the tth round of recommendation, θ is based on the data (A1, r1), (A2, r2), ..., (A t-1 ,r t-1 ) trained, r i It is the actual feedback received by the recommendation master server after executing the recommendation of the corresponding action data in the i-th round.

[0177] In some embodiments, the gradient of the nonlinear feedback function h(.) is The update process of the belief neural network is as follows: first, the belief neural network is initialized, and θ0 is initialized (θ0 is a parameter of the neural network). For 1≤l <L1,令 W {i,j} ~N(0,4 / m), for L1, let W l =(W T -W T ), W {i}~N(0, 2 / m), at the t-th round of recommendation, the confidence neural network obtains a set of candidate action data, calculates an upper confidence bound of the corresponding candidate action data set x for each candidate action data set x, see formula (1) and formula (2):

[0178]

[0179] U t,x = h'(x; θ t-1 ) + γ t-1 Var t (26);

[0180] Wherein, Var t is the uncertainty term corresponding to the candidate action data set x, h'(x; θ t-1 ) is the expected term corresponding to the candidate action data set x, θ t-1 characterizes the nonlinear feedback function of the confidence neural network is based on the t-1 rounds of recommendations obtained, and γ t-1 is the parameter based on the t-1 rounds of recommendations, based on the uncertainty term and the expected term, the upper confidence bound feature U t,x corresponding to the candidate action data set x is obtained, and the diversity feature C(x) of the candidate action data set x is determined, and the upper confidence bound feature (including reward feedback) and the diversity feature (constraint feedback used to represent the constraint violation degree) are aggregated by using the weighting coefficient λ of the upper confidence bound feature U t,x and the diversity feature C(x), to obtain a recommendation index, and the candidate action data set with the highest recommendation index is taken as the final decision of the t-th round of recommendation. The candidate action data set refers to an action set, which is usually identified by an L-dimensional vector. Each dimension of the vector is used to represent whether the information of the corresponding dimension is selected.

[0181] In some embodiments, when updating the parameters, for the parameter Z t , see formula (27) for updating:

[0182] Z t = Z t-1 + g(x t ; θ t-1 )g(x t ; θ t-1 ) T / m (27);

[0183] Wherein, Z t-1 is the parameter updated based on the t-1 rounds of recommendations, Z t is the parameter updated based on the t-th round of recommendations, g(x t ; θ t-1 ) is x tthe gradient obtained by substituting the nonlinear feedback function h'(x; θ t-1 ) into the gradient.

[0184] In some embodiments, when updating the parameters, for θ t , see formula (28) for updates:

[0185]

[0186] where, starting from θ = θ0, J-step gradient descent is performed on the loss function L(θ), θ t is the result of the last iteration, i.e., θ (0) = θ0, θ t = θ (j) , set γ t to 0.1 in the confidence neural network, which can be adjusted according to different situations, and the master server continuously updates the nonlinear feedback function h'(.) according to the above process of the confidence neural network. In each round of recommendation process, the master server will collect the candidate recommendation information set provided by each slave server, which can be represented by the candidate action data set. Assuming that action data of 1 represents that the information is selected, and action data of 0 represents that the information is not selected, there can be two candidate action data sets: (1, 1, 0) representing a candidate recommendation information set including the first information and the second information, and (0, 1, 1) representing a candidate recommendation information set including the second information and the third information, and U t,x -λC(x) instead of U t,x to evaluate each candidate recommendation information set (candidate action) to make a decision.

[0187] The following will be described in detail. When given a linear or quadratic objective function with known parameters and linearized diversity constraints, the Gurobi optimizer can output the optimal solution and the better solution that satisfies the constraints, respectively. Since the form of the nonlinear feedback function of the confidence neural network is unknown, if the Gurobi optimizer is used for solving, it may be necessary to linearize or approximate the nonlinear feedback function h' of the confidence neural network.

[0188] In some embodiments, the nonlinear feedback function h' of the confidence neural network is linearly estimated by the following technical solution: input the L column vectors of I L*L into the confidence neural network to obtain L outputs I is an identity matrix, and the L column vectors correspond to L information, respectively. According to the following linear integer programming problem is constructed, where x satisfies the diversity constraints, see formulas (29)-(31):

[0189]

[0190] where L is the number of information, K is the number of information of candidate recommended information set, x i is the action data of the i-th information, the action data of 1 represents that the i-th information is selected into the candidate recommended information set, the above problem is solved by using the optimizer Gurobi to obtain the candidate recommended information set, the solved result is K information selected from L information, the action data of K information is 1, and the diversity constraint is satisfied to make formula (29) reach maximum convergence, formula (29) is a linear estimation function.

[0191] In some embodiments, the nonlinear feedback function h' of the confidence neural network is twice estimated, assuming h'≈x T Qx, Q∈R L*L , Q=Q T , let e i be the i-th column vector of I L*L , and input into the confidence neural network to obtain the output {O ij} i,j∈[L],i≠j , for i,j∈[L], i≠j, let Q ij =O ij -(b i +b j ) / 2; for i∈[L], Q ii =b i , since , for i,j∈[L], i≠j, let Q ij =o ij -(b i +b j ) / 2. After obtaining the matrix Q, the following quadratic integer programming problem is established, and x in the quadratic integer programming problem satisfies the diversity constraint, see formulas (32)-(34):

[0192]

[0193] where L is the number of information, K is the number of information of candidate recommended information set, x i is the action data of the i-th information, the action data of 1 represents that the i-th information is selected into the candidate recommended information set, the above problem is solved by using the optimizer Gurobi to obtain the candidate recommended information set, the solved result is K information selected from L information, the action data of K information is 1, and the diversity constraint is satisfied to make formula (32) reach maximum convergence, formula (32) is a quadratic estimation function, and the maximum number of iterations is limited to 600-10000.

[0194] The following is a detailed introduction to the Wolpertinger sampler (hereinafter referred to as the second slave server). The Wolpertinger strategy is based on the action evaluation framework to train parameters through deep deterministic policy gradients. In each round of recommendation, the action network decision obtains an original action data set that may not belong to the action space. Therefore, it searches for Y (Y is an integer greater than or equal to 2) candidate action data sets that are closest to the original action data set in the action space, and obtains the action data (Q value) of the Y candidate action data sets through the evaluation network, and takes the candidate action data set with the largest action data as the recommendation decision for the tth round. In the recommendation system, let the action space be A1 is 0, indicating that the first information is not selected; A1 is 1, indicating that the first information is selected. In the process of the tth round of recommendation, the original action data set PA is calculated. t Afterwards, PA t Sort the components in descending order, set the value of the top K components to 1, and the values ​​of the other components to 0 to obtain a candidate action data set. Randomly swap two components with different values ​​in the candidate action data set to obtain a new candidate action data set. Repeat the above random swap operation Y-1 times to obtain Y-1 new candidate action data sets. The candidate action data set with the largest action value function value in all candidate action data sets is collected as the decision result of the recommendation in the tth round.

[0195] When diversity constraints are imposed on the candidate action data set, the benefit feedback needs to be adjusted. The benefit feedback needs to include reward feedback and constraint feedback. The trade-off coefficient between reward feedback and constraint feedback is very important. It will affect the accuracy of the evaluation network's action value estimation for different candidate action data sets and the training stability. Therefore, it is important to design a trade-off coefficient that can be adaptively adjusted as the training progresses. Training is performed through reward constraint strategy optimization. Suppose that in the tth round state s t Next take action a t Will get reward feedback r(s t ,a t ) and constraint feedback c(s t ,a t ), let the constraint function C(s t )=F(c(s t ,a t ),…,c(s N ,a N )), N is the total number of recommendation rounds, the F function is customized according to different situations, μ is the distribution obeyed by the initial state, and the initialized reward feedback is shown in formula (35):

[0196]

[0197] where S is the state space, π is the sampling basis of the candidate action data set, and the following problem is solved by reward-constrained policy optimization, see equation (36):

[0198] (36);

[0199] where, γ t is the parameter of the t-th round, r t is the reward feedback of the t-th round, μ(s) is the state feature, is the estimated reward feedback output by the evaluation network, and C(s) is the predicted constraint feedback of the candidate recommendation information set recommended in each round.

[0200] In some embodiments, the above-mentioned equation (36) problem is solved by taking the Lagrange relaxation method, i.e., the above-mentioned equation (36) problem is converted into the following optimization problem, see equation (37):

[0201]

[0202] The optimization problem described in equation (9) is to solve θ to maximize and then fix θ to solve λ to minimize The process of solving θ is to update the network of the action network, and solving λ is not in the same time dimension as solving θ, so the double time dimension method is used to solve the optimization problem in equation (9). On the level of fast time dimension, the parameters of the action evaluation framework are always updated to maximize the benefit J R , and on the level of slow time dimension, the Lagrange multiplier is also slowly updated to maximize J C The final goal of the action evaluation framework is to find a saddle point (θ * (λ * ), λ * ), and the variable weighting parameter λ and two evaluation networks are introduced in the reward-constrained policy optimization. One evaluation network is responsible for fitting the return about the actual reward (estimated reward feedback), and the other is responsible for fitting the return of the actual constraint (estimated constraint feedback). Then the two are weighted by λ to obtain the action value function value, see equations (38) and (39):

[0203]

[0204] where, is the estimated benefit feedback output by the evaluation network for each round of recommendation, r(s, a) is the estimated reward feedback output by the evaluation network for each round of recommendation, and c(s, a) is the estimated constraint feedback output by the evaluation network for each round of recommendation, is the value function value output by the evaluation network for multiple rounds of recommendations, It is the reward feedback obtained by evaluating the network output for multiple rounds of recommendations. It is the constrained feedback obtained from the evaluation network output for multiple rounds of recommendations.

[0205] The evaluation network, action network, and λ are updated sequentially through reward constraint strategy optimization. The learning rates (lr) of the three satisfy the following relationship: lr(λ) <lr(动作网络)<lr(评价网络),训练过程中包括两个时间维度,即包括两种循环,大循环是以迭代次数作为时间维度进行更新(更新λ),小循环是以推荐轮次作为时间维度进行更新(更新动作网络和评价网络),针对每次迭代,会进行多轮次推荐,即多次更新动作网络和评价网络后更新一次λ,针对动作评价框架进行K次迭代处理,并在每次迭代处理过程中进行T轮推荐,每轮推荐的过程中更新动作网络与评价网络的参数,在完成T轮推荐之后,则相当于完成了一次迭代,在完成一次迭代之后更新权衡系数,首先输入实际约束反馈c、反馈约束C、阈值α、评价网络、动作网络以及λ的学习率,初始化动作网络的参数θ,评价网络的参数v,拉格朗日乘子λ,首先依据迭代次数K进行循环计算,在每次迭代过程中,进行t轮推荐,完成了一次迭代,在完成一次迭代之后更新权衡系数的过程可以参见公式(40):

[0206]

[0207] Among them, λ k+1 is the parameter of the evaluation network updated after each iteration, λ k is the trade-off coefficient before updating, Γ λ is the projection operator, Γ λ Set to constrain λ to [0, λ max ] interval operators; Set to correspond to π θ The average constraint violation rate corresponding to the distributed candidate recommendation information set in the last T rounds; α is set as the upper bound of the constraint violation rate of the candidate recommendation information set, which needs to be determined according to the specific situation.

[0208] In the recommendation process, the candidate recommendation information set (candidate action data set) is a t , the state feature after executing the recommendation is s t+1 , the actual constraint feedback is c t , the actual reward feedback, actual constraint feedback and the value function value output by the evaluation network are used to determine the comprehensive value, see formula (41):

[0209]

[0210] in, is the comprehensive value, r t is the actual reward feedback, c t is the actual constraint feedback, γ is the parameter, is the corresponding state feature s t The value function value output by the action evaluation framework.

[0211] Based on the determined comprehensive value, the parameters of the evaluation network are updated, and the parameters of the action network are updated, see formulas (42) and (43):

[0212]

[0213] Among them, v k+1 is the parameter of the evaluation network updated after each recommendation, v k is the parameter of the evaluation network before updating, θ k is the parameter of the action network before updating, θ k+1 is the parameter of the action network updated after each recommendation. θ is the projection operator, Γ θ Set to the identity operator.

[0214] The initialization of the action network and the evaluation network follows the settings in the action-evaluation framework algorithm. The reward-constrained strategy optimization does not require the magnitude of the benefits, because the automatic update of λ enables the reward-constrained strategy optimization to adaptively perform feedback corrections. λ only changes in the reward-constrained strategy optimization algorithm. λ remains unchanged in the master server and other slave servers, and remains the trade-off coefficient between reward feedback and constraint feedback.

[0215] In the Wolpertinger sampler, the action network generates a set of raw action data, and recommends the final decision a to the tth round based on the raw action data set. t The conversion process follows the setting of the Wolpertinger strategy. In the update of the action evaluation framework, h' provided by the master server is used to calculate the candidate action data set a t The reward feedback r t .

[0216] The following will introduce the graph attention network sampler (hereinafter referred to as the third slave server) in detail. In a multi-agent environment, the complex game relationship between a large number of agents makes it very difficult to learn strategies. In addition, in the decision-making process, each agent does not need to interact with all agents at all times, but only needs to interact with neighbor agents. In the related art, it is only possible to determine which agents interact with each other through prior knowledge. When the system is very complex, it is very difficult to define the interaction based on rules. The information recommendation method based on artificial intelligence provided by the embodiments of the present application models the interaction relationship between each two agents, that is, judges whether the two agents exist interaction, and if the interaction exists, judges the importance of the interaction to the agent strategy.

[0217] In some embodiments, the multi-agent system is modeled as a graph network, that is, a fully connected topological graph. Each node in the graph represents an agent, and the edge between the nodes represents the interaction relationship between the two agents. Two attention mechanisms are used to infer the interaction mechanism between agents: hard attention mechanism: aims to disconnect irrelevant interaction edges. The hard attention mechanism is obtained by sampling and is not differentiable. The hard attention mechanism is improved to enable end-to-end learning. Soft attention mechanism: judges the importance weight of the interaction edge retained by the hard attention mechanism.

[0218] The graph attention network combines the above two attention mechanisms and reinforcement learning algorithms such as reinforcement learning or action evaluation framework to apply to the learning of multi-agent strategies. Referring to Figures 5A-5B , Figures 5A-5B is a model schematic diagram of the information recommendation system based on artificial intelligence provided by the embodiments of the present application, considering a locally observable environment. For the i-th agent, its local observation (O1, O2, O N-1 , O N ) is encoded into a feature vector (h1, h2, h N-1 , h N ) by a multi-layer perception. The multi-layer perception can be a long short-term memory artificial neural network (LSTM, Long Short-Term Memory). First, the hard attention mechanism is implemented by using a bidirectional long short-term memory artificial neural network to determine whether there is an interaction relationship between agents. For the i-th and j-th agents, the features of the agents i and j are combined to obtain (h i ,h j )((h i ,h1), …, (h i ,h N )), and (h i ,h j ) is input into a Bi-LSTM model to obtain h i,j =f(Bi-LSTM(hi h j )), where f is a fully connected layer. Since the output of the LSTM artificial neural network only depends on the input of the current time and the previous time, the input information of the later time is ignored, so that the information of part of the agents cannot be utilized, which is short-sighted and unreasonable. Therefore, a bidirectional LSTM artificial neural network is used to realize the hard attention mechanism. In addition, the hard attention mechanism involves a sampling process and cannot back-propagate the gradient, so the gumbel-softmax function is used to solve the back-propagation problem, and the real value between 0 and 1 of the edge between the agent i and the agent j is obtained Through the hard attention mechanism, the subgraph G i of the i-th agent can be obtained The soft attention mechanism is used to learn the weight of each edge in the subgraph G i The weight of the edge between the agent i and the agent j in G i is For multiple information, different hard attention values are obtained where e i and e j are the embeddings of the agent i and the agent j, respectively (e i and e j can be replaced by h i and h j ), W k and W q are the key linear mapping and the query linear mapping, respectively. W k converts e j into a key vector, and W q converts e i into a query vector W k and W q are the key linear mapping and the query linear mapping, respectively. W k converts e j into a key vector, and W q converts e i into a query vector.

[0219] In some embodiments, through the two-stage attention model, a reduced graph can be obtained, in which each agent is only connected to the agent (node) that needs to interact. The neighbor information x i (x1, x2, x3, x4) is obtained by weighting the neighbor features using the weight output by the soft attention mechanism, and finally the strategy of each agent a i = π(h i , x i) is the strategy of the ith agent, i.e., the candidate recommendation information set, where h i ,x i represents the observation feature of the agent and the contribution of other agents to the ith agent, respectively.

[0220] In some embodiments, each information of the original problem is regarded as an agent and is taken as a node of the graph attention network. The connection between different information in the message passing inference feedback function of the graph attention network is used to make a comprehensive decision for each node, so that the feature of each node is the feature vector of the information represented by the node (which can be extracted according to the information related information in the recommendation system data set), and the observation vector of each node is the training progress vector of the information represented by the node, which is composed of the selection rate of the information in the past to the current round, the average rate of the information being selected simultaneously with the inadmissible set of information, the average return and the standard deviation of the overall action in the round in which the information is selected, and the loss of the t-th round is -r t +λc t , where r t is the neural network function h' of the master server, and h' is the agent feedback calculated by the graph attention network for the t-th round decision, c t is the constraint violation rate of the t-th round decision (constraint feedback).

[0221] The graph attention network outputs a real number in the interval [0, 1] for each node. To obtain the final decision of each round, three methods can be adopted: for node i, if the output of the graph attention network at node i is greater than 0.5, the information i is selected; the real values output by the graph attention network for all nodes are sorted, and the information corresponding to the top K real values is selected; the sampling probability can also be calculated according to the output a i of node i, see formula (44):

[0222] U~uniform(0, 1), b i =a i -log(-logU) (44);

[0223] b i and b i obey the gumbel distribution corresponding to a i , and the information corresponding to the top K nodes in b i is selected, at this time, b iThe order of the top K values ​​in the graph is distinguished, and the ordered probabilities can be converted into unordered probabilities. That is, each new ranking is permuted K times, and the K different ordered probabilities are averaged to obtain the unordered probabilities. Since calculating K probabilities is computationally intensive, M permutations for each K piece of information can be randomly generated and the corresponding M probabilities are averaged. The probability of the final decision in each round is obtained by sampling the output of the graph attention network.

[0224] The following is a detailed introduction to the improved cross-entropy-deep reinforcement learning evolution sampler (hereinafter referred to as the fourth slave server). The evolutionary strategy is a black-box optimization technology. Its performance on modern reinforcement learning benchmarks is comparable to that of standard reinforcement learning technology. At the same time, it can overcome many inconveniences of reinforcement learning, such as no need for backpropagation, easier to expand in a distributed environment, less susceptible to sparse rewards, and fewer hyperparameters. The cross-entropy method is an evolutionary strategy that can be used to solve continuous and discrete optimization problems in parallel. In essence, the cross-entropy method is a search algorithm based on parameter perturbation. It gives the parameter space v some reasonable perturbations, searches and selects a better set among these perturbations (variants / offspring), and then uses cross-entropy to guide the update of v, so that these perturbation directions are closer to the target optimization direction. The cross-entropy method has strong universality, and its specific implementation steps are as follows: First, initialize. For any A t,i Whether it is 1 indicates whether the information i in round t is selected. Let A t,i Subject to the mean μ i Bernoulli distribution, for any Let the initial μ i =0.5. For Q∈{1,2,...,int(T / N)}(int(T / N) is the lower integer of T / N, N is the number of rounds of an epoch), t∈{QN+1,QN+1,...,min{(Q+1)N,T}}, on each component i∈[L], from the mean μ i A is obtained by sampling from the Bernoulli distribution t,i For t = min{(Q+1)N,T}, according to h'(.)-λc(.) is {A QN+1 ,A QN+1 ,...,A min{(Q+1)N,T}}, where h'(.) is the proxy revenue from the master server, c(.) is the diversity constraint violation rate of the vector, and the vectors with the top ρ percent scores are averaged to obtain the new (μ1,μ2,...,μ L ) vector, so that (μ1,μ2,...,μ L) generates as high a score as possible, we need to have a more accurate understanding of the layout of the candidate recommendation information set, which requires the length N of the training round to be as large as possible. However, a large N value will lead to (μ1,μ2,...,μ L ) updates slowly and converges slowly, making the cross entropy method unable to play an advantage in the online decision environment. Therefore, the number of training batches is divided into multiple time segments of length n, and at the critical point of each small segment, the proximal strategy is used to optimize the pair (μ1,μ2,...,μ L ) performs multi-step gradient descent processing. For j∈{1,2,...,int(N / n)}, when t=min{QN+(j+1)n,T}, the objective function is given by formula (45), and the descent gradient algorithm is adopted to update the parameters. See formula (45):

[0225]

[0226] Among them, u old is the parameter followed by sampling in the n rounds closest to the time t = min{QN+(j+1)n,T}, u new is the parameter followed by the sampling from min{QN+(j+1)n+1,T} rounds to the next n rounds. t,i =1, then P(A t,i |u i )=u i If A t,i =0, then P(A t,i |u i )=1-u i Every n rounds, update J t , and J t About U new Perform multi-step gradient descent to achieve timely parameter updates required for online decision-making. The best historical samples can be appropriately stored to enrich the current set of candidate recommendation information. A linear or exponential decay factor can also be set so that the new parameters are a combination of past parameters and the currently updated parameters.

[0227] The following is a detailed introduction to the teacher-student sampler. The best candidate recommendation information set in history (evaluated by h'(.)-λc(.)) is used as the teacher, and the other candidate recommendation information sets provided by the server are used as students. Interactions are carried out between teachers and students, and between students, to generate new candidate recommendation information sets. Let the teacher candidate recommendation information set be T and the student set be S. The teacher-student interaction is to select any student candidate recommendation information set A from the S set, let B = A + rand*(TA), rand is a random number in the interval [0,1], sort the components in A, set the components in the top K positions to 1, and set the other components to 0, to obtain a new candidate recommendation information set B. Repeat the above operation several times. The student-student interaction is to select any two student candidate recommendation information sets A and B from the S set. If h'(A)-λc(A) <h'(B)-λc(B),令C=A+rand*(B-A);否则C=A+rand*(A-B)。对C中分量进行排序,排在前K位的分量设为1,其他分量设为0,得到新的候选推荐信息集合C。贝塔分布在从服务器的作用如下:当一些从服务器输出0-1整数值向量A时,可对A的各分量进行扰动并将扰动值作为beta分布的参数,执行对beta分布的采样得到新候选推荐信息集合,例如,构造[0,1]区间实值向量B∈[0,1] L , for i∈[L], if A i =1, then let B i =1-τ; otherwise, let B i =τ, with (B i ,1-B i ) is the parameter of the beta distribution, and a real value C is randomly sampled from the beta distribution. i ,C1,C2,...,C L Construct a vector C, sort the components in C, set the components in the first K positions to 1, and set the other components to 0, and obtain a new candidate recommendation information set C.

[0228] In some embodiments, the data processing flow for the master and slave servers is as follows: A dataset on combination decision-making in a recommendation system or other fields is obtained. L items are selected as the L pieces of information s for the original problem, either through rough sorting or based on item popularity. The value of K is determined based on the number of combination recommendations in the dataset. The feature vector for each piece of information is provided by the dataset. If not, the feature vector can be calculated using singular value decomposition or other standard methods on the user-information interaction matrix. Feedback received by the master server after making a decision is directly provided by the dataset. Feedback received by the slave server after making a decision is the master server's revenue feedback h'(.) - λc(.), where the reward feedback is h'(.).

[0229] In some embodiments, the embodiments of the present application provide a framework for solving online combinatorial optimization problems in recommendation systems, integrating multiple samplers (from servers) to ensure that the action data output by the framework for multiple information meets the diversity constraints with a high probability and has good benefit feedback, utilizing the strong constraint processing capabilities of the optimization method corresponding to the sampler, the highly parallelizable capabilities of the evolutionary algorithm of the swarm intelligence agent, the strong fitting capabilities of the neural network and the online decision-making capabilities of the reinforcement learning method, and achieving a clever fusion when solving large-scale online combinatorial optimization problems.

[0230] In some embodiments, the objective function form of the combinatorial optimization problem is very complex and unknown, so it is difficult to output a better solution by relying on the optimization solver, and thus it is necessary to rely on deep learning and reinforcement learning methods. However, when there are many constraints that can be encoded into logical constraints, it is difficult for the neural network output to output a feasible solution or an approximate feasible solution. At this time, it is necessary to use a highly aggregated semantic loss function that builds a bridge between the neural output vector and the logical constraints. The semantic loss function propagates and aggregates the logical constraints on a probabilistic circuit that can be back-propagated, calculates the importance of each variable in all constraints, makes the reasoning process differentiable and retains the precise logical meaning of knowledge, and at the same time, the semantic loss function is highly Aggregation eliminates the need to manually assign different weights to each constraint. Constraint feedback (diversity features) in multiple slave servers and the diversity features in the master server can be replaced by a semantic loss function. The semantic loss function calculation process is as follows: data preprocessing is performed. For constraints that can be converted into pseudo-Boolean constraints (such as 0-1 linear constraints), the constraint programming solver can be used to convert the constraints into pseudo-Boolean constraints, and then the pseudo-Boolean constraints are converted into a conjunction normal form. The conjunction normal form is then converted into a probabilistic sentence decision diagram. The PyPSDD library is used to calculate the semantic loss of the candidate recommendation information set output by the neural network based on the probabilistic sentence decision diagram to replace the original diversity features.

[0231] The following continues to describe the exemplary structure of the artificial intelligence-based information recommendation device 255 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2As shown, the software modules in the artificial intelligence-based information recommendation device 255 stored in the memory 250 may include: an acquisition module 2551, used to acquire multiple candidate recommendation information sets, and determine the expected items and uncertain items of the information characteristics of each candidate recommendation information set; an aggregation module 2552, used to aggregate the expected items and uncertain items of each candidate recommendation information set to obtain the upper confidence bound characteristics of each candidate recommendation information set; a diversity module 2553, used to determine the diversity characteristics corresponding to each candidate recommendation information set; an index module 2554, used to determine the recommendation index of the corresponding candidate recommendation information set based on the upper confidence bound characteristics and constraint violation characteristics of each candidate recommendation information set; a recommendation module 2555, used to use the candidate recommendation information set with the highest recommendation index as the information set to be recommended, so as to perform a recommendation operation on the information set to be recommended.

[0232] In some embodiments, the acquisition module 2551 is also used to: perform at least one of the following processing to obtain multiple candidate recommendation information sets: obtain multiple candidate recommendation information sets based on a linear estimation function; obtain multiple candidate recommendation information sets based on a quadratic estimation function; obtain multiple candidate recommendation information sets through an action evaluation framework; obtain multiple candidate recommendation information sets by combining a soft attention mechanism and a hard attention mechanism; obtain multiple candidate recommendation information sets through a Bernoulli distribution.

[0233] In some embodiments, the acquisition module 2551 is also used to: perform mapping processing on the i-th column vector of the L column vectors of the unit matrix to obtain the mapping processing result corresponding to the i-th column vector; wherein the L column vectors correspond one-to-one to the L information; L is an integer greater than or equal to 2, and the value range of i satisfies 1≤i≤L; using the mapping processing result of the column vector of the corresponding information as the weight, weighted summing processing is performed on the action data of the L information to obtain a linear estimation function; wherein the action data represents whether the corresponding information is selected or not; determine the action data of the L information while satisfying the following conditions: when the action data of the L information are substituted into the linear estimation function, the value of the linear estimation function is the maximizing convergence value; the action data of the L information represent that at least one of the selected information in the L information satisfies the diversity constraint; and the at least one selected information in the L information is composed of a candidate recommendation information set.

[0234] In some embodiments, the acquisition module 2551 is further used to: perform mapping processing on the i-th column vector among the L column vectors of the unit matrix to obtain the mapping processing result corresponding to the i-th column vector, and use the mapping processing result corresponding to the i-th column vector as a matrix element; sum the i-th column vector and the j-th column vector among the L column vectors of the unit matrix, perform mapping processing on the summation processing result, and obtain the mapping processing result corresponding to the i-th column vector and the j-th column vector; wherein L is an integer greater than or equal to 2, the value range of i and j satisfies 1≤i, j≤L, and the values ​​of i and j are different; average the mapping processing result corresponding to the i-th column vector and the mapping processing result corresponding to the j-th column vector, and average the mapping processing result corresponding to the i-th column vector and the j-th column vector. The mapping processing result of the quantity is subtracted from the average processing result to obtain matrix elements; a matrix is ​​constructed according to the matrix elements; the transpose of the action data matrix corresponding to L information and the matrix are multiplied with the action data matrix to obtain a quadratic estimation function; wherein the action data matrix includes action data corresponding to the L information one by one, and the action data represents whether the corresponding information is selected or not; the action data of the L information is determined while satisfying the following conditions: when the action data of the L information are substituted into the quadratic estimation function, the value of the quadratic estimation function is the maximization convergence value; the action data of the L information represent that at least one of the selected information in the L information satisfies the diversity constraint; at least one of the selected information in the L information is composed of a candidate recommendation information set.

[0235] In some embodiments, the acquisition module 2551 is also used to: generate an action matrix with L column vectors through the action network in the action evaluation framework, and determine a candidate recommendation information set corresponding to the action matrix; wherein the column identifiers of the L column vectors correspond one-to-one to the L information, L is an integer greater than or equal to 2, and the value of the column vector represents the action data of the corresponding information; perform the following processing on the action matrix any number of times: swap any two different column vectors in the L column vectors in the action matrix to obtain a new action matrix, and determine a candidate recommendation information set corresponding to the new action matrix.

[0236] In some embodiments, the acquisition module 2551 is further used to: generate action data corresponding to each information through the action network in the action evaluation framework; sort the L information in descending order according to the action data of each information; update the action data of multiple information ranked at the top of the L information to one, and update the action data of other information to zero; wherein the other information is the information other than the multiple information ranked at the top of the L information; convert the updated action data of each information into a column vector of the corresponding information to obtain an action matrix with L column vectors.

[0237] In some embodiments, the acquisition module 2551 is further used to: initialize the evaluation network and the action network of the action evaluation framework; perform K iterative processing on the action evaluation framework, and perform the following processing during each iterative processing: perform T rounds of update processing on the action network and the evaluation network of the action evaluation framework according to the trade-off coefficients of the expected items and the diversity characteristics, and update the trade-off coefficients according to the results of the T-th round of update processing; wherein T and K are both integers greater than or equal to 2; and determine the action network obtained by the K-th iterative processing as the action network used to generate an action matrix with L column vectors.

[0238] In some embodiments, the acquisition module 2551 is also used to: perform T rounds of iterative processing on the action evaluation framework, and perform the following processing during each round of iterative processing: predict the candidate recommendation information set samples through the action network, and obtain the expected items and diversity characteristics of the corresponding candidate recommendation information set samples; determine the value function value of the corresponding candidate recommendation information set samples through the evaluation network, and determine the comprehensive value of the corresponding candidate recommendation information set samples based on the expected items, diversity characteristics, trade-off coefficients and value function values; obtain the error between the comprehensive value and the value function value, and update the parameters of the evaluation network according to the gradient term of the corresponding error; determine the penalty value function value of the corresponding candidate recommendation information set samples based on the expected items, diversity characteristics and trade-off coefficients, and update the parameters of the action network according to the gradient term of the corresponding penalty value function.

[0239] In some embodiments, the acquisition module 2551 is further used to: obtain local observation data corresponding to each information in the L information, and encode the local observation data into observation features; determine at least one interactive information in the L information that has an interactive relationship with the i-th information based on the hard attention mechanism and in combination with the observation features of each information; determine the interaction weight between each interactive information and the i-th information based on the soft attention mechanism, and determine the interaction features of all interactive information corresponding to the i-th information based on the interaction weight; determine the policy prediction value corresponding to the i-th information through the policy network based on the observation features and interaction features of the i-th information; wherein L is an integer greater than or equal to 2, i is an integer whose value increases from 1, and the value range of i satisfies 1≤i≤L; obtain a set of candidate recommended information based on the policy prediction value of each information in the L information.

[0240] In some embodiments, the obtaining module 2551 is further configured to: merge the observed feature of the i-th information with the observed feature of each other information different from the i-th information to obtain a merged feature corresponding to each other information; perform mapping processing on each merged feature by using a bidirectional long short-term memory artificial neural network, and perform maximum likelihood processing on the mapping processing result to obtain a hard attention value corresponding to each other information; and determine, as an interaction information having an interaction relationship with the i-th information, the other information whose hard attention value is greater than a hard attention threshold value, from the L information.

[0241] In some embodiments, the obtaining module 2551 is further configured to: for each interaction information, perform the following processing: obtain an i-th embedding feature of the i-th information, and perform linear mapping on the i-th embedding feature according to a query parameter of the soft attention mechanism to obtain a query feature corresponding to the i-th information; obtain an interaction embedding feature of the interaction information, and perform linear mapping on the interaction embedding feature according to a key parameter of the soft attention mechanism to obtain a key feature corresponding to the interaction information; determine a soft attention value that is exponentially positively correlated with the key feature, the query feature, and the hard attention value as an interaction weight corresponding to the interaction information; and perform weighting processing on the observed feature of each interaction information according to the interaction weight corresponding to the interaction information to obtain an interaction feature of all interaction information for the i-th information.

[0242] In some embodiments, the obtaining module 2551 is further configured to: perform any one of the following processing: obtain, from the L information, a plurality of information corresponding to a strategy prediction value greater than a strategy prediction threshold value, and sample K sampling information from the plurality of information to form a candidate recommended information set; perform descending order sorting processing on the L information according to the strategy prediction value of each information, and obtain K information ranked at the front to form the candidate recommended information set; and wherein K is the number of recommended information in the candidate recommended information set.

[0243] In some embodiments, the obtaining module 2551 is further configured to: obtain a training sample set, wherein the training sample set comprises N candidate recommendation information set samples corresponding to N rounds of historical recommendations, N being an integer greater than or equal to 2; divide the N rounds of historical recommendations to obtain a plurality of historical recommendation periods, wherein each historical recommendation period comprises M rounds of historical recommendations, M being an integer greater than 1 and less than N; initialize an objective function, wherein the objective function is used to represent the maximization of a penalty value function in the M rounds of historical recommendations, and the objective function comprises a Bernoulli distribution corresponding to a qth historical recommendation period and a Bernoulli distribution corresponding to a (q-1)th historical recommendation period, q being an integer greater than or equal to 2; in each historical recommendation period, perform the following processing: obtain the Bernoulli distribution corresponding to the historical recommendation period, and generate a candidate recommendation information set sample corresponding to each round of historical recommendation according to the Bernoulli distribution; determine a penalty value function value corresponding to each candidate recommendation information set sample, and substitute it into the objective function to perform gradient descent processing of the objective function for the Bernoulli distribution corresponding to the qth historical recommendation period, to obtain a Bernoulli distribution corresponding to a (q+1)th historical recommendation period; and generate a candidate recommendation information set based on the Bernoulli distribution of the last historical recommendation period.

[0244] In some embodiments, the obtaining module 2551 is further configured to: generate a new candidate recommendation information set according to a teacher-student mechanism and in combination with the obtained plurality of candidate recommendation information sets; or generate a new candidate recommendation information set according to a beta distribution sampling mechanism and in combination with the obtained plurality of candidate recommendation information sets.

[0245] In some embodiments, the obtaining module 2551 is further configured to: obtain an expected item and a diversity feature of each historical candidate recommendation information set, to determine a penalty value function value corresponding to each historical candidate recommendation information set, and determine a historical candidate recommendation information set with the highest corresponding penalty value function value as a teacher set, and determine each candidate recommendation information set as a student set; for any one student set, perform at least one of the following processing: map any one student set and the teacher set according to an operator to obtain a new candidate recommendation information set, or map any one student set and another student set different from any one student set according to an operator to obtain a new candidate recommendation information set.

[0246] In some embodiments, the obtaining module 2551 is further configured to: for each candidate recommendation information set, perform the following processing: performing perturbation processing on the action data of each recommendation information in the candidate recommendation information set to obtain a perturbation value of each action data of the candidate recommendation information set; performing perturbation processing on the action data of other information to obtain a perturbation value of each other information, wherein the other information is information other than the recommendation information in the L information, and L is an integer greater than or equal to 2; obtaining a beta distribution corresponding to each recommendation information based on the perturbation value of the corresponding recommendation information, and obtaining a beta distribution corresponding to each other information based on the perturbation value of the corresponding other information; sampling the beta distribution corresponding to each recommendation information to obtain sampled action data of the corresponding each recommendation information, and sampling the beta distribution corresponding to each other information to obtain sampled action data of the corresponding each other information; performing mixed descending sorting on the other information and the recommendation information based on the sampled action data of the corresponding each recommendation information and the sampled action data of the corresponding each other information, and obtaining the top K information in the sorting to constitute a new candidate recommendation information set; wherein K is the number of recommendation information in the candidate recommendation information set.

[0247] In some embodiments, the obtaining module 2551 is further configured to: performing forward propagation of the information features of each candidate recommendation information set in the confidence neural network to obtain an expected item corresponding to each candidate recommendation information set; obtaining a gradient function of the confidence neural network, and substituting the information features of each candidate recommendation information set into the gradient function to obtain an uncertainty item corresponding to each candidate recommendation information set.

[0248] In some embodiments, the diversity module 2553 is further configured to: performing multiple recommendation information extraction processing on each candidate recommendation information set to obtain multiple recommendation information subsets; wherein two recommendation information are extracted in each recommendation information extraction process, and each recommendation information subset includes the two recommendation information extracted in the corresponding recommendation information extraction process; obtaining the total number of recommendation information subsets and the number of recommendation information subsets that do not satisfy the diversity constraint, determining the ratio between the number of recommendation information subsets that do not satisfy the diversity constraint and the total number, and determining a diversity feature corresponding to the ratio.

[0249] The embodiments of the present application provide a computer program product or a computer program, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the information recommendation method based on artificial intelligence provided in the embodiments of the present application.

[0250] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the method provided by the embodiment of the present application, for example, the artificial intelligence-based information recommendation method shown in Figure 3.

[0251] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.

[0252] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0253] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0254] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0255] To summarize, through the embodiments of the present application, based on the information features of the candidate recommendation information set, the expected items and uncertain items for recommendation benefit prediction are characterized for the candidate recommendation information set, taking into account the contribution of information features to user behavior prediction, and ensuring a wide information coverage of the candidate recommendation information set through diversity features, so as to deeply mine the information of interest to the user, ensure the accuracy of information recommendation for subsequent information recommendation, and effectively avoid invalid recommendations, thereby saving computing resources related to recommendation logic in the server.

[0256] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An information recommendation method based on artificial intelligence, characterized in that: include: Acquire multiple candidate recommendation information sets, and determine an expected term and an uncertain term of information characteristics of each candidate recommendation information set, wherein the expected term represents a historical average return value, and the uncertain term represents a supremum value of uncertainty of the historical average return value; Aggregating the expected items and the uncertain items of each candidate recommendation information set to obtain an upper confidence bound feature of each candidate recommendation information set, wherein the upper confidence bound feature is obtained by weighted calculation based on the expected items and the uncertain items, and is used to predict the positive benefits after executing the recommendation operation; Determining a diversity feature corresponding to each of the candidate recommendation information sets; Determining a recommendation index corresponding to each candidate recommendation information set according to the upper confidence bound feature and the diversity feature of each candidate recommendation information set; The candidate recommendation information set with the highest recommendation index is used as the information set to be recommended, so as to perform a recommendation operation on the information set to be recommended.

2. The method according to claim 1, characterized in that The obtaining of multiple candidate recommendation information sets includes: Perform at least one of the following processes to obtain multiple candidate recommendation information sets: Acquire a plurality of candidate recommendation information sets according to a linear estimation function; Acquire a plurality of candidate recommendation information sets according to a quadratic estimation function; Acquire multiple candidate recommendation information sets through an action evaluation framework; Combining the soft attention mechanism with the hard attention mechanism to obtain multiple candidate recommendation information sets; A plurality of candidate recommendation information sets are obtained through Bernoulli distribution.

3. The method according to claim 2, characterized in that The obtaining of the plurality of candidate recommendation information sets according to the linear estimation function includes: Performing mapping processing on an i-th column vector among the L column vectors of the identity matrix to obtain a mapping processing result corresponding to the i-th column vector; Wherein, the L column vectors correspond one-to-one to the L information; L is an integer greater than or equal to 2, and the value range of i satisfies 1≤i≤L; Using the mapping result of the column vector of the corresponding information as the weight, the action data of L information is weighted and summed to obtain the linear estimation function; Wherein, the action data represents whether the corresponding information is selected or not; Determine the L pieces of information while satisfying the following action data: When the action data of the L information are substituted into the linear estimation function, the value of the linear estimation function is the maximum convergence value; The action data of the L information represent that at least one selected information from the L information satisfies the diversity constraint; At least one selected piece of information from the L pieces of information is used to form the candidate recommendation information set.

4. The method according to claim 2, characterized in that The obtaining of the plurality of candidate recommendation information sets according to the quadratic estimation function includes: Performing mapping processing on an i-th column vector among the L column vectors of the identity matrix to obtain a mapping processing result corresponding to the i-th column vector, and using the mapping processing result corresponding to the i-th column vector as a matrix element; performing a summation process on the i-th column vector and the j-th column vector of the L column vectors of the identity matrix, and performing a mapping process on the summation process result to obtain a mapping process result corresponding to the i-th column vector and the j-th column vector; Wherein, L is an integer greater than or equal to 2, the value range of i and j satisfies 1≤i, j≤L, and the values ​​of i and j are different; Averaging the mapping processing result corresponding to the i-th column vector and the mapping processing result corresponding to the j-th column vector, and subtracting the mapping processing results corresponding to the i-th column vector and the j-th column vector from the average processing result to obtain a matrix element; constructing a matrix based on the matrix elements; The transpose of the action data matrix corresponding to L information and the multiplication of the matrix with the action data matrix are performed to obtain a quadratic estimation function; The action data matrix includes action data corresponding to L pieces of information, and the action data represents whether the corresponding information is selected or not; Determine the L pieces of information while satisfying the following action data: When the action data of the L information are substituted into the quadratic estimation function, the value of the quadratic estimation function is a value that maximizes convergence; The action data of the L information represent that at least one selected information from the L information satisfies the diversity constraint; At least one selected piece of information from the L pieces of information is used to form the candidate recommendation information set.

5. The method according to claim 2, characterized in that The acquiring of the plurality of candidate recommendation information sets through the action evaluation framework includes: Generate an action matrix having L column vectors through the action network in the action evaluation framework, and determine a candidate recommendation information set corresponding to the action matrix; The column identifiers of the L column vectors correspond one-to-one to the L information, L is an integer greater than or equal to 2, and the values ​​of the column vectors represent the action data corresponding to the information; Perform the following process any number of times for the action matrix: Any two different column vectors in the L column vectors in the action matrix are swapped to obtain a new action matrix, and a candidate recommendation information set corresponding to the new action matrix is ​​determined.

6. The method according to claim 5, characterized in that The action matrix having L column vectors is generated by the action network in the action evaluation framework, including: Generate action data corresponding to each of the information through the action network in the action evaluation framework; sorting the L pieces of information in descending order according to the action data of each piece of information; Update the action data of the plurality of messages ranked first among the L messages to one, and update the action data of the other messages to zero; The other information is information other than the plurality of information ranked first among the L pieces of information; The updated action data of each of the information is converted into a column vector corresponding to the information to obtain an action matrix having the L column vectors.

7. The method according to claim 5, characterized in that Before generating the action matrix having L column vectors through the action network in the action evaluation framework, the method further includes: Initializing the evaluation network of the action evaluation framework and the action network; The action evaluation framework is iterated K times, and the following processing is performed during each iterative process: Performing T rounds of updating processing on the action network and the evaluation network of the action evaluation framework according to the trade-off coefficient between the expected item and the diversity feature, and updating the trade-off coefficient according to the result of the T rounds of updating processing; Wherein, T and K are both integers greater than or equal to 2; The action network obtained by the K-th iterative process is determined as the action network for generating an action matrix having L column vectors.

8. The method according to claim 7, characterized in that The step of performing T rounds of updating processing on the action network and the evaluation network of the action evaluation framework according to the trade-off coefficient between the expected item and the diversity feature includes: The action evaluation framework is iterated for T rounds, and the following processing is performed during each round of iterative processing: Predicting candidate recommendation information set samples through the action network, and obtaining expected items and diversity features corresponding to the candidate recommendation information set samples; Determining a value function value corresponding to the candidate recommendation information set sample through the evaluation network, and determining a comprehensive value corresponding to the candidate recommendation information set sample based on the desired item, the diversity feature, the trade-off coefficient, and the value function value; Obtaining an error between the comprehensive value and the value function value, and updating the parameters of the evaluation network according to a gradient term corresponding to the error; According to the expectation item, the diversity feature and the trade-off coefficient, a penalty value function value corresponding to the candidate recommendation information set sample is determined, and the parameters of the action network are updated according to the gradient item corresponding to the penalty value function.

9. The method according to claim 2, characterized in that The combining of the soft attention mechanism and the hard attention mechanism to obtain the plurality of candidate recommendation information sets includes: Obtaining local observation data corresponding to each of the L pieces of information, and encoding the local observation data into observation features; Determine, based on a hard attention mechanism and in combination with observation features of each of the information, at least one interactive information among the L information that has an interactive relationship with the i-th information; Determine an interaction weight between each of the interaction information and the i-th information according to a soft attention mechanism, and determine an interaction feature of all the interaction information corresponding to the i-th information according to the interaction weight; Determining a policy prediction value corresponding to the i-th information through a policy network according to the observed features and interactive features of the i-th information; Wherein, L is an integer greater than or equal to 2, i is an integer whose value increases from 1, and the value range of i satisfies 1≤i≤L; The candidate recommendation information set is obtained according to the policy prediction value of each information in the L information.

10. The method according to claim 9, characterized in that The step of determining, based on the hard attention mechanism and in combination with the observed features of each of the information, at least one interactive information among the L information that has an interactive relationship with the i-th information includes: Merging the observed feature of the i-th information with the observed feature of each other information different from the i-th information to obtain a merged feature corresponding to each other information; Mapping each of the merged features using a bidirectional long short-term memory artificial neural network, and performing maximum likelihood processing on the mapping results to obtain a hard attention value corresponding to each of the other information; Other information whose hard attention value is greater than the hard attention threshold is determined as interactive information among the L information that has an interactive relationship with the i-th information.

11. The method according to claim 9, characterized in that The step of determining, according to the soft attention mechanism, an interaction weight between each of the interaction information and the i-th information, and determining, according to the interaction weight, an interaction feature of all the interaction information with respect to the i-th information, includes: The following processing is performed for each interaction information: Obtaining an i-th embedded feature of the i-th information, and performing linear mapping on the i-th embedded feature according to the query parameter of the soft attention mechanism to obtain a query feature corresponding to the i-th information; Obtaining interactive embedding features of the interactive information, and performing linear mapping on the interactive embedding features according to key parameters of the soft attention mechanism to obtain key features corresponding to the interactive information; Determining a soft attention value that is exponentially positively correlated with the key feature, the query feature, and the hard attention value as an interaction weight corresponding to the interaction information; According to the interaction weight corresponding to the interaction information, weighted processing is performed on the observation feature of each interaction information to obtain the interaction feature of all the interaction information for the i-th information.

12. The method according to claim 9, characterized in that The acquiring the candidate recommendation information set according to the policy prediction value of each of the L pieces of information includes: Perform any of the following: Acquire multiple pieces of information whose corresponding policy prediction values ​​are greater than a policy prediction threshold from the L pieces of information, and sample K pieces of sampled information from the multiple pieces of information to form the candidate recommendation information set; Sorting the L pieces of information in descending order according to each of the information strategy prediction values, and obtaining the top K pieces of information to form the candidate recommendation information set; Wherein, K is the number of recommended information in the candidate recommended information set.

13. The method according to claim 2, characterized in that The obtaining of the plurality of candidate recommendation information sets through Bernoulli distribution includes: Obtaining a training sample set, wherein the training sample set includes N candidate recommendation information set samples corresponding one-to-one to N rounds of historical recommendations, where N is an integer greater than or equal to 2; Dividing the N rounds of historical recommendations to obtain multiple historical recommendation cycles, wherein each historical recommendation cycle includes M rounds of historical recommendations, where M is an integer greater than 1 and less than N; Initialize an objective function, wherein the objective function is used to characterize the maximization of the penalty value function value in the M rounds of historical recommendations, and the objective function includes a Bernoulli distribution corresponding to the qth historical recommendation cycle and a Bernoulli distribution corresponding to the q-1th historical recommendation cycle, where q is an integer greater than or equal to 2; In each of the historical recommendation cycles, the following processing is performed: Obtaining a Bernoulli distribution corresponding to the historical recommendation cycle, and generating a sample of candidate recommendation information sets corresponding to each round of the historical recommendation according to the Bernoulli distribution; Determine a penalty value function value corresponding to each sample of the candidate recommendation information set, and substitute the value into the objective function to perform a gradient descent process on the objective function with respect to the Bernoulli distribution corresponding to the qth historical recommendation period, thereby obtaining a Bernoulli distribution corresponding to the q+1th historical recommendation period; Generate a set of candidate recommendation information based on the Bernoulli distribution of the last historical recommendation cycle.

14. The method according to claim 2, characterized in that The method further comprises: Generate a new candidate recommendation information set based on the teacher-student mechanism and in combination with the obtained multiple candidate recommendation information sets; or According to the Beta distribution sampling mechanism, and in combination with the obtained multiple candidate recommendation information sets, a new candidate recommendation information set is generated.

15. The method according to claim 14, characterized in that The method of generating a new candidate recommendation information set based on the teacher-student mechanism and combining the obtained multiple candidate recommendation information sets includes: Obtaining the expected items and diversity characteristics of each historical candidate recommendation information set to determine a penalty value function value corresponding to each of the historical candidate recommendation information sets, and determining the historical candidate recommendation information set with the highest corresponding penalty value function value as the teacher set, and determining each candidate recommendation information set as the student set; For any set of students, perform at least one of the following operations: Map any one of the student sets and the teacher set according to the operator to obtain a new candidate recommendation information set, or The arbitrary student set and another student set different from the arbitrary student set are mapped according to the operator to obtain a new candidate recommendation information set.

16. The method according to claim 14, characterized in that The generating of a new candidate recommendation information set based on the Beta distribution sampling mechanism and combining the obtained multiple candidate recommendation information sets includes: The following processing is performed for each candidate recommendation information set: performing perturbation processing on the action data of each piece of recommendation information in the candidate recommendation information set to obtain a perturbation value of each piece of action data in the candidate recommendation information set; Performing perturbation processing on the action data of other information to obtain a perturbation value for each of the other information, wherein the other information is information other than the recommended information in the L information, and L is an integer greater than or equal to 2; Based on the disturbance value corresponding to each piece of the recommended information, obtaining a beta distribution corresponding to the recommended information, and based on the disturbance value corresponding to each piece of the other information, obtaining a beta distribution corresponding to the other information; Sampling from a beta distribution corresponding to the recommended information to obtain sampled action data corresponding to each piece of recommended information, and sampling from a beta distribution corresponding to the other information to obtain sampled action data corresponding to each piece of other information; Based on the sampled action data corresponding to each piece of recommended information and the sampled action data corresponding to each piece of other information, the other information and the recommended information are sorted in mixed descending order, and K pieces of information with the highest sorting order are obtained to form a new candidate recommended information set; Wherein, K is the number of recommended information in the candidate recommended information set.

17. The method according to claim 1, wherein The determining of expected items and uncertain items of information features of each candidate recommendation information set includes: Forward propagating the information features of each candidate recommendation information set in a belief neural network to obtain the expected items corresponding to each candidate recommendation information set; A gradient function of the belief neural network is obtained, and the information feature of each candidate recommendation information set is substituted into the gradient function to obtain an uncertainty item corresponding to each candidate recommendation information set.

18. The method according to claim 1, wherein The determining of the diversity feature corresponding to each candidate recommendation information set includes: Performing multiple recommendation information extraction processes on each candidate recommendation information set to obtain multiple recommendation information subsets; Wherein, two pieces of recommendation information are extracted in each recommendation information extraction process, and each of the recommendation information subsets includes the two pieces of recommendation information extracted in the corresponding recommendation information extraction process; The total number of the recommended information subsets and the number of the recommended information subsets that do not meet the diversity constraint are obtained, a ratio between the number of the recommended information subsets that do not meet the diversity constraint and the total number is determined, and a diversity feature corresponding to the ratio is determined.

19. An information recommendation device based on artificial intelligence, characterized in that: include: an acquisition module, configured to acquire a plurality of candidate recommendation information sets and determine an expected term and an uncertain term of information characteristics of each candidate recommendation information set, wherein the expected term represents a historical average return value, and the uncertain term represents a supremum value of uncertainty of the historical average return value; an aggregation module, configured to aggregate the expected items and uncertain items of each candidate recommendation information set to obtain an upper confidence bound feature of each candidate recommendation information set, wherein the upper confidence bound feature is obtained by weighted calculation based on the expected items and the uncertain items, and is used to predict the positive benefits after executing the recommendation operation; A diversity module, configured to determine a diversity feature corresponding to each of the candidate recommendation information sets; An index module, configured to determine a recommendation index corresponding to each candidate recommendation information set based on the upper confidence bound feature and the diversity feature of each candidate recommendation information set; The recommendation module is configured to take the candidate recommendation information set with the highest recommendation index as the information set to be recommended, so as to perform a recommendation operation on the information set to be recommended.

20. The device according to claim 19, characterized in that The acquisition module is further used to: initialize the evaluation network and the action network of the action evaluation framework; perform K iterative processing on the action evaluation framework, and perform the following processing during each iterative processing: perform T rounds of update processing on the action network and the evaluation network of the action evaluation framework according to the trade-off coefficient between the expected item and the diversity feature, and update the trade-off coefficient according to the result of the T-th round of update processing; wherein T and K are both integers greater than or equal to 2; and determine the action network obtained by the K-th iterative processing as the action network used to generate an action matrix with L column vectors.

21. The device according to claim 19, characterized in that The acquisition module is also used to: generate a new candidate recommendation information set based on the teacher-student mechanism and in combination with the multiple candidate recommendation information sets obtained; or generate a new candidate recommendation information set based on the Beta distribution sampling mechanism and in combination with the multiple candidate recommendation information sets obtained.

22. An electronic device, characterized in that: include: a memory for storing executable instructions; The processor is configured to implement the artificial intelligence-based information recommendation method according to any one of claims 1 to 18 when executing the executable instructions stored in the memory.

23. A computer-readable storage medium, characterized in that Executable instructions are stored for implementing the artificial intelligence-based information recommendation method according to any one of claims 1 to 18 when executed by a processor.

24. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the artificial intelligence-based information recommendation method according to any one of claims 1 to 18 is implemented.

Citation Information

Patent Citations

  • Recommendation method and device based on artificial intelligence, electronic equipment and storage medium

    CN111291266A

  • Information recommendation method and device based on artificial intelligence and electronic equipment

    CN111695037A