Information Processing Method, Apparatus, Computer-Readable Medium, and Electronic Device

Through the information processing method of obtaining and sorting candidate information in the information display system, the problem of underutilizing information display opportunities is solved, and more efficient network resource utilization is achieved.

CN113570395BActive Publication Date: 2025-06-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110088912.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-22
Publication Date
2025-06-10
Estimated Expiration
2041-01-22

AI Technical Summary

Technical Problem

In the prior art, when the information display system processes information delivered in different ways, the information display opportunities are not fully utilized, which leads to poor network resource utilization.

Method used

Through an information processing method, a collection of candidate information composed of multiple candidate information is obtained, including competition display information and agreement display information. The information sorting score of the competition display information is determined using the resource contribution, and the agreed display information is predicted through the policy network model to obtain the information sorting scores of each information, and finally the target information to be displayed is selected from the candidate information set.

Benefits of technology

It realizes the mixed sorting of agreed display information and competitive display information on the same standards, makes full use of the information display opportunities in the system, and improves the network resource utilization rate of information display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113570395B_ABST
    Figure CN113570395B_ABST
Patent Text Reader

Abstract

This application belongs to the field of artificial intelligence technology, and specifically relates to an information processing method, an information processing device, a computer-readable medium, and an electronic device. The method includes: obtaining a candidate information set composed of multiple candidate information according to an information display request, where the candidate information includes competitive display information that competes for a display opportunity according to the amount of resource expenditure and agreed display information with agreed display quantity requirements; determining an information ranking score for each piece of competitive display information according to the amount of resource expenditure, and the information ranking score is used to represent the display priority of the candidate information; performing a score prediction process on the agreed display information through a policy network model to obtain an information ranking score for each piece of agreed display information; the policy network model is a reinforcement learning model trained based on multiple parallel model training processes; selecting target information to be displayed from the candidate information set according to the information ranking score. This method can improve information processing efficiency and network resource utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to an information processing method, an information processing device, a computer-readable medium, and an electronic device. Background Art

[0002] In an information display scenario (such as an advertisement display scenario), an information provider can place information on an information display system in two ways, namely, by agreeing on the number of displays and by bidding.

[0003] In the related art, for two types of information placed in different ways, the information display system controls the display of these two types of information separately. For example, the information display system first predicts the information display opportunities in the system and allocates the predicted information display opportunities to two types of information placed in different ways; when an information display opportunity comes, the information display system selects one piece of information from the information corresponding to the placement method for display.

[0004] However, the above scheme of separately controlling the display of these two types of information will result in the underutilization of information display opportunities in the system, and thus the network resource utilization rate of information display is poor. Summary of the Invention

[0005] The purpose of this application is to provide an information processing method, an information processing device, a computer-readable medium, and an electronic device, which can at least overcome to some extent the problem of poor network resource utilization in the related art.

[0006] Other features and advantages of this application will become apparent through the following detailed description, or be partially learned through the practice of this application.

[0007] According to one aspect of the embodiments of this application, an information processing method is provided. The method includes: obtaining a candidate information set composed of multiple candidate information according to an information display request, where the candidate information includes competitive display information that competes for display opportunities according to the amount of resource paid and agreed display information with agreed display quantity requirements; determining an information ranking score for each piece of the competitive display information according to the amount of resource paid, where the information ranking score is used to represent the display priority of the candidate information; performing a score prediction process on the agreed display information through a policy network model to obtain an information ranking score for each piece of the agreed display information; the policy network model is a reinforcement learning model trained based on multiple parallel model training processes; and selecting target information to be displayed from the candidate information set according to the information ranking score.

[0008] According to one aspect of the embodiments of the present application, an information processing device is provided. The device includes: a candidate information acquisition module configured to acquire a candidate information set composed of a plurality of candidate information according to an information display request, where the candidate information includes competitive display information that competes for a display opportunity according to the amount of resource expenditure and agreed display information with agreed display quantity requirements; a first score acquisition module configured to determine an information ranking score for each of the competitive display information according to the amount of resource expenditure, and the information ranking score is used to represent the display priority of the candidate information; a second score acquisition module configured to perform score prediction processing on the agreed display information through a policy network model to obtain an information ranking score for each of the agreed display information; the policy network model is a reinforcement learning model trained based on a plurality of parallel model training processes; a target information selection module configured to select target information to be displayed from the candidate information set according to the information ranking score.

[0009] In some embodiments of the present application, based on the above technical solution, the information processing device further includes: a sample exploration module configured to acquire a plurality of sample sets respectively maintained by a plurality of parallel sample exploration processes, where the sample set includes training samples obtained by the sample exploration process performing policy exploration on a sample environment related to a historical information display request; a model training module configured to read training samples from the sample set based on a plurality of parallel model training processes, and perform score prediction processing on the training samples through the policy network model to obtain a loss error corresponding to the training samples; a parameter update module configured to update the network parameters of the policy network model according to the loss error.

[0010] In some embodiments of the present application, based on the above technical solution, the sample exploration module includes: a set acquisition unit configured to acquire a sample information set corresponding to a historical information display request, and form a sample environment with the sample information in the sample information set; a policy exploration unit configured to perform policy exploration on the sample environment respectively through a plurality of parallel sample exploration processes to obtain training samples corresponding to the historical information display request, where the training samples include environmental state data, an information display policy corresponding to the environmental state data, and an information display benefit determined according to the environmental state data and the information display policy; a sample storage unit configured to store the training samples explored by the sample exploration process into the sample set maintained by the sample exploration process.

[0011] In some embodiments of the present application, based on the above technical solutions, the sample storage unit includes: a sample quantity monitoring subunit, configured to monitor the quantity of training samples obtained by the sample exploration process to determine whether the quantity of the training samples reaches a preset quantity threshold; a sample writing subunit, configured to write the training samples into a sample set shared memory corresponding to the sample exploration process when it is monitored that the quantity of the training samples reaches the preset quantity threshold, and the sample set shared memory corresponds one-to-one with the sample set maintained by the sample exploration process.

[0012] In some embodiments of the present application, based on the above technical solutions, the sample writing subunit includes: a data storage quantity acquisition subunit, configured to acquire the data storage quantity of the sample set shared memory corresponding to the sample exploration process; a sample sequential writing subunit, configured to sequentially write the training samples into the blank area of the sample set shared memory when the data storage quantity does not reach the maximum capacity of the sample set shared memory; a sample random overwrite subunit, configured to randomly write the training samples into any storage location of the sample set shared memory when the data storage quantity reaches the maximum capacity of the sample set shared memory, so that the training samples randomly overwrite the existing data in the sample set shared memory.

[0013] In some embodiments of the present application, based on the above technical solutions, the sample writing subunit includes: a data storage quantity acquisition subunit, configured to acquire the data storage quantity of the sample set shared memory corresponding to the sample exploration process; a sample sequential writing subunit, configured to sequentially write the training samples into the blank area of the sample set shared memory when the data storage quantity does not reach the maximum capacity of the sample set shared memory; a sample sequential overwrite subunit, configured to sequentially write the training samples into the sample set shared memory when the data storage quantity reaches the maximum capacity of the sample set shared memory, so that the training samples sequentially overwrite the existing data in the sample set shared memory according to the data writing time.

[0014] In some embodiments of the present application, based on the above technical solutions, the information processing device further includes: a status monitoring module, configured to monitor in real time the data storage quantity and the data writing status of the sample set shared memory; an identification bit assignment module, configured to assign an identification bit for the status of the sample set shared memory according to the monitored data storage quantity and data writing status, and the identification bit for the status is used to indicate whether the data in the sample set shared memory is readable.

[0015] In some embodiments of the present application, based on the above technical solutions, the identification bit assignment module includes: a first assignment unit configured to assign a first value to the status identification bit of the sample set shared memory when it is detected that data is being written into the sample set shared memory; the status identification bit with the first value is used to indicate that the sample set shared memory is in a state where data cannot be read; a second assignment unit configured to assign a first value to the status identification bit of the sample set shared memory when it is detected that the data writing is completed and the data storage amount does not reach the maximum capacity of the sample set shared memory; a third assignment unit configured to assign a second value to the status identification bit of the sample set shared memory when it is detected that the data writing is completed and the data storage amount reaches the maximum capacity of the sample set shared memory; the status identification bit with the second value is used to indicate that the sample set shared memory is in a state where data can be read.

[0016] In some embodiments of the present application, based on the above technical solutions, the model training module includes: a status polling unit configured to poll the sample sets maintained by each sample exploration process based on multiple parallel model training processes to determine whether the sample sets are in a state where data can be read; a data reading unit configured to read data from the sample sets when the sample sets are in a state where data can be read.

[0017] In some embodiments of the present application, based on the above technical solutions, the parameter update module includes: an error gradient calculation unit configured to calculate the error gradients of the policy network models maintained by each of the model training processes respectively according to the loss errors obtained by training with multiple parallel model training processes; a network parameter update unit configured to write the error gradients calculated by each of the model training processes into the model parameter shared memory to update the network parameters of the policy network models stored in the model parameter shared memory according to the error gradients.

[0018] In some embodiments of the present application, based on the above technical solutions, the policy network model includes a current policy network model and a target policy network model that is the training target of the current policy network model; the network parameter update unit includes: a first parameter update sub-unit configured to update the network parameters of the current policy network model stored in the model parameter shared memory according to the error gradients; a second parameter update sub-unit configured to update the network parameters of the target policy network according to the network parameters of the current policy network model when a preset target update condition is satisfied.

[0019] In some embodiments of the present application, based on the above technical solutions, the policy network model includes a current policy network model and a target policy network model that serves as the training objective of the current policy network model. The current policy network model includes a current policy generation network for generating an information display policy and a current policy evaluation network for evaluating the information display policy. The target policy network model includes a target policy generation network that serves as the training objective of the current policy generation network and a target policy evaluation network that serves as the training objective of the current policy evaluation network. The model training module includes: a first benefit prediction unit configured to perform score prediction processing on the training sample through the current policy network model to obtain a current policy benefit corresponding to the training sample; a second benefit prediction unit configured to perform score prediction processing on the training sample through the target policy network model to obtain a target policy benefit corresponding to the training sample; a first error mapping unit configured to perform mapping processing on the current policy benefit based on a first loss function to obtain a first loss error for updating the parameters of the current policy generation network of the current policy network model; and a second error mapping unit configured to perform mapping processing on the current policy benefit and the target policy benefit based on a second loss function to obtain a second loss error for updating the parameters of the current policy evaluation network of the current policy network model.

[0020] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the information processing method in the above technical solutions.

[0021] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the information processing method in the above technical solutions by executing the executable instructions.

[0022] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the information processing method in the above technical solutions.

[0023] In the technical solution provided by the embodiments of the present application, a reinforcement learning model is trained based on multiple parallel model training processes, which can eliminate training bottlenecks and significantly improve the model training speed, thereby improving the information processing efficiency of processing candidate information using the policy network model. By using the trained reinforcement learning model to predict scores for the agreed display information, an information ranking score with the same dimension as the competing display information can be determined, so that the agreed display information and the competing display information can be mixedly ranked on the same standard. Based on the comprehensive comparison of the agreed display information and the competing display information, the mixed control of the two types of information is realized, so that the information display opportunities in the system can be fully utilized, and further the network resource utilization rate of information display is improved.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 FIG. shows an exemplary system architecture block diagram of an information display system applying the technical solution of the embodiments of the present application.

[0027] Figure 2 FIG. shows a flowchart of the steps of an information processing method in an embodiment of the present application.

[0028] Figure 3 FIG. shows a system framework diagram of the application of the embodiments of the present application to the mixed display of advertisements.

[0029] Figure 4 FIG. shows a flowchart of the steps of updating the parameters of the policy network model based on reinforcement learning in an embodiment of the present application.

[0030] Figure 5 FIG. shows a schematic diagram of the framework structure of reinforcement learning in an embodiment of the present application.

[0031] Figure 6 FIG. shows the model framework of a distributed reinforcement learning model in an embodiment of the present application.

[0032] Figure 7 FIG. shows a schematic diagram of the structure of the reinforcement learning model maintained by the model training process Learner in the embodiments of the present application.

[0033] Figure 8 A structural block diagram of an information processing apparatus provided by an embodiment of the present application is schematically shown.

[0034] Figure 9 A structural block diagram of a computer system of an electronic device suitable for implementing an embodiment of the present application is schematically shown. Detailed implementation manners

[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0036] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will recognize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0037] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0038] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0039] Embodiments of the present application relate to a solution for controlling an information display strategy through artificial intelligence technology in information display scenarios such as advertisement placement and advertisement playback.

[0040] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0041] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0042] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0043] Reinforcement Learning (RL), also known as reward learning, evaluation learning, or enhancement learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem that an agent achieves maximum reward or realizes a specific goal by learning strategies during the interaction with the environment. Among them, an agent is a computational entity with its own computational resources and a locally self-contained behavior control mechanism. This computational entity can be, for example, a process, a thread, a computer system, a simulator, a robot, etc. An agent can decide and control its own behavior based on its internal state and the environmental information it perceives without direct external manipulation.

[0044] Reinforcement learning is developed from theories such as animal learning and parameter perturbation adaptive control. Its basic principle is:

[0045] If an agent's behavior strategy leads to a positive reward (reinforcement signal) in the environment, then the agent's tendency to use this behavior strategy in the future will be strengthened. The agent's goal is to find the optimal strategy in each discrete state to maximize the expected discounted reward.

[0046] Reinforcement learning regards learning as a trial-and-error evaluation process. The agent selects an action for the environment. After the environment accepts the action, its state changes and a reinforcement signal (reward or punishment) is generated and fed back to the agent. The agent selects the next action based on the reinforcement signal and the current state of the environment. The principle of selection is to increase the probability of receiving positive reinforcement (reward). The selected action not only affects the immediate reinforcement value, but also affects the state of the environment at the next moment and the final reinforcement value.

[0047] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, etc. I believe that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0048] Figure 1 An exemplary system architecture block diagram of an information display system using the technical solution of the embodiment of the present application is shown. Figure 1 As shown, the information display system 100 may include a terminal device 110 , a network 120 and a server 130 .

[0049] The terminal device 110 may be an electronic device having a network connection function and having an information display application corresponding to the server 130 installed, such as a smart phone, a tablet computer, a laptop computer, a desktop computer, an e-book reader, smart glasses, a smart watch, etc.

[0050] In an embodiment of the present application, the above-mentioned information display application may include any application that provides information recommendation locations, such as, including but not limited to video playback applications, video live broadcast applications, news applications, reading applications, music playback applications, social applications, game applications, communication applications, browser applications, and applications that come with the terminal system (such as the negative one screen), etc.

[0051] The server 130 is a server that can provide background data support for information display applications installed on the terminal device 110. For example, it can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0052] Network 120 can be a communication medium of various connection types capable of providing a communication link between terminal device 110 and server 130. For example, it can be a wired communication link or a wireless communication link. The network is usually the Internet, but it can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, custom and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0053] According to the implementation requirements, the system architecture in the embodiments of the present application can have any number of terminal devices, networks and servers. For example, server 130 can be a server group composed of multiple server devices. In addition, the technical solutions provided by the embodiments of the present application can be applied to terminal device 110, can also be applied to server 130, or can be jointly implemented by terminal device 110 and server 130. The present application makes no special limitation on this.

[0054] The following will make a detailed description of the technical solutions such as the information processing method, information processing device, computer-readable medium and electronic device provided by the embodiments of the present application in combination with specific implementation manners.

[0055] Figure 2 The step flowchart of the information processing method in an embodiment of the present application is shown. This information processing method can be executed by terminal device 110 or server 130 in the above information display system, or can also be jointly executed by terminal device 110 and server 130. As Figure 2 shown, this information processing method mainly includes the following steps S210 to step S240.

[0056] Step S210: acquiring a candidate information set consisting of a plurality of candidate information according to the information display request, wherein the candidate information includes competitive display information that competes for display opportunities according to resource payment amounts and agreed display information that has agreed display quantity requirements.

[0057] In some optional implementations, the above-mentioned information display request may be a request generated when an information display position appears in an information display application of a terminal device. For example, when a user uses a video playback application to watch a video, it is generally necessary to play an advertisement of a specified length before the video is played. Therefore, when a user clicks to play a video on a terminal device, a corresponding information display request may be generated on the terminal device, and the terminal device may obtain a corresponding candidate information set based on the information display request. The acquisition method may be to search for candidate information in a local database of the terminal device to form a candidate information set, or to send an information display request to a server, and the server returns the candidate information set to the terminal device.

[0058] In some optional implementations, for an information display request, multiple preferred information can be screened out from all the information that can be displayed in the system to form a candidate information set corresponding to the information display request. In the advertising display scenario, the above candidate information set can also be called a refined queue corresponding to the information display request in the advertising sequence.

[0059] In the application scenario of advertising, the agreed display information can be contract advertising, and the competitive display information can be bidding advertising. Among them, contract advertising is a contract signed between the advertiser and the media party, and the media party plays a predetermined amount of advertisements (generally refers to the exposure volume of the advertisement scheduled by the advertiser, such as the scheduled number of days of exposure) to the users of the specified type of the advertiser within a specified time; if the contract is reached, the advertiser pays a certain fee to the media, and if the playback volume does not meet the standard, the media needs to compensate the advertiser; no additional fees will be charged if the playback volume exceeds the scheduled amount. Bidding advertising is that advertisers will give a bid for requests with the same orientation. There will be multiple advertisers bidding for the same request, and the advertiser with the highest bid will win the competition and get the exposure of this request.

[0060] Step S220: determining the information ranking score of each competing display information according to the resource contribution amount, where the information ranking score is used to indicate the display priority of the candidate information.

[0061] In some optional implementations, the competitive display information determines the corresponding information ranking score according to its respective resource payment amount, and the information ranking score is positively correlated with the resource payment amount. The competitive display information with a larger resource payment amount has a higher information ranking score, and the probability of obtaining the information display opportunity is also greater. For example, in the application scenario of advertising delivery, for bidding ads, the higher the advertiser's bid, the greater the probability of obtaining the advertising delivery opportunity.

[0062] Step S230: Perform score prediction processing on the agreed display information through the policy network model to obtain the information ranking scores of each piece of agreed display information; the policy network model is a reinforcement learning model trained based on multiple parallel model training processes.

[0063] For the agreed display information, due to the lack of a quantization index corresponding to the resource expenditure, the embodiment of the present application adopts the mechanism of reinforcement learning to explore the policy of the current information display environment through the policy network model, and then predicts the corresponding information ranking scores for the agreed display information. Since the information display environment is complex and changeable, and the number of candidate information is large, the reinforcement learning model trained by the embodiment of the present application based on multiple parallel model training processes can eliminate the training bottleneck and greatly improve the model training speed, thereby improving the information processing efficiency of using the policy network model to process candidate information.

[0064] Step S240: Select the target information to be displayed from the candidate information set according to the information ranking scores.

[0065] After obtaining the information ranking scores of the competing display information and the agreed display information, the target information to be displayed can be selected from the candidate information set according to the high and low scores of the information ranking scores. The information selection method can be, for example, directly selecting one or more candidate information with the highest scores as the target information; it can also be to assign corresponding information selection probabilities to each candidate information according to the information ranking scores, and then extract the target information from the candidate information set according to the probabilities.

[0066] In the information processing method provided by the embodiment of the present application, by performing score prediction on the agreed display information through the pre-trained policy network model, the information ranking scores with the same dimension as the competing display information can be determined, so that the agreed display information and the competing display information can be mixedly ranked on the same standard. On the basis of comprehensively comparing the agreed display information and the competing display information, the mixed control of the two types of information is realized, so that the information display opportunities in the system can be fully utilized, and further the network resource utilization rate of information display is improved.

[0067] Taking the application scenario of advertising placement as an example, the server records the requests for obtaining advertisements sent by each terminal within a period of time (i.e., the above-mentioned historical information display requests), and obtains multiple advertisements that match these requests from the contract advertisements and the auction advertisements to form a sample environment, and each advertisement in the sample environment corresponds to its own state data; through exploration in the sample environment for reinforcement learning, the policy network model can be obtained.

[0068] When an advertising display opportunity appears in a certain terminal at a later time, the terminal sends a request to obtain an advertisement to the server; after receiving the request, the server obtains multiple advertisements that match the request from contract advertisements and auction advertisements and forms a refined queue. Among them, for auction advertisements, the advertising revenue ecpm (expected cost per mile, the expected revenue per thousand impressions) can be predicted based on the bid of the advertiser. In the embodiments of the present application, ecpm can be used as the information sorting score of the auction advertisement. For contract advertisements, in the embodiments of the present application, the corresponding information sorting score is obtained by predicting the score through a pre-trained policy network model. After sorting the auction advertisements and contract advertisements together according to the information sorting score, the advertisement with the highest information sorting score can be selected and pushed to the terminal device for display.

[0069] Figure 3 shows the system framework diagram of the embodiments of the present application applied to mixed advertising display. As Figure 3 shown, the mixed model is in the central position, and the system input includes broadcast control parameters, TrackLog exposure data, and inventory data. The model gives parameters for auction advertisements and contract advertisements and passes them into the dictionary structure of the Feature Server, and finally takes effect in the Mixer.

[0070] Generally speaking, Figure 3 the system framework shown is divided into three main parts, namely data processing 310, mixed model 320, and online system 330. The above parts will be introduced one by one below:

[0071] The data processing 310 part includes three modules: data source, data transmission, and data processing, which complete the processing operation from the original data to the algorithm input.

[0072] The inventory data comes from the inventory prediction service, which is a detailed prediction of the future using past data, accurate to the mapping of each access request (Page View, PV) and each advertisement, and can reflect the inventory of each order on a given day. The bipartite graph is calculated based on the inventory data. Through the bipartite graph, two data can be obtained: the playback probability of the contract advertisement and the playback curve of the current day. The former gives a reference for contract quantity guarantee, and the latter gives the occupied space of the contract.

[0073] The logs are divided into two types, one is the request-level data track_log, and the other is the exposure-level data joined_exposure.

[0074] Through track_log, we can obtain the refined ranking queue of each request. Through the refined ranking queue within a time period and the expected revenue per thousand impressions (Expected Cost Per Mile, ECPM) of all ads in the queue, predicted click-through rate, filtering conditions, support strategies and other data, the reinforcement learning algorithm can simulate the online competition environment through this data. If the length of the time period (Δt) is small enough, it can be assumed that the distribution of bidding contracts in the first Δt is the same as or similar to the distribution of bidding contracts in the last Δt.

[0075] Through joined_exposure, we can get which ad is actually exposed for each request, as well as the corresponding billing and ecpm information. The reinforcement learning algorithm can obtain feedback on online advertising through this data.

[0076] The playback control of the contract is affected by a variety of playback control parameters, such as rate (probability of entering the sorting queue), theta (playback probability), etc., which are key information to assist in adjusting the contract guarantee. In related technologies, the internal sorting basis of contract push information is the playback probability (Theta, a playback control parameter). For example, if contract push information A and B both match a request, A's theta is 0.3, and B's theta is 0.6, then the playback probability of A is 0.3, and the playback probability of B is 0.6. Theta can be considered a known quantity, and the calculation method is theta = Dj / Sj, where Dj is the scheduled amount of the contract push information (the exposure of the contract push information scheduled by the advertiser), and Sj is the current inventory of the contract push information. The playback control parameter rate is a parameter for controlling the playback of advertisements. Rate = 0.5 means that this push information has a 50% chance of entering the sorting queue. Inventory means: push information is targeted. For example, if a push information is targeted at Shanghai, then only users in Shanghai can be exposed to this push information. Inventory refers to the number of visits by all users that can match this push information (not the number of users, because users may visit more than once).

[0077] The online system 330 has two parts, namely the feature server FeatureServer and the mixer Mixer. The feature server FeatureServer is referred to as fs. The scores of each advertisement obtained in this application will be transmitted to fs with other parameters (Theta, Rate). After integration, fs waits for the request of Mixer. Mixer is a complex system. The part related to the embodiment of this application is the mixing module. When a request arrives, Mixer will receive the advertising queues of the bidding and contract, and then request the display scores of each advertisement from fs, and obtain the final displayed advertisement.

[0078] In the embodiments of the present application, the policy network model is used to predict scores for the agreed display information. The prediction efficiency and accuracy of the policy network model determine the overall benefit of the final information display. Figure 4 FIG. shows a flowchart of steps for updating parameters of the policy network model based on reinforcement learning in an embodiment of the present application. As Figure 4 shown, before obtaining a candidate information set composed of multiple candidate information according to an information display request, the parameters of the policy network model can be updated according to the following steps S410 to S430.

[0079] Step S410: Obtain multiple sample sets respectively maintained by multiple parallel sample exploration processes. The sample set includes training samples obtained by the sample exploration process performing policy exploration on a sample environment related to a historical information display request.

[0080] Step S420: Read training samples from the sample set based on multiple parallel model training processes, and perform score prediction processing on the training samples through the policy network model to obtain loss errors corresponding to the training samples.

[0081] Step S430: Update the network parameters of the policy network model according to the loss errors.

[0082] A complete policy network model includes two parts: a policy generation network and a policy evaluation network. The policy generation network is used to generate an information display policy according to the current environmental state and predict the information display benefit that can be obtained by executing the information display policy in the current environmental state. The policy evaluation network is used to evaluate the information display policy generated by the policy generation network according to the overall benefit situation.

[0083] In the embodiments of the present application, the sample exploration process and the model training process are separated. The sample exploration process is to control the policy generation network to perform policy exploration on the sample environment through multiple parallel sample exploration processes to obtain training samples, and each sample exploration process does not affect each other. Therefore, a large number of training samples can be generated exponentially. The model training process is to control the complete policy network model to use the training samples for training to complete the parameter update of the model through multiple parallel model training processes, and each model training process is also independent of each other. Therefore, the training speed of the model can be increased exponentially.

[0084] In an optional implementation manner, the method for obtaining multiple sample sets respectively maintained by multiple parallel sample exploration processes in step S410 may include the following steps S411 to S413.

[0085] Step S411: Obtain a sample information set corresponding to the historical information display request, and form a sample environment with the sample information in the sample information set;

[0086] Step S412: Perform policy exploration on the sample environment through multiple parallel sample exploration processes respectively to obtain training samples corresponding to the historical information display request. The training samples include environmental state data, an information display policy corresponding to the environmental state data, and an information display benefit determined according to the environmental state data and the information display policy.

[0087] Step S413: Save the training samples obtained by the sample exploration processes to the sample set maintained by the sample exploration processes.

[0088] The training samples obtained by the sample exploration processes include three parts: environmental state data, an information display policy, and an information display benefit. After the sample information is displayed according to the information display policy, the overall sample environment will change. Therefore, the new environmental state data after the sample environment changes can be updated according to the environmental state data and the information display policy.

[0089] The environmental state data reflects the reasons for the Agent of the sample exploration process to make specific actions. The environmental state data must be able to sufficiently represent the current sample environment so that there is sufficient distinguishability between different environmental states. In a possible implementation manner, the above environmental state data may include the overall shortage rate of the agreed display information in the system and the average resource expenditure of the competing display information.

[0090] In another possible implementation manner, in order to provide sufficient distinguishability, the environmental state data may include at least one of information-level data, overall data, and traffic dimension feature data.

[0091] The information-level data includes at least one of: the identifier of the corresponding information, the identifier of the corresponding information display position, the number of times the corresponding information has been played, the playback volume requirement of the corresponding information, the playback speed of the corresponding information, and the upper limit of the playback volume of the corresponding information.

[0092] The overall data includes at least one of: the overall shortage rate of the agreed display information in the system, the average click-through rate of the agreed display information in the system, the average click-through rate of the competing display information in the system, and the average resource expenditure of the competing display information in the system.

[0093] The traffic dimension features include at least one of: the geographical data matching the corresponding information display request, the gender data matching the corresponding information display request, and the age data matching the corresponding information display request.

[0094] Herein, only the information included in the above information-level data, overall data, and traffic dimension feature data is taken as an example in the embodiments of the present application. The above information-level data, overall data, and traffic dimension feature data include but are not limited to the data listed above.

[0095] In the embodiments of the present application, the policy network model can process the environmental state data of each sample information through its internal policy generation network to obtain the display scores of each sample information; then, the policy generation model can generate an information display policy based on the display scores of each sample information.

[0096] In the embodiments of the present application, under the control of the sample exploration process, there is a policy generation network that takes the environmental state data of the sample information as input, obtains the display score of the sample information in the current state, and outputs the corresponding information display policy accordingly.

[0097] Among them, the above information display policy is also called the action in the reinforcement learning model. In reinforcement learning, the action can be obtained through multi-classification, binary classification, or regression.

[0098] 1) Multi-classification:

[0099] The action setting method of multi-classification is the most intuitive. Each decision of the Agent is to select an advertisement in the current state. However, there is a problem of too many classification targets. Taking the news video traffic as an example, the number of contract advertisements online on the same day is several thousand, and even up to twenty thousand during some special periods. There are even more competing advertisements; such a large classification model is very difficult to converge, unless the sample size is huge. However, it is very difficult for the Agent to return enough samples. Even if the advertisements that can be associated with each request can be selected to greatly reduce the number of classifications in each training, it is still very difficult to train.

[0100] 2) Binary classification:

[0101] If the advertisement mixing is regarded as a binary classification problem of selecting competing bids or contracts, in this mode, the convergence of the P network becomes easy. However, there are the following problems in this mode:

[0102] First, it is very difficult to decide which advertisement to select specifically after selecting a contract or a competing bid. Among them, for competing bids, the order with the highest bid can be selected, while for contracts, it can only be output through another model. That is to say, a secondary selection of the playback probability of the advertisements in the queue is required, and the model complexity is relatively high.

[0103] Second, it is difficult to go online. Since there is no separate service running the mixing model, let the Mixer transfer the refined queue to this service and return the order that should be placed. Therefore, currently, only the scores corresponding to the advertisement and traffic dimensions can be output, and it is very difficult to achieve this with the above binary classification method.

[0104] 3) Regression:

[0105] In the solution shown in the embodiments of the present application, the policy generation network is changed once to become a regression model. The input is an environmental state for each advertising order, and the output is the display score of the advertising order in the current state. The Agent sorts the mixed queue according to the display scores of each advertising order and selects the advertisement with the highest score.

[0106] In the embodiments of the present application, after the selected target sample information is displayed by simulating the information display policy, the state changes of each piece of information in the sample environment are realized, so as to update the environmental state data of the sample information in the sample environment.

[0107] Among them, updating the environmental state data of the sample information in the sample environment includes updating the state of each sample information (corresponding to the information-level data in the above environmental state data), and updating the overall state of each sample information in the sample environment (corresponding to the overall data in the above environmental state data).

[0108] The design of information display reward is the core part of reinforcement learning. A good reward should be able to reflect the task goal and be easy to converge. In the mixed sorting task mode involved in the embodiments of the present application, the overall system reward includes three parts: contract guaranteed quantity, guaranteed click-through rate, and increased auction ecpm. Among them, the contract guaranteed quantity is represented by the overall shortage rate of the agreed display information in the sample environment, the guaranteed click-through rate is represented by the score of the average click-through rate of the agreed display information in the sample environment, and the increased auction ecpm is represented by the score of the average resource expenditure of the competitive display information in the sample environment.

[0109] Among them, the overall shortage rate of the agreed display information in the sample environment can be obtained from the played volume and the should-be-played volume of each agreed display information; the average click-through rate of the above-mentioned agreed display information can be obtained by taking the average of the predicted click-through rates of the agreed display information; the average resource expenditure of the above-mentioned competitive display information can be obtained by taking the average of the predicted resource expenditures of the agreed display information.

[0110] Among them, the above-mentioned resource expenditure can be the bid price of the auction advertisement for the display opportunity corresponding to each information display request.

[0111] In the embodiments of the present application, there is a linear relationship between the click-through rate and the ecpm. The traffic with a higher bid price often has a higher click-through rate, which is more convenient to compare on one dimension, while there is no such relationship between the contract guaranteed quantity and the increased ecpm. For example, a 20% shortage and an average ecpm of 18.9 yuan are almost impossible to compare. That is to say, shortage and ecpm are two completely different concepts and it is difficult to normalize them to the same dimension. This is the difficulty of setting the mixed sorting reward.

[0112] In an alternative embodiment, the embodiments of the present application can process the overall shortfall rate, average click-through rate, and average resource expenditure by setting weights W to obtain the score lackScore of the overall shortfall rate, the score ctrScore of the average click-through rate, and the score ecpm of the average resource expenditure, that is, by experience, the reward is normalized to the same dimension. For example, the information display benefit can be calculated using the following formula:

[0113]

[0114] where n is the number of consecutive historical information display requests, also known as the exploration steps; Gamma is the decay coefficient for each step. The higher this coefficient, the more the model values long-term benefits, and the lower this coefficient, the more the model values short-term benefits. W lack 、W ctr and W ecpm are the corresponding preset weights.

[0115] The processing method of normalizing the reward to the same dimension by experience has great limitations. The setting of weights cannot adapt to all situations. As the state (shortfall, ecpm) changes, the weights should change accordingly, but it is difficult to find a formula to represent this change.

[0116] To solve the problem of comparability between guaranteeing quantity and ecpm, the embodiments of the present application consider the essence of reinforcement learning, that is, reinforcement learning is a process of continuously simulating human decision-making through the "memory ability" of the model, and recording the optimal decision corresponding to the state. Returning to the problem of mixed arrangement itself, the purpose of the mixed arrangement model is to improve the ecpm of bidding while ensuring that the contract exposure volume is equal to the contract priority strategy. Then, the reinforcement learning goal can be set as the improvement ratio of the mixed arrangement strategy to the contract priority strategy. The ratios are of the same dimension, thus solving this problem. Among them, the contract priority strategy means that when there are available contract advertisements, the contract advertisement with the highest display score is preferentially selected, and when there are no available contract advertisements, the bidding advertisement with the highest display score is selected.

[0117] In an alternative embodiment, the policy generation network includes a priority policy network and a mixed arrangement policy network; the sample display policy of the sample information includes the priority display policy output by the priority policy network and the mixed arrangement display policy output by the mixed arrangement policy network; the priority display policy is a policy of preferentially selecting the target sample information from the agreed display information; the mixed arrangement display policy is a policy of mixing and sorting each sample information based on the display score and selecting the target sample information. On this basis, the method for determining the information display benefit in the embodiments of the present application can include:

[0118] Obtain a first gain parameter score based on first state data, where the first state data is the state data before and after updating the state data of sample information in the sample environment through the priority display strategy;

[0119] Obtain a second gain parameter score based on second state data, where the second state data is the state data before and after updating the state data of sample information in the sample environment through the mixed display strategy;

[0120] Obtain the improvement ratio of the second gain parameter score relative to the first gain parameter score, so as to determine the information display benefit based on the improvement ratio.

[0121] Figure 5 The schematic diagram of the reinforcement learning framework structure in an embodiment of the present application is shown. As Figure 5 shown, the first agent 501 corresponds to the priority policy network, and its policy for selecting information is the contract priority policy, which will return a reward_base. The policy of the second agent 502 is the mixed display policy, which will also return a reward. Among them, the improvement of the reward relative to the reward_base will be returned as the final reward function value.

[0122] When the reinforcement learning goal is set to the improvement ratio of the mixed display policy to the contract priority policy, the formula for determining the information display benefit is as follows:

[0123]

[0124] Among them, n is the number of consecutive historical information display requests, also known as the exploration steps; Gamma is the decay coefficient for each step. The higher this coefficient, the more the model values long-term benefits, and the lower this coefficient, the more the model values short-term benefits. lackBase, ctrBase, and ecpmBase are the lack quantity, click-through rate, and expected revenue per thousand impressions obtained by the first agent 501 based on the contract priority policy at each step of exploration, respectively. lack, ctr, and ecpm are the lack quantity, click-through rate, and expected revenue per thousand impressions obtained by the second agent 502 based on the mixed display policy at each step of exploration, respectively. W lack 、W ctr and W ecpm are the preset weights corresponding to the lack quantity, click-through rate, and expected revenue per thousand impressions, respectively.

[0125] In one embodiment of the present application, saving the training samples obtained by the sample exploration process to the sample set maintained by the sample exploration process in step S413 may include: monitoring the number of training samples obtained by the sample exploration process to determine whether the number of training samples reaches a preset number threshold; when it is monitored that the number of training samples reaches the preset number threshold, writing the training samples to the sample set shared memory corresponding to the sample exploration process, and the sample set shared memory corresponds one-to-one with the sample set maintained by the sample exploration process.

[0126] In one embodiment of the present application, the method of writing the training samples to the sample set shared memory corresponding to the sample exploration process may include:

[0127] Obtaining the data storage amount of the sample set shared memory corresponding to the sample exploration process;

[0128] When the data storage amount does not reach the maximum capacity of the sample set shared memory, writing the training samples sequentially to the blank area of the sample set shared memory;

[0129] When the data storage amount reaches the maximum capacity of the sample set shared memory, writing the training samples randomly to any storage location in the sample set shared memory so that the training samples randomly overwrite the existing data in the sample set shared memory.

[0130] When the sample set shared memory is already full of data, the embodiments of the present application adopt a random overwrite method to overwrite and update the training samples in the sample set shared memory, which can maintain the diversity of the training samples.

[0131] In one embodiment of the present application, the method of writing the training samples to the sample set shared memory corresponding to the sample exploration process may include:

[0132] Obtaining the data storage amount of the sample set shared memory corresponding to the sample exploration process;

[0133] When the data storage amount does not reach the maximum capacity of the sample set shared memory, writing the training samples sequentially to the blank area of the sample set shared memory;

[0134] When the data storage amount reaches the maximum capacity of the sample set shared memory, writing the training samples sequentially to the sample set shared memory so that the training samples overwrite the existing data in the sample set shared memory in the order of data writing time.

[0135] When the sample set shared memory is already full of data, the embodiments of the present application adopt a sequential overwrite method to overwrite and update the training samples in the sample set shared memory in the order of writing time, which can improve the timeliness of the training samples.

[0136] Multiple parallel sample exploration processes can write the explored training samples into the sample set shared memory corresponding to each of them, and parallel model training threads can read the training samples from the sample set shared memory according to the training needs. To avoid conflicts in writing and reading training samples, embodiments of the present application can configure corresponding status flag bits for each sample set shared memory. On this basis, embodiments of the present application can monitor the data storage amount and data writing status of the sample set shared memory in real time; assign values to the status flag bits of the sample set shared memory according to the monitored data storage amount and data writing status, and the status flag bits are used to indicate whether the data in the sample set shared memory is readable.

[0137] In an embodiment of the present application, assigning values to the status flag bits of the sample set shared memory according to the monitored data storage amount and data writing status may include:

[0138] When it is monitored that data is being written into the sample set shared memory, assign the status flag bit of the sample set shared memory to a first value; the status flag bit with the first value is used to indicate that the sample set shared memory is in a state where the data is not readable;

[0139] When it is monitored that the data writing is completed and the data storage amount does not reach the maximum capacity of the sample set shared memory, assign the status flag bit of the sample set shared memory to a first value;

[0140] When it is monitored that the data writing is completed and the data storage amount reaches the maximum capacity of the sample set shared memory, assign the status flag bit of the sample set shared memory to a second value; the status flag bit with the second value is used to indicate that the sample set shared memory is in a state where the data is readable.

[0141] In an embodiment of the present application, the step of reading training samples from the sample set in step S420 based on multiple parallel model training processes may include: polling the sample sets maintained by each sample exploration process based on multiple parallel model training processes to determine whether the sample sets are in a state where the data is readable; when the sample set is in a state where the data is readable, read the data from the sample set.

[0142] In an embodiment of the present application, the step of updating the network parameters of the policy network model according to the loss error in step S430 may include: calculating the error gradients of the policy network models maintained by each model training process respectively according to the loss errors obtained by training with multiple parallel model training processes; writing the error gradients calculated by each model training process into the model parameter shared memory to update the network parameters of the policy network model stored in the model parameter shared memory according to the error gradients.

[0143] In one embodiment of the present application, the policy network model includes a current policy network model and a target policy network model that is the training target of the current policy network model; updating the network parameters of the policy network model stored in the model parameter shared memory according to the error gradient includes: updating the network parameters of the current policy network model stored in the model parameter shared memory according to the error gradient; when a preset target update condition is satisfied, updating the network parameters of the target policy network according to the network parameters of the current policy network model.

[0144] In one embodiment of the present application, the policy network model includes a current policy network model and a target policy network model that is the training target of the current policy network model. The current policy network model includes a current policy generation network for generating an information display policy and a current policy evaluation network for evaluating the information display policy. The target policy network model includes a target policy generation network that is the training target of the current policy generation network and a target policy evaluation network that is the training target of the current policy evaluation network.

[0145] The method for obtaining the loss error corresponding to the training sample by performing score prediction processing on the training sample through the policy network model may include: performing score prediction processing on the training sample through the current policy network model to obtain the current policy return corresponding to the training sample; performing score prediction processing on the training sample through the target policy network model to obtain the target policy return corresponding to the training sample; performing mapping processing on the current policy return based on the first loss function to obtain the first loss error for updating the parameters of the current policy generation network of the current policy network model; performing mapping processing on the current policy return and the target policy return based on the second loss function to obtain the second loss error for updating the parameters of the current policy evaluation network of the current policy network model.

[0146] Figure 6 shows the model framework of the distributed reinforcement learning model in one embodiment of the present application. As Figure 6 shown, under the main process of reinforcement learning, there are multiple parallel sample exploration processes Agent and multiple parallel model training processes Learner distributed. The model training process Learner is an independently running process for performing training operations on the complete policy network model to update its model parameters. The policy network model includes a policy generation network and a policy evaluation network.

[0147] The sample exploration process Agent can continuously explore the environment through the policy generation network to generate policy samples to be evaluated, while the model training process Learner can evaluate the policy samples obtained through exploration through the policy evaluation network and update the network parameters of the policy generation network and the policy evaluation network based on the evaluation results. In the embodiments of the present application, the sample exploration process Agent and the model training process Learner independently execute the sample exploration operation and the model training operation, which can realize the parallelization of sample exploration and model training, enabling the two processes to continuously run. Even if any one of the processes has problems such as process blocking or slow process, it will not affect the overall reinforcement learning process of the policy network model, so the training efficiency of the model can be greatly improved.

[0148] The number of sample exploration process Agents is not unique, and multiple can be started simultaneously. For example, 5 parallel sample exploration process Agents can be started simultaneously. The sample exploration process Agent maintains a separate policy generation network Actor, a log data environment (En), and a tensorflow context (tf context). The sample exploration process Agent inherits the Actor network parameters trained by the model training process Learner. The parameters are randomly initialized at the beginning and continuously explore the environment, writing the samples into the corresponding s_memory. When writing the data, the status flag is set to 0, and after the data is written and s_memory is full, the status flag is set to 1.

[0149] After each data writing is completed, the sample exploration process Agent will pull the latest Actor network parameters from the ANP_memory, update its own Actor network, and then continue to explore the environment with the latest parameters.

[0150] s_memory is the shared memory corresponding to the sample pool. Each Agent corresponds to a separate sample pool, which is jointly maintained by the Agent and the Learner. There is a flag bit flag in s_memory as a lock. It is set to 1 after each data writing and storage is completed, and set to 0 when writing is in progress or not full.

[0151] LACNP_Memory (Learner ActorCriticNetParam Memory) is the shared memory corresponding to the complete network parameters, responsible for maintaining the network parameters of the Actor and Critic trained by the Learner.

[0152] ANP_Memory (ActorNetParam Memory) is the shared memory corresponding to the Actor network parameters, responsible for maintaining the latest Actor network parameters. After the Learner training generates the latest Actor parameters, they can be written into ANP_Memory, and the Agent can read from it and update its own Actor network parameters.

[0153] The shared memory is a one-dimensional array with a fixed type that only supports basic C language types (int, float, char, etc.). Whether it is samples or tensorflow network parameters, they need to be processed and encoded into a specified format before being stored in the shared memory. This process is called serialization. Similarly, when pulling data, it also needs to be decoded and processed to parse it into data that can be processed by the Learner and Agent processes.

[0154] The model training process Learner maintains a complete reinforcement learning model (including the policy generation network Actor and the policy evaluation network Critic), a log data environment, and a tensorflow context, but does not maintain the network parameters of the reinforcement learning model. All these network parameters are stored in the LACNP_Memory shared memory.

[0155] The Learner is only responsible for training the network, completing the training steps in the reinforcement learning process, and maintaining three shared memory structures: s_memory, ANP_memory, and LACNP_Memory. The Learner continuously polls the s_memory maintained by the Agent. Once data can be read, it reads the data and conducts training. It should be noted that during the training process, the Learner is only responsible for calculating the gradients and then using the gradients to update the network parameters stored in LACNP_Memory.

[0156] In the embodiments of this application, when multiple parallel sample exploration processes Agent are used for environment exploration, policy samples can be continuously explored. Moreover, different sample exploration processes Agent may also explore different policy samples when facing the same environmental data. Therefore, the diversity of policy samples can be improved while increasing the efficiency of policy sample generation.

[0157] Meanwhile, the embodiments of the present application adopt multiple parallel model training processes, namely Learner, to train a complete policy network model, which changes the traditional irreversible training mechanism based on time linearity in reinforcement learning technology. This is equivalent to performing model training simultaneously on multiple parallel timelines, and then summarizing the network parameters corresponding to each model training process through shared content. Therefore, the training progress of the model can be further accelerated. In addition, the multiple parallel model training processes, Learner, operate independently of each other, and the environmental data and policy samples they face are all different. Thus, the effect of hybrid training the policy network models under different training progress can be achieved through the policy samples generated under different training progress. This training method that disrupts the timeline can improve the training efficiency while avoiding the problem of model overfitting and enhancing the robustness of the model.

[0158] Figure 7 shows a schematic structural diagram of the reinforcement learning model maintained by the model training process Learner in the embodiments of the present application. As Figure 7 shown, four networks, Actor, Actor_, Critic, and Critic_, are maintained in Learner. Among them, Actor is responsible for generating actions, that is, order scoring, and Critic is responsible for evaluating this scoring. Actor_ and Critic_ are the target networks of Actor and Critic respectively. The parameters of these two networks are slowly updated from the Actor and Critic networks and can be considered as stable versions of these two networks.

[0159] The Critic network evaluates the reward of the action given by the Actor network. At the beginning, the Critic does not know what the real reward is and needs to give a target. The embodiments of the present application use the Critic_ network to achieve this purpose. Assuming that the reward given by the Critic_ network is correct, then the goal of the Critic is to continuously approach this target. Therefore, the loss error of the Critic can be expressed as loss(Q, Q_). The Critic_ needs an Action_ to give the reward evaluation, and naturally an Actor_ network is required for the same reason. Our goal is to make the action given by the Actor network maximize the reward, so the reward is Q, and the loss error of the Actor network is loss(Q).

[0160] Leaner can calculate the gradient by taking the derivative after inputting the sample through these four networks and two losses, and then update the parameters stored in LACNP_Memory. Taking the parameter θ i as an example, the sample is x i , and the learning rate is α. The update formula of the network parameters is as follows:

[0161]

[0162] Each process Leaner performs parallel computing Then update the parameter θ i , since the calculation method in the above formula is very fast, within 1 ms, the main computational workload is above, so using multiple parallel threads will greatly improve the training speed.

[0163] After a preset number of training steps, the main process will write the Actor parameters in LACNP_Memory to ANP_Memory for the process Agent to call.

[0164] In an application scenario of this application, the steps for distributed training are as follows:

[0165] Main process:

[0166] (1) Obtain playback data and calculate data such as current playable.

[0167] (2) Initialize all shared memories and parameters.

[0168] (3) Instantiate Agent and Leaner.

[0169] (4) Start N Agent processes and M Leaner processes.

[0170] (5) After the model converges, output the score corresponding to each order.

[0171] Agent process:

[0172] (1) Continuously explore the environment. Write K samples to the shared memory each time they are explored. The writing method is circular writing. When it is full, overwrite the oldest sample from the head. Before it is full, the shared memory flag is set to 0.

[0173] (2) Pull the latest network parameters to update its own Actor network.

[0174] Learner process:

[0175] (1) Poll the s_memory corresponding to each agent

[0176] (2) If it is found that a certain mem is readable, a batch of data will be randomly read

[0177] (3) Calculate the gradient and update the parameters of each network in LACNP_Memory

[0178] (4) Calculate the total return under the current parameters for each round and save the optimal parameters

[0179] (5) Write the latest network parameters into the shared memory every X steps

[0180] (6) After the model converges, calculate the score using the optimal parameters

[0181] Among them, N represents the number of Agent processes started, M is the number of Learner processes, K is the number of samples written at one time, and X is the number of steps to write the network parameters to ANP_Memory every few steps. For example, N = 5, M = 10, K = 2000, X = 1000.

[0182] In the embodiments of the present application, both the model training step and the exploration step of generating samples are fully parallelized, greatly improving the operation efficiency of the model. The training speed of the Learner is increased by 5 times. Since the parameters need to be updated in the shared memory, there are some communication and locking overheads, and the overall improvement is slightly lower than the Learner overhead. However, compared with the traditional sequential reinforcement learning method, the model parameter update speed can be increased by 2 - 3 times.

[0183] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0184] The following introduces the device embodiments of the present application, which can be used to execute the information processing method in the above embodiments of the present application. Figure 8 Schematically shows the structural block diagram of the information processing device provided by the embodiments of the present application. As Figure 8 shown, the information processing device 800 mainly may include: a candidate information acquisition module 810, configured to acquire a candidate information set composed of multiple candidate information according to an information display request, where the candidate information includes competitive display information that competes for a display opportunity according to the resource expenditure amount and agreed display information with agreed display quantity requirements; a first score acquisition module 820, configured to determine an information ranking score for each of the competitive display information according to the resource expenditure amount, where the information ranking score is used to represent the display priority of the candidate information; a second score acquisition module 830, configured to perform a score prediction process on the agreed display information through a policy network model to obtain an information ranking score for each of the agreed display information; the policy network model is a reinforcement learning model trained based on multiple parallel model training processes; a target information selection module 840, configured to select target information to be displayed from the candidate information set according to the information ranking score.

[0185] In some embodiments of the present application, based on the above embodiments, the information processing device 800 further includes: a sample exploration module configured to obtain a plurality of sample sets respectively maintained by a plurality of parallel sample exploration processes, where the sample sets include training samples obtained by the sample exploration processes through policy exploration of a sample environment related to a historical information display request; a model training module configured to read training samples from the sample sets based on a plurality of parallel model training processes, and perform score prediction processing on the training samples through the policy network model to obtain a loss error corresponding to the training samples; and a parameter update module configured to update network parameters of the policy network model according to the loss error.

[0186] In some embodiments of the present application, based on the above embodiments, the sample exploration module includes: a set acquisition unit configured to obtain a sample information set corresponding to a historical information display request, and form a sample environment with the sample information in the sample information set; a policy exploration unit configured to perform policy exploration on the sample environment respectively through a plurality of parallel sample exploration processes to obtain training samples corresponding to the historical information display request, where the training samples include environmental state data, an information display policy corresponding to the environmental state data, and an information display benefit determined according to the environmental state data and the information display policy; and a sample storage unit configured to store the training samples explored by the sample exploration processes in the sample sets maintained by the sample exploration processes.

[0187] In some embodiments of the present application, based on the above embodiments, the sample storage unit includes: a sample quantity monitoring subunit configured to monitor the sample quantity of the training samples explored by the sample exploration processes to determine whether the sample quantity of the training samples reaches a preset quantity threshold; and a sample writing subunit configured to, when it is monitored that the sample quantity of the training samples reaches the preset quantity threshold, write the training samples into a sample set shared memory corresponding to the sample exploration process, where the sample set shared memory corresponds one-to-one to the sample sets maintained by the sample exploration processes.

[0188] In some embodiments of the present application, based on the above embodiments, the sample writing subunit includes: a data storage amount acquisition subunit configured to acquire the data storage amount of the sample set shared memory corresponding to the sample exploration process; a sample sequential writing subunit configured to sequentially write the training samples into the blank area of the sample set shared memory when the data storage amount does not reach the maximum capacity of the sample set shared memory; a sample random overwrite subunit configured to randomly write the training samples into any storage location of the sample set shared memory when the data storage amount reaches the maximum capacity of the sample set shared memory, so that the training samples randomly overwrite the existing data in the sample set shared memory.

[0189] In some embodiments of the present application, based on the above embodiments, the sample writing subunit includes: a data storage amount acquisition subunit configured to acquire the data storage amount of the sample set shared memory corresponding to the sample exploration process; a sample sequential writing subunit configured to sequentially write the training samples into the blank area of the sample set shared memory when the data storage amount does not reach the maximum capacity of the sample set shared memory; a sample sequential overwrite subunit configured to sequentially write the training samples into the sample set shared memory when the data storage amount reaches the maximum capacity of the sample set shared memory, so that the training samples sequentially overwrite the existing data in the sample set shared memory according to the data writing time.

[0190] In some embodiments of the present application, based on the above embodiments, the information processing device further includes: a status monitoring module configured to monitor in real time the data storage amount and the data writing status of the sample set shared memory; an identification bit assignment module configured to assign an identification bit of the status of the sample set shared memory according to the monitored data storage amount and the data writing status, and the identification bit of the status is used to indicate whether the data in the sample set shared memory is readable.

[0191] In some embodiments of the present application, based on the above embodiments, the identification bit assignment module includes: a first assignment unit configured to assign a first numerical value to the status identification bit of the sample set shared memory when it is detected that data is being written into the sample set shared memory; the status identification bit with the first numerical value is used to indicate that the sample set shared memory is in a state where data cannot be read; a second assignment unit configured to assign a first numerical value to the status identification bit of the sample set shared memory when it is detected that the data writing is completed and the data storage amount does not reach the maximum capacity of the sample set shared memory; a third assignment unit configured to assign a second numerical value to the status identification bit of the sample set shared memory when it is detected that the data writing is completed and the data storage amount reaches the maximum capacity of the sample set shared memory; the status identification bit with the second numerical value is used to indicate that the sample set shared memory is in a state where data can be read.

[0192] In some embodiments of the present application, based on the above embodiments, the model training module includes: a status polling unit configured to poll the sample sets maintained by each sample exploration process based on multiple parallel model training processes to determine whether the sample sets are in a state where data can be read; a data reading unit configured to read data from the sample sets when the sample sets are in a state where data can be read.

[0193] In some embodiments of the present application, based on the above embodiments, the parameter update module includes: an error gradient calculation unit configured to calculate the error gradients of the policy network models maintained by each of the model training processes respectively according to the loss errors obtained by training with multiple parallel model training processes; a network parameter update unit configured to write the error gradients calculated by each of the model training processes into the model parameter shared memory to update the network parameters of the policy network model stored in the model parameter shared memory according to the error gradients.

[0194] In some embodiments of the present application, based on the above embodiments, the policy network model includes a current policy network model and a target policy network model that is the training target of the current policy network model; the network parameter update unit includes: a first parameter update subunit configured to update the network parameters of the current policy network model stored in the model parameter shared memory according to the error gradients; a second parameter update subunit configured to update the network parameters of the target policy network according to the network parameters of the current policy network model when a preset target update condition is satisfied.

[0195] In some embodiments of the present application, based on the above embodiments, the policy network model includes a current policy network model and a target policy network model that serves as the training target of the current policy network model. The current policy network model includes a current policy generation network for generating an information display policy and a current policy evaluation network for evaluating the information display policy. The target policy network model includes a target policy generation network that serves as the training target of the current policy generation network and a target policy evaluation network that serves as the training target of the current policy evaluation network. The model training module includes: a first benefit prediction unit configured to perform score prediction processing on the training sample through the current policy network model to obtain a current policy benefit corresponding to the training sample; a second benefit prediction unit configured to perform score prediction processing on the training sample through the target policy network model to obtain a target policy benefit corresponding to the training sample; a first error mapping unit configured to perform mapping processing on the current policy benefit based on a first loss function to obtain a first loss error for updating the parameters of the current policy generation network of the current policy network model; a second error mapping unit configured to perform mapping processing on the current policy benefit and the target policy benefit based on a second loss function to obtain a second loss error for updating the parameters of the current policy evaluation network of the current policy network model.

[0196] The specific details of the information processing device provided in each embodiment of the present application have been described in detail in the corresponding method embodiments, and will not be elaborated here.

[0197] Figure 9 Schematically shown is a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.

[0198] It should be noted that Figure 9 The computer system 900 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0199] Such as Figure 9As shown, computer system 900 includes a central processing unit 901 (CPU), which can perform various appropriate actions and processes according to a program stored in read-only memory 902 (ROM) or a program loaded from storage section 908 into random access memory 903 (RAM). In random access memory 903, various programs and data required for system operation are also stored. The central processing unit 901, read-only memory 902, and random access memory 903 are connected to each other via bus 904. Input / output interface 905 (Input / Output interface, i.e., I / O interface) is also connected to bus 904.

[0200] The following components are connected to input / output interface 905: input section 906 including a keyboard, a mouse, etc.; output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; storage section 908 including a hard disk, etc.; and communication section 909 including a network interface card such as a local area network card, a modem, etc. Communication section 909 performs communication processing via a network such as the Internet. Drive 910 is also connected to input / output interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on drive 910 as needed so that a computer program read from it can be installed into storage section 908 as needed.

[0201] Specifically, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via communication section 909, and / or installed from removable medium 911. When the computer program is executed by central processing unit 901, various functions defined in the system of the present application are executed.

[0202] It should be noted that the computer-readable medium shown in the embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0203] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0204] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0205] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by the way of software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0206] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0207] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. An information processing method, characterized in that, the method includes: obtaining a candidate information set composed of multiple candidate information according to an information display request, where the candidate information includes competitive display information that competes for display opportunities according to the amount of resource expenditure and agreed display information with agreed display quantity requirements; determining an information sorting score for each of the competitive display information according to the amount of resource expenditure, where the information sorting score is used to represent the display priority of the candidate information; performing score prediction on the agreed display information through a pre-trained policy network model to obtain an information sorting score for each of the agreed display information having the same dimension as the competitive display information; the policy network model is a reinforcement learning model trained based on multiple parallel model training processes; selecting target information to be displayed from the candidate information set according to the information sorting score of the competitive display information and the information sorting score of the agreed display information.

2. The information processing method according to claim 1, characterized in that, before obtaining a candidate information set composed of multiple candidate information according to an information display request, the method further includes: obtaining multiple sample sets respectively maintained by multiple parallel sample exploration processes, where the sample set includes training samples obtained by the sample exploration process performing policy exploration on a sample environment related to a historical information display request; reading training samples from the sample set based on multiple parallel model training processes, and performing score prediction processing on the training samples through the policy network model to obtain a loss error corresponding to the training samples; updating the network parameters of the policy network model according to the loss error.

3. The information processing method according to claim 2, characterized in that, the obtaining multiple sample sets respectively maintained by multiple sample exploration processes includes: obtaining a sample information set corresponding to a historical information display request, and forming a sample environment with the sample information in the sample information set; respectively performing policy exploration on the sample environment through multiple parallel sample exploration processes to obtain training samples corresponding to the historical information display request, where the training samples include environmental state data, an information display policy corresponding to the environmental state data, and an information display benefit determined according to the environmental state data and the information display policy; saving the training samples obtained by the sample exploration process to the sample set maintained by the sample exploration process.

4. The information processing method according to claim 3, characterized in that, the saving the training samples obtained by the sample exploration process to the sample set maintained by the sample exploration process includes: monitoring the number of samples of the training samples obtained by the sample exploration process to determine whether the number of samples of the training samples reaches a preset number threshold; when it is monitored that the number of samples of the training samples reaches the preset number threshold, writing the training samples into a sample set shared memory corresponding to the sample exploration process, where the sample set shared memory corresponds one-to-one to the sample set maintained by the sample exploration process.

5. The information processing method according to claim 4, wherein, the step of writing the training samples into the sample set shared memory corresponding to the sample exploration process includes: obtaining the data storage amount of the sample set shared memory corresponding to the sample exploration process; when the data storage amount does not reach the maximum capacity of the sample set shared memory, sequentially writing the training samples into the blank area of the sample set shared memory; when the data storage amount reaches the maximum capacity of the sample set shared memory, randomly writing the training samples into any storage location of the sample set shared memory, so that the training samples randomly overwrite the existing data in the sample set shared memory.

6. The information processing method according to claim 4, wherein, the step of writing the training samples into the sample set shared memory corresponding to the sample exploration process includes: obtaining the data storage amount of the sample set shared memory corresponding to the sample exploration process; when the data storage amount does not reach the maximum capacity of the sample set shared memory, sequentially writing the training samples into the blank area of the sample set shared memory; when the data storage amount reaches the maximum capacity of the sample set shared memory, sequentially writing the training samples into the sample set shared memory, so that the training samples sequentially overwrite the existing data in the sample set shared memory according to the data writing time.

7. The information processing method according to claim 4, wherein, the method further includes: real-time monitoring of the data storage amount and data writing status of the sample set shared memory; assigning a status flag bit to the sample set shared memory according to the monitored data storage amount and data writing status, and the status flag bit is used to indicate whether the data in the sample set shared memory is readable.

8. The information processing method according to claim 7, wherein, the step of assigning a status flag bit to the sample set shared memory according to the monitored data storage amount and data writing status includes: when it is monitored that data is being written into the sample set shared memory, assigning the status flag bit of the sample set shared memory to a first value; the status flag bit with the first value is used to indicate that the sample set shared memory is in a state where the data is not readable; when it is monitored that the data writing is completed and the data storage amount does not reach the maximum capacity of the sample set shared memory, assigning the status flag bit of the sample set shared memory to the first value; when it is monitored that the data writing is completed and the data storage amount reaches the maximum capacity of the sample set shared memory, assigning the status flag bit of the sample set shared memory to a second value; the status flag bit with the second value is used to indicate that the sample set shared memory is in a state where the data is readable.

9. The information processing method according to claim 2, wherein, the step of reading training samples from the sample set based on multiple parallel model training processes includes: Polling the sample set maintained by each sample exploration process based on multiple parallel model training processes to determine whether the sample set is in a state where data can be read; When the sample set is in a state where data can be read, read data from the sample set.

10. The information processing method according to claim 2, wherein, Updating the network parameters of the policy network model according to the loss error includes: Calculating the error gradients of the policy network models maintained by each of the model training processes respectively according to the loss errors obtained by training with multiple parallel model training processes; Writing the error gradients calculated by each of the model training processes into the model parameter shared memory to update the network parameters of the policy network model stored in the model parameter shared memory according to the error gradients.

11. The information processing method according to claim 10, wherein, The policy network model includes a current policy network model and a target policy network model that is the training target of the current policy network model; updating the network parameters of the policy network model stored in the model parameter shared memory according to the error gradient includes: Updating the network parameters of the current policy network model stored in the model parameter shared memory according to the error gradient; When a preset target update condition is satisfied, updating the network parameters of the target policy network according to the network parameters of the current policy network model.

12. The information processing method according to claim 2, wherein, The policy network model includes a current policy network model and a target policy network model that is the training target of the current policy network model. The current policy network model includes a current policy generation network for generating an information display policy and a current policy evaluation network for evaluating the information display policy. The target policy network model includes a target policy generation network that is the training target of the current policy generation network and a target policy evaluation network that is the training target of the current policy evaluation network; The process of performing score prediction processing on the training sample through the policy network model to obtain a loss error corresponding to the training sample includes: Performing score prediction processing on the training sample through the current policy network model to obtain a current policy return corresponding to the training sample; Performing score prediction processing on the training sample through the target policy network model to obtain a target policy return corresponding to the training sample; Performing mapping processing on the current policy return based on a first loss function to obtain a first loss error for updating the parameters of the current policy generation network of the current policy network model; Performing mapping processing on the current policy return and the target policy return based on a second loss function to obtain a second loss error for updating the parameters of the current policy evaluation network of the current policy network model.

13. An information processing device, wherein, The device includes: A candidate information acquisition module, configured to acquire a candidate information set composed of multiple candidate information according to an information display request, where the candidate information includes competitive display information that competes for a display opportunity according to the amount of resource contribution and agreed display information with agreed display quantity requirements; A first score acquisition module, configured to determine an information ranking score for each of the competitive display information according to the amount of resource contribution, where the information ranking score is used to represent the display priority of the candidate information; A second score acquisition module, configured to perform score prediction on the agreed display information through a pre-trained policy network model to obtain an information ranking score for each of the agreed display information having the same dimension as the competitive display information; the policy network model is a reinforcement learning model trained based on multiple parallel model training processes; A target information selection module, configured to select target information to be displayed from the candidate information set according to the information ranking score of the competitive display information and the information ranking score of the agreed display information.

14. A computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the information processing method according to any one of claims 1 to 12 is implemented.

15. An electronic device, characterized in that, it includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the information processing method according to any one of claims 1 to 12 by executing the executable instructions.

16. A computer program product, including computer instructions, characterized in that, when the computer instructions are executed by a processor, the information processing method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Method, device and system of multi-learning subject parallel training model

    CN104980518A

  • Advertisement delivery method and system, server, and computer readable storage medium

    CN108898436A