Information regulation method and device

Through the regulatory model of reinforcement learning training, dynamic regulation is performed using the interactive information and effects of the regulatory information set, the problem that search result information sorting is difficult to meet the needs of multiple parties is solved, and personalized and efficient search result regulation is achieved.

CN114048250BActive Publication Date: 2025-07-18BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111366346.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-07-18
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

In the prior art, in Internet services, the number of search result information is large and the order is difficult to meet the needs of the search platform, search result information owner and user at the same time, resulting in poor regulation effect.

Method used

The regulation model of reinforcement learning training is adopted, the regulation information set is used as an agent, the interaction information is used as the state and regulation parameters as the action, and the regulation effect is used as the reward to build a regulation model to achieve dynamic regulation of the search result information sequence.

Benefits of technology

It realizes personalized regulation of search result information based on real-time interactive information and regulatory effects, improves the accuracy and efficiency of regulation, and avoids the loss of regulation efficiency in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114048250B_ABST
    Figure CN114048250B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an information regulation method and apparatus. A specific implementation of the method includes: in response to receiving a search request, determining a search result information sequence corresponding to the search request; obtaining the current regulation attributes of the regulation information set, where the regulation attributes include interaction information and regulation effects; according to the regulation attributes, using a regulation model pre-trained based on reinforcement learning to obtain regulation parameters, where the regulation model uses the regulation information set as an agent, the interaction information of the regulation information set as a state, the regulation parameters as actions, and the regulation effects as rewards; according to the regulation parameters, updating the positions of the regulation information in the regulation information set in the search result information sequence. This implementation helps to achieve stable regulation of search result information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to information regulation methods and apparatuses. Background Art

[0002] Currently, various online services provided by the Internet (such as search, online shopping, etc.) usually first receive user requirements (such as search keywords entered by the user), and then return corresponding search result information to the user for selection. With the explosive growth of the amount of information on the Internet, in many cases, the number of search result information returned to the user based on the user requirements may be very large.

[0003] Generally, the sorting of search result information is determined by considering the requirements of the search platform, the owner of the search result information, and the user at the same time, so as to satisfy the requirements of the three parties as much as possible. For example, it satisfies the value stability of the search platform, the traffic requirements of the owner of the search result information, and the search requirements of the user at the same time. Summary of the Invention

[0004] Embodiments of the present disclosure propose information regulation methods and apparatuses.

[0005] In a first aspect, embodiments of the present disclosure provide an information regulation method, the method including: in response to receiving a search request, determining a search result information sequence corresponding to the search request; obtaining a current regulation attribute of a regulation information set, where the regulation attribute includes interaction information and regulation effect; according to the regulation attribute, obtaining a regulation parameter by using a regulation model pre-trained based on reinforcement learning, where the regulation model takes the regulation information set as an agent, the interaction information of the regulation information set as a state, the regulation parameter as an action, and the regulation effect as a reward; and according to the regulation parameter, updating the position of the regulation information in the regulation information set in the search result information sequence.

[0006] In a second aspect, embodiments of the present disclosure provide an information regulation apparatus, the apparatus including: a search unit configured to, in response to receiving a search request, determine a search result information sequence corresponding to the search request; an obtaining unit configured to obtain a current regulation attribute of a regulation information set, where the regulation attribute includes interaction information and regulation effect; a determining unit configured to obtain a regulation parameter by using a regulation model pre-trained based on reinforcement learning according to the regulation attribute, where the regulation model takes the regulation information set as an agent, the interaction information of the regulation information set as a state, the regulation parameter as an action, and the regulation effect as a reward; and a regulation unit configured to update the position of the regulation information in the regulation information set in the search result information sequence according to the regulation parameter.

[0007] In a third aspect, embodiments of the present disclosure provide an electronic device, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.

[0008] In a fourth aspect, embodiments of the present disclosure provide a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0009] The information regulation method and device provided by the embodiments of the present disclosure first determine a corresponding search result information sequence according to a search request, then use a pre-trained regulation model to determine corresponding regulation parameters according to the current interaction information and regulation effect of a regulation information set, and then use the regulation parameters to update the positions of the regulation information in the regulation information set in the search result information sequence. Specifically, the regulation model is constructed with the regulation information set as the agent, the interaction information of the regulation information set as the state, the regulation parameters as the actions, and the regulation effect as the reward, so as to transform the regulation problem of the regulation information into a sequential decision-making problem through abstract modeling, and perform real-time dynamic regulation on the regulation information according to the regulation parameters determined by the real-time interaction information and regulation effect of the regulation information set. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Other features, objects, and advantages of the present disclosure will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0011] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;

[0012] Figure 2 is a flowchart of an embodiment of the information regulation method according to the present disclosure;

[0013] Figure 3 is a flowchart of another embodiment of the information regulation method according to the present disclosure;

[0014] Figure 4 is a flowchart of still another embodiment of the information regulation method according to the present disclosure;

[0015] Figure 5 is a schematic diagram of an application scenario of the information regulation method according to the embodiments of the present disclosure;

[0016] Figure 6 is a schematic structural diagram of an embodiment of the information regulation device according to the present disclosure;

[0017] Figure 7It is a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners

[0018] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant invention and do not limit the invention. Additionally, it should be noted that for the convenience of description, only parts related to the relevant invention are shown in the drawings.

[0019] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and embodiments.

[0020] Figure 1 An exemplary architecture 100 is shown, which can apply the embodiments of the information regulation method or information regulation device of the present disclosure.

[0021] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0022] The terminal devices 101, 102, 103 interact with the server 105 through the network 104 to receive or send messages, etc. Various client applications may be installed on the terminal devices 101, 102, 103. For example, browser applications, search applications, instant messaging tools, social platforms, shopping applications, information flow applications, etc.

[0023] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices, including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (for example, multiple software or software modules for providing distributed services), or it may be implemented as a single software or software module. No specific limitation is made here.

[0024] Server 105 may be a server that provides various services, such as a backend server that provides service support for client applications installed on terminal devices 101, 102, and 103. In response to receiving search requests sent by terminal devices 101, 102, and 103, server 105 may determine a search result information sequence corresponding to the search requests, obtain a regulation parameter by using a regulation model according to the current regulation attribute of the regulation information set, then update the position of the regulation information in the regulation information set in the search result information sequence according to the regulation parameter, and return the updated search result information sequence as the final search result to terminal devices 101, 102, and 103.

[0025] It should be noted that the information regulation method provided by the embodiments of the present disclosure is generally executed by server 105. Correspondingly, the information regulation device is generally disposed in server 105.

[0026] It should also be pointed out that search applications may also be installed in terminal devices 101, 102, and 103. Terminal devices 101, 102, and 103 may also determine a search result information sequence corresponding to a search request based on the search application, obtain a regulation parameter by using a regulation model according to the current regulation attribute of the regulation information set, then update the position of the regulation information in the regulation information set in the search result information sequence according to the regulation parameter, and use the updated search result information sequence as the final search result. At this time, the information regulation method may also be executed by terminal devices 101, 102, and 103. Correspondingly, the information regulation device may also be disposed in terminal devices 101, 102, and 103. At this time, the exemplary system architecture 100 may not include server 105 and network 104.

[0027] It should be noted that server 105 may be hardware or software. When server 105 is hardware, it may be implemented as a distributed server cluster composed of multiple servers, or may be implemented as a single server. When server 105 is software, it may be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.

[0028] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0029] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 2 Continue to refer to

[0030] Step 201: In response to receiving a search request, determine a search result information sequence corresponding to the search request.

[0031] In this embodiment, the search request can be any request for requesting the return of search result information. The search result information can be various types of information, such as images, texts, audios, videos, etc.

[0032] Generally, the search result information can be determined by matching the search request with the information in a preset information set, and the information that matches the search request in the information set is used as the search result information. Among them, the information set can be determined according to the actual application scenario. The search result information sequence can be formed by arranging several pieces of information that match the search request in a certain order (such as the matching degree with the search request, generation time, etc.).

[0033] As an example, the search request includes search keywords. At this time, items with a correlation degree greater than a preset threshold with the search keywords can be determined from the item set according to the search keywords, and the display information corresponding to the determined items is obtained. Then, a display information sequence is formed in descending order of the correlation degree corresponding to the determined items as the search result information sequence of the search request.

[0034] The execution entity of the method for processing user requests (such as Figure 1 the server 105 shown, etc.) can receive the search request directly input by the user, or can receive the search requests sent by other terminal devices (such as Figure 1 the terminal devices 101, 102, 103 shown, etc.).

[0035] Step 202: Obtain the current regulation attribute of the regulation information set.

[0036] In this embodiment, the regulation information set can be composed of several pieces of regulation information. The regulation information can be any information to be regulated. The regulation information set can be specifically specified by technicians in advance according to actual application requirements. For example, the regulation information set can be several pieces of information to be regulated under the same category.

[0037] Generally, the regulation target for information can be set according to the actual application scenario or application requirements. For example, the regulation target can be that the traffic of the information meets preset conditions. Among them, the traffic can be any information interaction attribute, such as exposure volume, click-through rate, conversion rate, etc.

[0038] Among them, the regulation target can usually be achieved by regulating the position of information in the information sequence of search results. Since the information sequence of search results presented to users may contain a large amount of information, the interaction attributes corresponding to information at different positions in the information sequence of search results may vary. For example, the exposure of information at the front position in the information sequence of search results is generally greater than that of information at the rear position.

[0039] The regulation attributes of the regulation information set can include interaction information and regulation effects. The above-mentioned execution entity can obtain the current regulation attributes of the regulation information set from local or other storage devices. Among them, the interaction information can refer to various interaction information related to the regulation information set. For example, the interaction information can include the overall click-through rate, overall conversion rate, etc. of the regulation information set. The regulation effect can be used to describe the current regulation result of the regulation information set, and can be specifically represented by various representation methods. For example, the regulation effect can include the current regulation target completion information (such as the regulation target completion rate) of the regulation information set, regulation distribution information (such as the regulation target completion rate of each time period, etc.).

[0040] As an example, in the scenario of traffic regulation for information, the regulation effect can include the overall traffic completion rate, which can be specifically determined by the quotient of the currently actually completed traffic and the currently expected completed traffic. The regulation effect can also include the traffic distribution situation of each time period (such as each day), and the traffic distribution situation of each time period can be specifically determined by the quotient of the traffic actually completed in that time period and the traffic expected to be completed in that time period.

[0041] Step 203: Obtain regulation parameters according to the regulation attributes by using a regulation model pre-trained based on reinforcement learning.

[0042] In this embodiment, reinforcement learning (RL), also known as re-inforcement learning, evaluation learning or enhanced learning, is one of the paradigms and methodologies of machine learning. Reinforcement learning is generally used to describe and solve the problem that an agent (Agent) learns a strategy to maximize the reward or achieve a specific goal during the interaction with the environment.

[0043] The regulation model is constructed based on the idea of using the regulation information set as the agent, the interaction information of the regulation information set as the state (State), the regulation parameters as the action (Action), and the regulation effect as the reward (Reward). Based on this, the information regulation problem is transformed into a sequential decision-making problem that can be solved by using reinforcement learning. The regulation model can adopt various existing regulation model structures, such as DQN (Deep Q-learning), DRQN (Deep Recurrent Q-Learning Network), etc.

[0044] The policy of the regulation model can be set by technicians according to the actual application scenario.

[0045] As an example, the policy of the regulation model can be defined as follows:

[0046] E[(r + γmaxQ(s′, a′, w) - Q(s, a, w)) 2

[0047] Where s represents the current state, a represents the current action, s′ represents the next state, a′ represents the next action, γ represents the damping coefficient, r represents the current reward. E() represents taking the expectation. The damping coefficient can be preset by technicians. Generally, the value range of the damping coefficient can be between 0 and 1. w represents the parameter of the regulation model.

[0048] Step 204: Update the position of the regulation information in the search result information sequence according to the regulation parameter.

[0049] In this embodiment, after obtaining the regulation parameter, the positions of the information in the search result information sequence, that is, the arrangement order of the information, can be adjusted to obtain the updated search result information sequence.

[0050] Specifically, for the information in the regulation information set, if the search result information sequence includes this information, the current position of this information in the search result information sequence can be obtained, and then according to the regulation parameter, various methods can be used to adjust the current position of this information in the search result information sequence. For example, the sum of the current position of this information in the search result information sequence and the regulation parameter can be determined as the adjusted position of this information in the search result information sequence.

[0051] In some optional implementation manners of this embodiment, after obtaining the regulation attribute of the regulation information set, the regulation parameter of the regulation information set can be determined from the preset regulation parameter set by using the regulation model.

[0052] Where the regulation parameter set can be determined according to the target regulation parameter and the regulation frequency. The regulation frequency can refer to the number of regulations per unit time. Both the target regulation parameter and the regulation frequency can be flexibly set by technicians according to the actual application scenario or application requirements. Specifically, the regulation parameter set can include the target regulation parameter. At the same time, the regulation parameter set can also include the new regulation parameter obtained by updating the target regulation parameter based on the regulation frequency. The regulation parameter set can also include the regulation parameter specified in advance by technicians to control the value range of the regulation parameters in the regulation parameter set.

[0053] ​For example, the correspondence between the regulation frequency and the step size of the regulation parameter can be preset by a technician, so that the step size of the regulation parameter corresponding to the regulation frequency can be obtained. Then, the target regulation parameter is updated with the determined step size of the regulation parameter to obtain a set of regulation parameters. As an example, if the target regulation parameter is 1 and the step size is 0.05, the set of regulation parameters can be (0.9, 0.95, 1.0, 1.05, 1.10).

[0054] By presetting a set of regulation parameters such as the regulation frequency, the regulation parameters are discretized and controlled within a finite discrete space, and then the regulation result can be controlled to a certain extent to ensure the rationality of the regulation effect.

[0055] In some alternative implementation manners of this embodiment, the reward function of the regulation model can be determined according to the difference between the target regulation effect and the actual regulation effect of the regulation information set.

[0056] As an example, the reward function can be designed as follows:

[0057] C - a×ReLU(C - bC0)

[0058] Where C represents the currently accumulated completed regulation target, C0 represents the expected completed regulation target up to the current time. a and b are control parameters, which can be specifically preset by a technician. Generally, the values of a and b can be greater than 1. RELU represents the rectified linear unit function.

[0059] Using the difference between the target regulation effect and the actual regulation effect of the regulation information set for reward and punishment control can further ensure the rationality of the regulation effect, make the regulation effect as consistent with the regulation target as possible, and avoid the situation where the regulation effect is too different from the regulation target.

[0060] In the prior art, usually a quality score is set for each search result information according to the needs of three parties, and then the quality score is used as the basis for the display sorting of each search result information. The method provided in the above embodiment of the present disclosure constructs a regulation model by using the regulation information set as an agent, the interaction information of the regulation information set as the state, the regulation parameter as the action, and the regulation effect as the reward, so as to transform the regulation problem of the regulation information into a sequential decision-making problem through abstract modeling. The regulation model dynamically determines the current regulation parameter in real time according to the real-time interaction information and the regulation effect of the regulation information set, and then updates the position of the regulation information in the search result information sequence corresponding to the user's search request according to the determined regulation parameter, thereby realizing the real-time dynamic regulation of the regulation information.

[0061] Further referring to Figure 3 , which shows the flow 300 of another embodiment of the information regulation method. The flow 300 of the information regulation method includes the following steps:

[0062] Step 301: In response to receiving a search request, determine a search result information sequence corresponding to the search request.

[0063] Step 302: Obtain the current regulation attribute of the regulation information set and the current regulation attribute of each regulation information in the regulation information set.

[0064] In this embodiment, for each regulation information in the regulation information set, the regulation attribute of the regulation information may include at least one of the following: interaction information and value information. Among them, the interaction information may refer to various information related to the interaction between the regulation information and the user. For example, the interaction information may include, but is not limited to, click-through rate, exposure volume, conversion rate, etc. The value information may be used to indicate the value of the regulation information. For example, the value information may include, but is not limited to, price, cost, etc. The current regulation attribute of each regulation information may refer to the current real-time interaction information and / or value information of the regulation information.

[0065] Step 303: According to the regulation attribute, use a regulation model pre-trained based on reinforcement learning to obtain regulation parameters.

[0066] Step 304: For each regulation information in the regulation information set, update the position of the regulation information in the search result information sequence according to the regulation parameters and the current regulation attribute of the regulation information.

[0067] In this embodiment, for each regulation information, the position of the regulation information in the search result information sequence may be updated by various methods according to the regulation parameters corresponding to the regulation information set to which it belongs and the current regulation attribute of the regulation information.

[0068] For example, according to actual application requirements, the relationship between the interaction information and / or value information of the regulation information and the regulation target may be determined, and then the position of the regulation information in the search result information sequence may be updated according to this relationship.

[0069] As an example, for each regulation information, the product of the regulation parameters corresponding to the regulation information set to which it belongs and the current regulation attribute of the regulation information may be calculated first as the regulation parameter corresponding to the regulation information, and then the sum of the regulation parameter corresponding to the regulation information and the position of the regulation information in the search result information sequence may be calculated as the position of the regulation information in the updated search result information sequence.

[0070] For the content not specifically described in this embodiment, reference may be made to Figure 2 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.

[0071] Based on the method provided by the above embodiments of the present disclosure, which determines the regulation parameters according to the real-time interaction information and regulation effects of the regulation information set by using the regulation model, the position of the regulation information in the search result information is comprehensively regulated by combining the real-time interaction information and / or real-time value information of each information in the regulation information set to give personalized information, so as to realize personalized regulation of different regulation information, refine the granularity of information regulation from the global level to individual information, and use the offline-trained regulation model to obtain the regulation parameters and then combine real-time data for information regulation to avoid loss of regulation efficiency.

[0072] Further referring to Figure 4 , which shows the flow 400 of another embodiment of the information regulation method. The flow 400 of the information regulation method includes the following steps:

[0073] Step 401, in response to receiving a search request, determine a search result information sequence corresponding to the search request.

[0074] Step 402, obtain the current regulation attributes of each regulation information set in at least two regulation information sets.

[0075] In this embodiment, at least two regulation information sets can be preset according to actual application scenarios or application requirements. For example, different regulation information sets can respectively belong to different owners (such as different merchants or brands, etc.). The regulation objectives of different regulation information sets can be the same or different, and can be specifically set according to the actual requirements of each regulation information.

[0076] Step 403, according to the regulation attributes of each regulation information set, use the regulation model to obtain the regulation parameters of each regulation information set.

[0077] In this embodiment, the regulation model can be used to obtain the regulation parameters corresponding to each regulation information set.

[0078] Step 404, for each regulation information set, update the position of the regulation information in the regulation information set in the search result information sequence according to the regulation parameters of the regulation information set.

[0079] In this embodiment, for each regulation information set, the position of the regulation information in the regulation information set in the search result information sequence can be determined according to the regulation parameters corresponding to the regulation information set. Thus, the updated position of each regulation information in each regulation information set in the search result information sequence can be obtained, thereby forming a new search result information sequence.

[0080] In some alternative implementation manners of this embodiment, a balance parameter may also be obtained, and then for each set of regulation information, the position of the regulation information in the search result information sequence in this set of regulation information may be updated according to the balance parameter and the regulation parameter corresponding to this set of regulation information.

[0081] Among them, the balance parameter may be determined according to the differences between each set of regulation information for balancing the differences between each set of regulation information. Among them, the differences between each set of regulation information may refer to the differences in various attributes. For example, the differences include but are not limited to: differences in the brands to which they belong, differences in value, and so on. The balance parameter may specifically be preset by technicians or may be continuously updated and adjusted according to the actual regulation effect.

[0082] Specifically, for each set of regulation information, various methods may be adopted to update the position of the regulation information in the search result information sequence in this set of regulation information according to the balance parameter and the regulation parameter corresponding to this set of regulation information.

[0083] For example, the product of the balance parameter and the regulation parameter corresponding to this set of regulation information may be determined first, and then for each piece of regulation information in this set of regulation information, the sum of the obtained product and the position of this piece of regulation information in the search result information sequence may be used as the updated position of this piece of regulation information in the search result information sequence.

[0084] Since there are large differences in various attributes of various different sets of regulation information, but it is necessary to regulate the regulation information in each set of regulation information, therefore, by setting a balance parameter to balance the dimensional differences between different sets of regulation information, it is possible to more reasonably regulate the regulation information in different sets of regulation information as a whole and avoid the volatility between various information regulations.

[0085] In some alternative implementation manners of this embodiment, a value parameter may be obtained, and then for each set of regulation information, the position of the regulation information in the search result information sequence in this set of regulation information may be updated according to the value parameter and the regulation parameter corresponding to this set of regulation information.

[0086] Among them, the value parameter may be determined according to the influence of the regulation effect of the set of regulation information on the target value. Among them, the target value may be flexibly determined according to the actual application scenario. For example, the target value may include but is not limited to the value of the brand to which the set of regulation information belongs, the value of the third-party platform that displays the information in the set of regulation information, and so on.

[0087] The impact of the regulation effect on the target value can refer to the increase or loss of the target value. Different regulation effects will affect the mutual information of each regulation information, and thus affect the target value. Based on this, by setting value parameters and adjusting the value parameters in a timely manner according to the regulation effect, the impact of each regulation on the target value can be controlled, avoiding situations such as excessive fluctuations in the target value caused by the regulation of the regulation information concentrated in the regulation information, so as to ensure the stability of the target value.

[0088] After obtaining the regulation parameters of the regulation information set, various factors can be considered according to the actual application requirements, and the information in the regulation information set can be comprehensively updated in the position of the search result information sequence by using the corresponding various parameters.

[0089] As an example, the position update of the regulation information in the regulation information set can be designed as follows:

[0090] α0×α×pctr×pcvr×price×log(rk0)×β + rk0

[0091] Among them, rk0 represents the initial position of the regulation information in the search result information sequence. α represents the regulation parameter. α0 represents the balance parameter. β represents the value parameter. pctr represents the real-time click-through rate of the regulation information, pcvr represents the real-time conversion rate of the regulation information, and price represents the real-time value of the regulation information. log() represents the natural logarithm. At this time, the natural logarithm of the initial position of the regulation information, the real-time click-through rate, the real-time conversion rate, the real-time value, and the product of the regulation parameter, the balance parameter, and the value parameter can be calculated first, and then the sum of the obtained product and the initial position is used as the updated position.

[0092] Continue to refer to Figure 5 , Figure 5 is a schematic application scenario 500 of the information regulation method according to this embodiment. In Figure 5 In the application scenario, the server can receive the search keywords input by the user, and then determine the search result information sequence 501 according to the user's search keywords. As shown in the figure, the search result information sequence 501 successively includes information "A", information "B", information "C", and so on. Then the service can determine the updated positions of the respective regulation information in the search result information sequence 501.

[0093] Taking the information "A" 5011 as an example, the server can first obtain the real-time regulation attributes 502 of the information set "A0" to which the information "A" 5011 belongs (such as click-through rate, conversion rate, regulation effect, etc.), and then use the pre-trained regulation model to obtain the regulation parameters 504 corresponding to the information set "A0" based on the regulation attributes 502. Then, the server can obtain the preset value parameter 505, the preset balance parameter 506, the current position 507 of the information "A" 5011 in the search result information sequence 501, and the real-time regulation attributes 508 of the information "A" (such as click-through rate, conversion rate, value, etc.). After that, the service can determine the updated position 509 corresponding to the information "A" based on the regulation parameters 504, the value parameter 505, the balance parameter 506, the current position 507, and the regulation attributes 508. Then, the server can update the search result information sequence according to the updated positions corresponding to the respective regulation information in the search result information sequence 501, form a new search result information sequence 510 and use it as the final search result information sequence, and then display it to the user.

[0094] For the content not specifically described in this embodiment, reference can be made to Figure 2 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.

[0095] The method provided in the above embodiments of the present disclosure uses a regulation model to determine the corresponding regulation parameters according to the real-time interaction information and regulation attributes such as regulation effects of multiple different regulation information sets respectively, and then updates the positions of the regulation information in each regulation information set in the search result information sequence according to the respectively corresponding regulation parameters, so as to achieve the comprehensive regulation of multiple different regulation information sets. Moreover, through balance parameters, value parameters, etc., the natural differences between each regulation information set and the change range of the target value can be reasonably controlled, which helps to achieve the accuracy of the regulation results of each regulation information set.

[0096] Further referring to Figure 6 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an information regulation device. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0097] Such as Figure 6As shown in the figure, the information regulation device 600 provided in this embodiment includes a search unit 601, an acquisition unit 602, a determination unit 603, and a regulation unit 604. Among them, the search unit 601 is configured to determine a search result information sequence corresponding to a search request in response to receiving the search request; the acquisition unit 602 is configured to acquire the current regulation attribute of the regulation information set, where the regulation attribute includes interaction information and regulation effect; the determination unit 603 is configured to obtain a regulation parameter by using a regulation model pre-trained based on reinforcement learning according to the regulation attribute, where the regulation model takes the regulation information set as an agent, the interaction information of the regulation information set as a state, the regulation parameter as an action, and the regulation effect as a reward; the regulation unit 604 is configured to update the position of the regulation information in the regulation information set in the search result information sequence according to the regulation parameter.

[0098] In this embodiment, in the information regulation device 600: the specific processing of the search unit 601, the acquisition unit 602, the determination unit 603, and the regulation unit 604 and the technical effects brought by them can respectively refer to Figure 2 The relevant descriptions of steps 201, step 202, step 203, and step 204 in the corresponding embodiments will not be elaborated here.

[0099] In some optional implementation manners of this embodiment, the above method further includes: acquiring the current regulation attribute of each regulation information in the regulation information set, where the regulation attribute of each regulation information includes at least one of the following: interaction information, value information; and the above updating the position of the regulation information in the regulation information set in the search result information sequence according to the regulation parameter includes: for each regulation information in the regulation information set, updating the position of the information in the search result information sequence according to the regulation parameter and the current regulation attribute of the regulation information.

[0100] In some optional implementation manners of this embodiment, the above obtaining the regulation parameter by using a regulation model pre-trained based on reinforcement learning according to the regulation attribute includes: determining the regulation parameter of the regulation information set from a preset regulation parameter set according to the regulation attribute, where the regulation parameter set is determined according to the target regulation parameter and the regulation frequency.

[0101] In some optional implementation manners of this embodiment, the reward function of the above regulation model is determined according to the difference between the target regulation effect and the actual regulation effect of the regulation information set.

[0102] In some optional implementation manners of this embodiment, the number of regulation information sets is at least two; and the above obtaining the regulation parameter by using a regulation model pre-trained based on reinforcement learning according to the regulation attribute includes: obtaining the regulation parameter of each regulation information set by using the regulation model according to the regulation attribute of each regulation information set.

[0103] In some alternative implementation manners of this embodiment, the above method further includes: obtaining a balance parameter, where the balance parameter is determined according to the differences between the respective regulation information sets; and updating the positions of the regulation information in the regulation information sets in the search result information sequence according to the regulation parameters, including: for each regulation information set, updating the positions of the regulation information in the regulation information set in the search result information sequence according to the balance parameter and the regulation parameters of this regulation information set.

[0104] In some alternative implementation manners of this embodiment, the above method further includes: obtaining a value parameter, where the value parameter is determined according to the influence of the regulation effect of the regulation information set on the target value; and updating the positions of the regulation information in the regulation information sets in the search result information sequence according to the regulation parameters, including: updating the positions of the regulation information in the regulation information set in the search result information sequence according to the value parameter and the regulation parameters.

[0105] The apparatus provided in the above embodiment of the present disclosure, through the search unit, in response to receiving a search request, determines a search result information sequence corresponding to the search request; the obtaining unit obtains the current regulation attributes of the regulation information set, where the regulation attributes include interaction information and regulation effect; the determining unit, according to the regulation attributes, uses a regulation model pre-trained based on reinforcement learning to obtain regulation parameters, where the regulation model uses the regulation information set as the agent, the interaction information of the regulation information set as the state, the regulation parameters as the action, and the regulation effect as the reward; the regulation unit updates the positions of the regulation information in the regulation information set in the search result information sequence according to the regulation parameters, thereby transforming the regulation problem of the regulation information into a sequential decision-making problem through abstract modeling, and dynamically regulating the regulation information in real time according to the regulation parameters determined by the real-time interaction information and regulation effect of the regulation information set.

[0106] Reference is made below to Figure 7 , which shows a schematic structural diagram of an electronic device (such as Figure 1 the server in Figure 7 ) 700 suitable for implementing the embodiments of the present disclosure.

[0107] As Figure 7As shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0108] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 7 an electronic device 700 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 7 Each block shown in may represent a device or, as needed, multiple devices.

[0109] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the method of the embodiment of the present disclosure are executed.

[0110] It should be noted that the computer-readable medium described in the embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0111] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: in response to receiving a search request, determine a search result information sequence corresponding to the search request; obtain the current regulation attribute of the regulation information set, where the regulation attribute includes interaction information and regulation effect; according to the regulation attribute, use a regulation model pre-trained based on reinforcement learning to obtain regulation parameters, where the regulation model takes the regulation information set as an agent, the interaction information of the regulation information set as a state, the regulation parameters as actions, and the regulation effect as a reward; according to the regulation parameters, update the position of the regulation information in the regulation information set in the search result information sequence.

[0112] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0114] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. The described units may also be provided in a processor. For example, it may be described as: a processor includes a search unit, an acquisition unit, a determination unit, and a regulation unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the search unit may also be described as "the unit that determines the search result information sequence corresponding to the search request in response to receiving the search request".

[0115] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.

Claims

1. An information regulation method, comprising: Responding to receiving a search request, determining a search result information sequence corresponding to the search request; Obtaining the current regulation attributes of the regulation information set, wherein the regulation attributes include interaction information and regulation effect; According to the regulation attributes, using a regulation model pre-trained based on reinforcement learning to determine a regulation parameter from a preset set of regulation parameters, wherein the regulation model uses the regulation information set for the reinforcement learning of an intelligent agent, the interaction information of the regulation information set as the state, the regulation parameter as the action, and the regulation effect as the reward, and the set of regulation parameters is determined according to the target regulation parameter and the regulation frequency; According to the regulation parameter and the current position of the regulation information in the search result information sequence in the regulation information set, determining the adjusted position of the regulation information in the search result information sequence.

2. The method according to claim 1, wherein, The method further comprises: Obtaining the current regulation attributes of each regulation information in the regulation information set, wherein the regulation attributes of each regulation information include at least one of the following: interaction information, value information; and The updating the position of the regulation information in the search result information sequence according to the regulation parameter includes: For each regulation information in the regulation information set, updating the position of the information in the search result information sequence according to the regulation parameter and the current regulation attributes of the information.

3. The method according to claim 1, wherein The reward function of the regulation model is determined according to the difference between the target regulation effect and the actual regulation effect of the regulation information set.

4. The method according to any one of claims 1 to 3, wherein The number of the regulation information sets is at least two; And The obtaining the regulation parameter by using the regulation model pre-trained based on reinforcement learning according to the regulation attributes includes: obtaining the regulation parameter of each regulation information set by using the regulation model according to the regulation attributes of each regulation information set.

5. The method according to claim 4, wherein, The method further comprises: Obtaining a balance parameter, wherein the balance parameter is determined according to the differences between the regulation information sets; and The updating the position of the regulation information in the search result information sequence according to the regulation parameter includes: for each regulation information set, updating the position of the regulation information in the search result information sequence according to the balance parameter and the regulation parameter of the regulation information set.

6. The method according to claim 5, wherein The method further comprises: Obtaining a value parameter, wherein the value parameter is determined according to the influence of the regulation effect of the regulation information set on the target value; and The updating the position of the regulation information in the search result information sequence according to the regulation parameter includes: Updating the position of the regulation information in the search result information sequence according to the value parameter and the regulation parameter.

7. An information regulation device, wherein, The device comprises: A search unit configured to respond to receiving a search request and determine a search result information sequence corresponding to the search request; An obtaining unit configured to obtain the current regulation attributes of the regulation information set, wherein the regulation attributes include interaction information and regulation effect; A determination unit, configured to determine a regulation parameter from a preset set of regulation parameters according to the regulation attribute by using a regulation model pre-trained based on reinforcement learning, where the regulation model uses an interaction information set for the reinforcement learning of an intelligent agent, the interaction information of the interaction information set is used as a state, the regulation parameter is used as an action, and the regulation effect is used as a reward, and the set of regulation parameters is determined according to a target regulation parameter and a regulation frequency; A regulation unit, configured to determine an adjusted position of the regulation information in the search result information sequence according to the regulation parameter and the current position of the regulation information in the search result information sequence in the regulation information set; 8. The apparatus according to claim 7, wherein The obtaining unit is further configured to: Obtain the current regulation attribute of each regulation information in the regulation information set, where the regulation attribute of each regulation information includes at least one of the following: interaction information, value information; and The regulation unit is further configured to: for each regulation information in the regulation information set, update the position of the information in the search result information sequence according to the regulation parameter and the current regulation attribute of the regulation information; 9. The apparatus according to claim 7, wherein The reward function of the regulation model is determined according to the difference between the target regulation effect and the actual regulation effect of the regulation information set; 10. The device according to any one of claims 7-9, wherein, The number of the regulation information sets is at least two; and The determination unit is further configured to: obtain the regulation parameter of each regulation information set by using the regulation model according to the regulation attribute of each regulation information set; 11. The apparatus according to claim 10, wherein The obtaining unit is further configured to: Obtain a balance parameter, where the balance parameter is determined according to the difference between the regulation information sets; and The regulation unit is further configured to: for each regulation information set, update the position of the regulation information in the regulation information set in the search result information sequence according to the balance parameter and the regulation parameter of the regulation information set; 12. The device according to claim 11, wherein, The obtaining unit is further configured to: Obtain a value parameter, where the value parameter is determined according to the influence of the regulation effect of the regulation information set on a target value; and The regulation unit is further configured to: update the position of the regulation information in the regulation information set in the search result information sequence according to the value parameter and the regulation parameter; 13. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6; 14. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Information search method and device

    CN112579897A

  • Model training method and device and information display method and device

    CN113343130A