Joint broadcast method, system, device and medium
Through the joint broadcasting method, Markov decision-making and reinforcement learning model are used to adjust the user's broadcasting strategy based on the user's weight of interest and broadcasting time, which solves the problem of the platform's resource in the push of multiple works and maximizes the profits.
Patent Information
- Application Number
- CN202510570362.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-06
AI Technical Summary
It is difficult for the platform to balance user selection in the push and streaming of multiple works, resulting in resource tilt and affecting the platform and author's maximization of benefits.
The joint broadcasting method is adopted, and the Markov decision-making and reinforcement learning model is used to adjust the user's broadcasting strategy based on the user's interest weight and broadcasting duration to maximize the total incentive.
The revenue of the platform and authors has been improved, and by dynamically adjusting the broadcasting strategy, the total incentive is maximized in scenarios where the user has the right to choose.
Smart Images

Figure CN120091063B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of audio recommendation, and in particular, to a combined broadcast method, system, device, and medium. Background Art
[0002] In order to maintain the activity of content creators (i.e., authors) and users (i.e., viewers or listeners) on platforms (such as Bilibili, QQ Reading, etc.), the platforms usually need to push multiple works simultaneously. This can not only enhance the user's stickiness to the platform but also give continuous creative motivation to content creators.
[0003] In the push operation of multiple works by the platform, the decision-making power often lies in the hands of users. That is, users can decide whether to switch the currently broadcast work according to various factors (such as environment, incentives, etc.). The platform cannot influence the user's choice, but the platform needs to balance the push progress of each work to avoid excessive resource tilt. Therefore, how the platform controls and adjusts the works broadcast by users to maximize the benefits of pushing multiple works on the platform is an urgent problem to be solved currently. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this article. This overview is not intended to limit the scope of protection of the claims.
[0005] The main purpose of the embodiments of the present application is to propose a combined broadcast method, system, device, and medium, which can adjust the combined broadcast strategy of multiple audios and maximize the benefits of pushing multiple works on the platform.
[0006] To achieve the above object, a first aspect of the embodiments of the present application provides a combined broadcast method, which is applied to a combined broadcast network. The combined broadcast network includes a platform and multiple user terminals. The platform is used to control multiple user terminals to execute a combined broadcast strategy for multiple audios. The combined broadcast strategy includes: enabling each user terminal to broadcast one of the multiple audios, and any one of the audios can be broadcast on one or more user terminals. The method includes:
[0007] According to the combined broadcast strategy at the current moment, count the push incentives of each of the multiple audios; wherein, the push incentive of the target audio is related to the interest weight of the target user terminal in the target audio. The target audio is any one of the multiple audios, and the target user terminal is the user terminal that broadcasts the target audio among the multiple user terminals. The interest weight refers to the proportion of the interest value of the target user terminal in the target audio in the sum of the interest values of the target user terminal in the multiple audios, and the interest value of the user terminal in the audio changes with time;
[0008] Use the push stream incentives of the multiple audios counted respectively as the state factors of the Markov decision at the current moment, and build the action factors of the Markov decision. Among them, the action factors include influence factors that can enable the client to switch or stop the played audio;
[0009] Construct a reinforcement learning model, and input the state factors at the current moment into the reinforcement learning model to obtain the action factors at the current moment. Based on the action factors at the current moment and on the basis of preset conditions, obtain the state factors at the next moment, and so on until the action factors at each moment are obtained. Based on the action factors at each moment, adjust the joint broadcast strategy of the platform for the multiple clients; the preset condition is to maximize the total incentives of the multiple audios, where the total incentives are related to the push stream incentives of the multiple audios respectively.
[0010] A joint broadcast method provided by this application has at least the following beneficial effects:
[0011] This method sets the push stream incentives for the audios of the platform and the author. Since the user's interest weight in the audio can affect the push stream incentive for the author, the interest weight is determined by the interest value, and the interest value can measure the user's interest degree in the audio at different times. Then, in the case of a greater degree of interest, more push stream incentives will be given to the author. Therefore, the set of push stream incentives corresponding to each audio is used as the state factor of the Markov decision, and the push stream effect of the audio is characterized by the state factor; then, an influence factor for playing the audio to the user is set. Since this factor can influence the user's choice of the played audio, the influence factor is used as the action factor of the Markov decision to predict the state factors at different times. Finally, a total incentive related to the push stream incentive is set as the reward (Markov decision). Based on the above-set parameters, this method can use the reinforcement learning model to learn and use different action factors to adjust the state factors under the condition of maximizing the total incentives of the multiple audios, so as to realize the adjustment of the platform's selection of audio by the user under the scenario where only the user has the right to choose when pushing the multiple audios, and improve the maximum benefits of the platform and the author.
[0012] In some embodiments, before counting the push stream incentives of the multiple audios respectively according to the joint broadcast strategy at the current moment, the following steps are further included:
[0013] Obtain the interest values of the multiple clients for the multiple audios at historical moments;
[0014] Use the interest value at the historical moment as time-series data, and input the time-series data into a time-series prediction model to predict the interest values of the multiple clients for the multiple audios at the current moment based on the time-series prediction model;
[0015] Calculate the interest weight at the current moment according to the interest value at the current moment.
[0016] In some embodiments, the time-series prediction model is an autoregressive integrated moving average model.
[0017] In some embodiments, statistically calculating the streaming incentive of the target audio includes:
[0018] Determine the interest weight of the target client for the target audio at the current moment;
[0019] Determine the broadcast duration of the target audio on the target client;
[0020] Obtain the streaming incentive of the target audio according to the product of the interest weight and the broadcast duration.
[0021] In some embodiments, before building the action factor of the Markov decision, the method further includes:
[0022] Determine the cost unit price for any client to broadcast any audio, and the streaming unit price of the platform to push any audio to any client;
[0023] Multiply the broadcast duration of any audio on any client by the streaming unit price to obtain a positive incentive;
[0024] Multiply the broadcast duration of any audio on any client by the cost unit price to obtain a negative incentive;
[0025] Subtract the negative incentive from the positive incentive to obtain the user incentive for any client;
[0026] Use the user incentive as the influence factor.
[0027] In some embodiments, the total incentive is expressed by the following function C:
[0028]
[0029] where r1 and r2 are weight values, V j is the streaming incentive corresponding to the jth audio, is the streaming incentive of the jth audio under the adjustment of the action factor a t and P jis the expected push stream incentive for the j-th audio. k(x, j) is the strategy selected by the x-th client in the joint broadcast strategy, and the value of k(x, j) is 0 or 1. min is the function to take the minimum value; n is the number of multiple audios, m is the number of multiple clients, and s t is the state factor at the current moment t, and U i,j is the user incentive for the i-th client to broadcast the j-th audio.
[0030] In some embodiments, the reinforcement learning model is a deep Q network.
[0031] To achieve the above object, a second aspect of the embodiments of the present application provides a joint broadcast system, which is applied to a joint broadcast network. The joint broadcast network includes a platform and multiple clients. The platform is used to control the multiple clients to execute the joint broadcast strategy of multiple audios. The joint broadcast strategy includes: making each of the clients broadcast one of the multiple audios, and any one of the audios can be broadcast on one or more of the clients. The system includes:
[0032] A data statistics module, configured to count the push stream incentives of the multiple audios respectively according to the joint broadcast strategy at the current moment; wherein, the push stream incentive of the target audio is related to the interest weight of the target client in the target audio; the target audio is any one of the multiple audios, and the target client is the client that broadcasts the target audio among the multiple clients; the interest weight refers to the proportion of the interest value of the target client in the target audio in the sum of the interest values of the target client in the multiple audios, and the interest value of the client in the audio changes with time;
[0033] An action factor state factor construction module, configured to use the counted push stream incentives of the multiple audios respectively as the state factors of the Markov decision at the current moment, and construct the action factors of the Markov decision, wherein the action factors include influence factors that can make the client switch or stop the broadcast of the audio;
[0034] An action factor adjustment module, configured to construct a reinforcement learning model, input the state factor at the current moment into the reinforcement learning model to obtain the action factor at the current moment, and based on the action factor at the current moment, obtain the state factor at the next moment based on a preset condition, and so on until the action factors at each moment are obtained. Based on the action factors at each moment, the platform adjusts the joint broadcast strategy of the multiple clients; the preset condition is to maximize the total incentive of the multiple audios, where the total incentive is related to the push stream incentives of the multiple audios respectively.
[0035] It can be understood that the beneficial effects of the second to fourth aspects compared with the related art are the same as those of the first aspect compared with the related art. For the relevant descriptions, reference can be made to the first aspect above, and details will not be elaborated here. Description of the Drawings
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the related art descriptions. Obviously, the drawings in the following descriptions are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 is a schematic flowchart of a joint broadcast method provided by an embodiment of the present application;
[0038] Figure 2 is a schematic diagram of a Markov decision provided by an embodiment of the present application;
[0039] Figure 3 is a schematic diagram of a deep Q-network provided by an embodiment of the present application;
[0040] Figure 4 is a schematic structural diagram of a joint broadcast system provided by an embodiment of the present application;
[0041] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0042] In order to make the purpose, technical solutions and advantages of the present application clearer, the following will further elaborate on the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] Problems existing at the current stage:
[0044] In order to maintain the activity of content creators (i.e., authors) and users (i.e., viewers or listeners) on platforms (such as Bilibili, QQ Reading, etc.), the platform usually needs to push multiple works simultaneously. This can not only enhance the user's stickiness to the platform but also give continuous creative motivation to content creators.
[0045] In the push - streaming operation of the platform for multiple works, the decision - making power often lies in the hands of users. That is, users can decide whether to switch the currently broadcast work based on various factors (such as environment, incentives, etc.), and the platform cannot influence the users' choices. However, the platform needs to balance the push - streaming progress of each work to avoid excessive resource tilt. Therefore, how the platform controls and adjusts the works broadcast by users to maximize the benefits of pushing multiple works on the platform is an urgent problem to be solved currently.
[0046] To solve the above - mentioned technical problems, such as Figure 1 , this application provides a combined broadcast method;
[0047] This method is applied to a combined broadcast network, and the combined broadcast network includes:
[0048] (1) Multiple user terminals; each user terminal can only broadcast one audio at each moment, that is, it cannot broadcast two audios simultaneously. Among them, users use the user terminal to broadcast audio.
[0049] (2) The platform, where authors are stationed. Each author will upload the above - mentioned audio works on the platform. Therefore, there are multiple audios on the platform. Here, the audio refers to the audio that needs to be push - streamed.
[0050] Multiple user terminals execute a combined broadcast strategy, and the combined broadcast strategy includes: making each user terminal broadcast one of the multiple audios, and one audio can be broadcast on one or more user terminals.
[0051] User terminals can, driven by factors such as incentives and environment, broadcast, switch, or stop the audio.
[0052] When the platform push - streams the audio, the author can obtain push - streaming incentives (the push - streaming incentives can be given to users to enhance the creative motivation). For example, incentives can be given in segments according to the broadcast duration of the audio, or in order to increase stickiness, for two users who broadcast the same audio, more incentives can be given to the user with a higher weight of interest in the audio.
[0053] Moreover, it is generally considered that when the audio is push - streamed, if the push - streamed audio is popular among users, then the author believes that this is a relatively more popular audio, or has received accurate push - streaming. This can guide the author's creative direction and also enhance the users' dependence on the platform.
[0054] Therefore, in order to maximize the benefits of pushing multiple works (for multiple authors or platforms), in the scenario where only the user terminal has the decision - making power, the platform needs to adjust the content broadcast by the user terminal.
[0055] The method includes:
[0056] Step S110: When obtaining the combined broadcast policy at the current moment, statistically calculate the push incentives for each of the multiple audio files. Among them, the push incentive for the target audio file is related to the interest weight of the target client for the target audio file. The target audio file is any one of the audio files, and the target client is the client that broadcasts the target audio file. The interest weight refers to the proportion of the interest value of the target client for the target audio file in the interest values of the target client for the multiple audio files, and the interest weight can change over time.
[0057] Step S120: Use the statistically calculated push incentives for each of the multiple audio files as the state factors of the Markov decision at the current moment, and construct the action factors of the Markov decision. Among them, the action factors include influence factors that can enable the client to switch or stop the broadcast of the audio file.
[0058] Step S130: Construct a reinforcement learning model, input the state factors at the current moment into the reinforcement learning model to obtain the action factors at the current moment, and based on the action factors at the current moment and on the basis of preset conditions, obtain the state factors at the next moment. The preset condition is to maximize the total incentive of the multiple audio files, where the total incentive is related to the push incentives for each of the multiple audio files.
[0059] Step S140: And so on, until the action factors for each moment are obtained, and based on the action factors for each moment, adjust the combined broadcast for multiple clients.
[0060] In step S110, assume that the number of multiple clients is m and the number of multiple audio files is n.
[0061] K = {k(1,j), k(2,j), k(3,j),..., k(m,j)};
[0062] Define the policy: It represents the audio file that the platform pushes to each corresponding client. If the first user selects the jth audio file, then k(1,j) = 1; if the first user does not select the jth work, then k(1,j) = 0.
[0063] Among them, the push incentive can be:
[0064]
[0065] Among them, V i,j is the push incentive generated by the ith client for the jth audio file, is the interest weight of the ith client for the jth audio file at the current moment t. Here, the interest weights at different moments may be different, and t i,j is the broadcast duration of the jth audio file on the ith client.
[0066] In this embodiment, the push flow incentive is the product of the interest weight and the broadcast duration. Since the interest weight can measure the user's interest in the audio at the client side, generally, the higher the interest, the more precise the audio delivery is proven. At this time, the corresponding push flow incentive can be given according to the interest weight. Generally speaking, the broadcast duration is one of the assessment indicators on each platform. Therefore, the corresponding push flow incentive can be given according to the broadcast duration. It should be noted that the interest weight changes with time, which is determined by factors such as the user's habits and environment at the client side.
[0067] This embodiment uses the product of the interest weight and the broadcast duration as the push flow incentive, which can improve the accuracy of defining the push flow incentive, and further improve the accuracy of subsequent adjustments.
[0068] In step S120, multiple push flow incentives are used as the push flow degree defined by the platform for multiple audios, that is:
[0069] V = {V 1 , V 2 , V 3 ,... V n};
[0070] Among them, V is the push flow degree (that is, the set of push flow incentives), and V 1 , V 2 , V 3 to V n are push flow incentives.
[0071] This embodiment takes the set V of push flow incentives as the state factor. At different times, there are decision changes at the client side (such as changing the audio). Therefore, the state factor values at different times will be different.
[0072] This embodiment sets an action factor, where the action factor is: an influence factor that enables the client side to switch or stop the broadcast audio. For example,
[0073] The influence factor at least includes user incentives. Among them, the user incentive function is:
[0074]
[0075] Among them, U i,j is the user incentive for the i-th client to broadcast the j-th audio, is the push flow unit price of the j-th audio for the i-th client, is the cost unit price of the i-th client to broadcast the j-th audio.
[0076] Among them, the influence factor can also be the remaining performance of the client side (such as battery power, CPU occupancy), user time, etc., which are not discussed in this embodiment.
[0077] In step S130, a reinforcement learning model is constructed, where the reinforcement learning model can be a Deep Q-Network.
[0078] The Deep Q-Network (DQN) is a method that combines deep learning and reinforcement learning.
[0079] Among them, DQN approximates the Q function by using a deep neural network. DQN introduces an experience replay mechanism. By storing past experiences in a buffer and randomly sampling during training, it improves the utilization rate of data and reduces the correlation between samples. To further stabilize the training process, DQN uses two neural networks with the same structure but different parameters: one for predicting Q values and the other for calculating target Q values. The parameters of the target network are updated regularly, which helps reduce the instability during training.
[0080] It should be noted that the training process of the reinforcement learning model will be described in subsequent embodiments.
[0081] In this embodiment, the state factor at the current moment is input into the reinforcement learning model to obtain the action factor at the current moment. Based on the action factor at the current moment and on the basis of preset conditions, the state factor at the next moment is obtained, and so on until the action factor at each moment is obtained. The preset condition is to maximize the total excitation of multiple audio, and the total excitation is related to the streaming excitation of each of the multiple audio.
[0082] Finally, based on the action factor at each moment, the joint broadcast of multiple audio by multiple clients can be adjusted.
[0083] In the model, the platform can, under preset conditions, intervene in the state factors at different moments by adjusting the action factor, which is the core for adjusting multiple audio.
[0084] This method has at least the following beneficial effects:
[0085] This method sets the streaming incentives for the platform and the authors' audio. Since the interest weight of the user in the audio can affect the streaming incentives of the platform for the authors, the interest weight is determined by the interest value, and the interest value can measure the degree of interest of the user in the audio at different times. Then, in the case of a greater degree of interest, more streaming incentives will be given to the author. Therefore, the set of streaming incentives corresponding to each audio is used as the state factor of the Markov decision, and the streaming effect of the audio is characterized by the state factor. Then, an influence factor for the user to broadcast the audio is set. Since this factor can influence the selection of the broadcast audio by the user side, the influence factor is used as the action factor of the Markov decision to predict the state factor at different times. Finally, the total incentive related to the streaming incentive is set as the reward (Markov decision). Based on the above-set parameters, this method can use the reinforcement learning model to learn and utilize different action factors to adjust the state factor in the case of maximizing the total incentive of multiple audios, so as to adjust the selection of the audio by the user side by the platform in the scenario where only the user side has the right to choose when multiple audios are streamed, and improve the maximum revenue of the platform and the authors.
[0086] In some embodiments, the interest weight refers to the proportion of the interest value of the target user side in the target audio at the current moment among the interest values of the target user side in multiple audios;
[0087] Before obtaining the streaming incentives of each of the multiple audios by counting in the case of obtaining the joint broadcast policy at the current moment, the following steps S210 to S230 are further included:
[0088] Step S210, obtaining the interest values of multiple user sides in multiple audios at historical moments;
[0089] Step S220, taking the interest values at historical moments as time series data, and inputting the time series data into the time series prediction model to predict the interest values of multiple user sides in multiple audios at the current moment based on the time series prediction model;
[0090] Step S230, calculating the interest weight at the current moment according to the interest value at the current moment.
[0091] Furthermore, the function C corresponding to the total incentive includes:
[0092]
[0093] Among them, r1 and r2 are weight values, V j is the streaming incentive corresponding to the jth audio, is the streaming incentive of the jth audio under the adjustment of the action factor a t , and P jis the expected push stream incentive for the j-th audio, k(x,j) is the strategy selected by the x-th client in the joint broadcast strategy, k(x,j) takes values of 0 or 1, and min is the function to take the minimum value; n is the number of multiple audios, m is the number of multiple clients, and s t is the state factor at the current moment t.
[0094] The total incentive of this embodiment includes:
[0095] The first part:
[0096]
[0097] This part shows that the sum of the push stream incentives of the joint broadcast strategy adjusted according to the action factor a t and the sum of the push stream incentives not adjusted according to the action factor a t The difference between them, means taking and P j The minimum value between them, means taking V j and P j The minimum value between them.
[0098] The second part:
[0099]
[0100] This part indicates the product of the user incentives given by the platform to the clients and the number of clients.
[0101] This embodiment defines the total incentive through the incentives of the first part and the second part, which can improve the accuracy of the total incentive definition and further strengthen the adjustment effect.
[0102] This application provides a joint broadcast method, which is applied to the following scenarios:
[0103] The platform (such as Bilibili, QQ Reading, etc.) has the task of pushing streams for multiple works (such as short videos, audios, e-books, etc.). The clients broadcast the pushed works and generate push stream incentives (such as integral incentives for authors). It should be noted that the platform will not actively stop the push stream.
[0104] Multiple clients: Each client can only broadcast one work within a time stamp, and the client can change the work at any time. During the process of the client broadcasting the work, the client will generate client incentives (such as coin incentives).
[0105] Among them, the work selected by the client is affected by the client's interest weight in the work, and the interest weight is measured by the similarity between the client and the work.
[0106] For example, an e - book reading app needs to push n pieces of audio to m client terminals to achieve joint broadcasting on m client terminals. Among them, each client terminal can only broadcast one piece of audio at a certain time stamp, and any user has the option to replace the audio and stop the audio broadcast.
[0107] The purpose of platform push - streaming is:
[0108] Firstly, the e - book reading app generates a push - streaming incentive for the author to encourage the author to continue creating and maintain the activity of the account.
[0109] Secondly, when the client terminal broadcasts the pushed - streamed audio, the generated user incentive can be used to increase user stickiness to the platform.
[0110] Firstly, since users have selectivity for audio (that is, the platform cannot stop push - streaming and switch the stream, while users can adjust in a timely manner based on incentives, etc.). The platform needs to maximize the generation of push - streaming incentives. To generate the maximum push - streaming incentives (total incentives maximized), it is necessary to select an appropriate joint - broadcasting strategy, that is, how to push n pieces of audio to m client terminals, and how the platform adjusts the choices of m client terminals.
[0111] To achieve the purpose, this method includes the following steps:
[0112] Step S910, parameter setting;
[0113] The parameters mainly include the following items:
[0114] (1) The user incentive U for the client terminal to select and broadcast audio i,j .
[0115] The user incentive is composed of the incentive minus the cost. Among them, the incentive can be determined by the product of the broadcast duration and the push - streaming unit price (of course, it can be independently set according to different platforms. For example, the platform assesses the broadcast duration of audio, so the push - streaming unit price is set according to the platform), and the cost is the product of the broadcast duration and the cost unit price (such as traffic consumption, etc.). The following is the function of the user incentive (case - sensitive).
[0116]
[0117] Among them, U i,j is the user - side incentive for the i - th client terminal to broadcast the j - th audio, t i,j is the broadcast duration, is the push - streaming unit price of the j - th audio for the i - th client terminal, is the cost unit price for the i - th client terminal to broadcast the j - th audio.
[0118] (2) The push - streaming incentive V generated by the client terminal for broadcasting the audio i,j: It can be determined based on the user client's interest weight for the audio and the broadcast duration, as follows: the push stream incentive function
[0119]
[0120] Among them, V i,j is the push stream incentive generated by the i-th user client for the j-th audio, is the interest weight of the i-th user client for the j-th audio at time t, and the value ranges from 0 to 1.
[0121] It should be noted that the interest weight here is related to time t, that is, at different times t, the user client's interest in the audio is different.
[0122] (3) During the joint broadcast of n audios by the platform, the push stream incentives promoted by m user clients are statistically calculated:
[0123] V = {V 1 , V 2 , V 3 ,... V n};
[0124] Among them, V is the set of push stream incentives for n audios, and V 1 to V n are the push stream incentives for each audio (it should be noted that the i superscript in V i,j is omitted here).
[0125] Step S920, construct the decision function in the Markov decision-making.
[0126]
[0127] Among them,
[0128] t ∈ {1, 2,..., T};
[0129] s t = V t ;
[0130] K = {k(1, j), k(2, j), k(3, j),..., k(m, j)};
[0131] Among them, Q(s t , a t ) is the decision function, E is the mapping function corresponding to the mathematical expectation, ω t is the weight value at time t, which is set according to experience. For example, if the author believes that the content of the audio becomes more and more important as the broadcast time progresses, ω t can be set to increase with the increase of t, but the value ranges from 0 to 1; the value of T is determined according to the audio, and the upper limit value is infinity; s tThe state factor value representing time t, a t The action factor value representing time t, with a value of {U i,j}, K is the policy, indicating that the platform pushes the corresponding audio to each client. If the first client selects the j-th audio, then k(1, j) = 1; if the first client does not select the j-th audio, then k(1, j) = 0, and so on. L1 is the expected total incentive at time t+1 under the policy K and when the state factor value at time t is s t , and the action factor value is a t . The total incentive at time t+1 is C K (s t , a t ) is the total incentive under the policy K and when the state factor value is s t , and the action factor value is a t .
[0132] The purpose of constructing the Markov decision in this embodiment is:
[0133] In addition to the state factor and the action factor, this embodiment also sets an incentive. Q(s t , α t ) is the decision function, which plays a role in loss calculation in the reinforcement learning model.
[0134] Here, the total incentive C K (s t , a t ) is defined as follows:
[0135]
[0136] Among them, r1 and r2 are weight values (which can be based on experience). Among them, V j is the push stream incentive corresponding to the j-th audio, is the push stream incentive of the j-th audio under the adjustment of the action factor value a t , and P j is the expected push stream incentive of the j-th audio. k(x, j) is the strategy selected by the x-th client in the policy K, and min is the minimum value selection function, that is, selects the minimum value, or selects the minimum value from V j , P j .
[0137] Step S930, construct a Q network and a target Q network based on the Markov decision, and build a reinforcement learning model;
[0138] As Figure 3 , the reinforcement learning model includes:
[0139] The Q network takes the state factor as input and outputs the action factor. The Q network interacts with the environment. The target Q network takes the state factor and the action factor as input and outputs the total incentive generated after the input state factor selects the input action factor. The environment is used to interact with the policy. For example, the current action factor is input into the environment for execution, and the state factor and total incentive at the next moment output by the environment are obtained. The experience pool; the loss function;
[0140] The following introduces the training process of the model:
[0141] Step 1: Data acquisition.
[0142] Obtain the policy K at the current moment t. The initialization process is not described here as it is well-known in the field.
[0143] Step 2: Input the state factor s corresponding to the policy K at the current moment t t into the Q network to obtain the action factor a output by the Q network t .
[0144] Step 3: Simulate the output action factor a t in the environment to obtain the state factor s at the next moment t + 1 t+1 and the total incentive C t+1 .
[0145] Step 4: Input the state factor s t+1 , the total incentive C t+1 , the state factor s t , and the action factor a t into the experience pool.
[0146] Step 5: Soft-update the Q network based on the data in the experience pool (samples corresponding to the policy K at moment t).
[0147] Step 6: Update the loss of the target Q network:
[0148]
[0149] where x and X are the number of samplings in the experience pool and the total amount of samples, respectively.
[0150] By minimizing the loss, the value Q(s t , a t |θ Q ) that the policy K can identify for the target Q network is maximized.
[0151] The target Q network fits the value function Q(s t , a t |θ Q) provides a criterion for evaluating the policy optimization of the current t. Among them, the training objective is to maximize the value of the policy under the target Q network.
[0152] It should be noted that the parameter settings of the Q network and the target Q network will not be introduced here.
[0153] Step 7: Repeat the above steps until the number of iterations or other iteration end conditions are reached.
[0154] Finally, it is possible to use the reinforcement learning model to adjust the appropriate action factor a at different times.
[0155] This method sets the push stream incentive for the author's audio. Since the user's interest weight for the audio at the user side is a factor determining the push stream incentive, the interest weight can measure the user's interest level in the audio at different times, and can give the author more push stream incentives when the interest level is higher. Therefore, the set of push stream incentives corresponding to each audio is used as the state factor of the Markov decision, and the push stream effect of the audio is characterized by the state factor; then the user incentive for playing the audio to the user is set. Since the user incentive is an influencing factor that can cause the user side to switch or stop playing the said audio, the user incentive is used as the action factor of the Markov decision to predict the state factor at different times. Finally, the total incentive related to the push stream incentive is set as the reward (Markov decision). Based on the above-set parameters, this method can learn to adjust the state factor using different action factors by maximizing the total incentive of multiple audios, so as to realize the adjustment of the audio selected by the user side by the platform in the scenario of many-to-many push stream tasks where only the user side has the right to choose, and improve the maximum benefits of the platform and the author.
[0156] On the other hand, is a key factor that can change with time and can also determine the user side's choice of playing the audio.
[0157] This application provides the following calculation method:
[0158] Here, a time series prediction model is adopted, with the interest values corresponding to different audios at different times by the user side as the input, to predict the user side's interest weight (taking values from 0 to 1) for different audios at a specific time.
[0159] Meanwhile, there is also interaction between user terminals. Therefore, when making predictions, the user terminal dimension is taken into account, that is, the interest values corresponding to multiple audio at different times for multiple user terminals are used as input. The interest value is the preference value that the user terminal has for the audio. Different platforms have different statistical methods for interest values. For example, QQ Reading gives the user terminal's interest value for the audio based on the broadcast duration of similar audio.
[0160] For example, taking the interest values of a user terminal for an audio at each moment as a time series data, and constructing a time series prediction model (such as an autoregressive integrated moving average model) based on each time series data:
[0161]
[0162] Among them, represents the preference score at the predicted time t, x t-1 to x t-p represent the respective elements of the time series, α represents the coefficient of the p-order time series prediction model, β represents the coefficient of the q-order time series model, ε represents the influencing factor, and u is the constant term.
[0163] Calculate the interest values corresponding to each time series at the prediction time according to each time series prediction model.
[0164] Then calculate the interest weight (i.e., the weight) based on the interest value.
[0165] Among them, the interest weight is the weight of a user terminal's interest in an audio in the user terminal's interest in all audio. By predicting the preference score of a user terminal for a certain audio at time t through the time series prediction model, the interest weight of the user terminal for the audio at the prediction time can be obtained based on the proportion of the user terminal's preference score for the audio at the prediction time to all audio at the prediction time.
[0166] The method for calculating the interest weight at time t of this method is to obtain the interest values of multiple user terminals for each audio at multiple times, and then take a user terminal and an audio as a time series. Based on the time series prediction model, the interest weights corresponding to each time series data at the prediction time are calculated using each time series data. By combining the time series prediction model for prediction, it fully takes into account the characteristic that the user terminal's interest in the audio changes with time, and can improve the accuracy of interest weight calculation.
[0167] Such as Figure 4, an embodiment provides a combined broadcast system applied to a combined broadcast network. The combined broadcast network includes multiple client terminals and multiple audio files. The multiple client terminals execute a combined broadcast strategy, and the combined broadcast strategy includes: enabling each client terminal to broadcast one of the multiple audio files, and one audio file can be broadcast on one or more client terminals. The system includes:
[0168] The data statistics module 1100 is used to, when obtaining the combined broadcast strategy at the current moment, statistically calculate the push stream incentives of each of the multiple audio files. Among them, the push stream incentive of the target audio file is related to the interest weight of the target client terminal in the target audio file. The target audio file is any one of the audio files, and the target client terminal is the client terminal that broadcasts the target audio file. The interest weight refers to the proportion of the interest value of the target client terminal in the target audio file among the interest values of the target client terminal in the multiple audio files, and the interest weight can change over time.
[0169] The action factor and state factor construction module 1200 is used to use the statistically calculated push stream incentives of each of the multiple audio files as the state factors of the Markov decision at the current moment, and construct the action factors of the Markov decision. Among them, the action factors include influence factors that can enable the client terminal to switch or stop the broadcast audio file.
[0170] The action factor adjustment module 1300 is used to construct a reinforcement learning model, input the state factors at the current moment into the reinforcement learning model to obtain the action factors at the current moment, and based on the action factors at the current moment, obtain the state factors at the next moment based on preset conditions, and so on until the action factors at each moment are obtained. Based on the action factors at each moment, the combined broadcast of the multiple client terminals is adjusted. The preset condition is to maximize the total incentive of the multiple audio files, and the total incentive is related to the push stream incentives of each of the multiple audio files.
[0171] It should be noted that the combined broadcast system provided in this embodiment and the above combined broadcast method are based on the same inventive concept. Therefore, the relevant content of the above combined broadcast method also applies to the content of the combined broadcast system. Therefore, it will not be elaborated here.
[0172] As Figure 5 , this embodiment of the present application also provides an electronic device, and this electronic device includes:
[0173] At least one memory;
[0174] At least one processor;
[0175] At least one program;
[0176] The program is stored in the memory, and the processor executes at least one program to implement the combined broadcast method described above in the present disclosure.
[0177] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0178] The following provides a detailed introduction to the electronic device according to the embodiments of the present application.
[0179] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0180] The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute the combined broadcast method of the embodiments of the present application.
[0181] The input / output interface 1800 is used to implement information input and output;
[0182] The communication interface 1900 is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);
[0183] The bus 2000 transmits information between the various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900);
[0184] Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 achieve communication connections with each other inside the device through the bus 2000.
[0185] The embodiments of the present application also provide a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the above-mentioned combined broadcast method.
[0186] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0187] The embodiments described in this application are for more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0188] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0190] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0191] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0192] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the associated objects before and after. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0193] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0194] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0195] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0196] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0197] The above has specifically described the preferred implementation of the embodiments of this application. However, the embodiments of this application are not limited to the above implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the embodiments of this application. These equivalent deformations or substitutions are all included within the scope defined by the claims of the embodiments of this application.
Claims
1. A combined broadcast method, characterized in that, Applied to a joint broadcast network, the joint broadcast network includes a platform and multiple user terminals. The platform is used to control the multiple user terminals to execute a joint broadcast strategy for multiple audio files. The joint broadcast strategy includes: enabling each of the user terminals to broadcast one of the multiple audio files, and any one of the audio files can be broadcast on one or more of the user terminals. The method includes: According to the joint broadcast strategy at the current moment, calculate the push incentives for each of the multiple audio files; wherein, the push incentive for the target audio file is related to the interest weight of the target user terminal in the target audio file. The target audio file is any one of the multiple audio files, and the target user terminal is the user terminal that broadcasts the target audio file among the multiple user terminals. The interest weight refers to the proportion of the interest value of the target user terminal in the target audio file in the sum of the interest values of the target user terminal in the multiple audio files, and the interest value of the user terminal in the audio file changes over time; Use the calculated push incentives for each of the multiple audio files as the state factors of the Markov decision at the current moment, and construct the action factors of the Markov decision. Among them, the action factors include influence factors that can cause the user terminal to switch or stop broadcasting the audio file; Construct a reinforcement learning model, input the state factors at the current moment into the reinforcement learning model to obtain the action factors at the current moment, and based on the action factors at the current moment, and based on preset conditions, obtain the state factors at the next moment, and so on, until the action factors at each moment are obtained. Based on the action factors at each moment, the platform adjusts the joint broadcast strategy for the multiple user terminals; the preset condition is to maximize the total incentive of the multiple audio files, where the total incentive is related to the push incentives for each of the multiple audio files.
2. The combined broadcast method according to claim 1, wherein Before calculating the push incentives for each of the multiple audio files according to the joint broadcast strategy at the current moment, the following steps are also included: Obtain the interest values of the multiple user terminals in the multiple audio files at historical moments; Use the interest values at the historical moments as time series data, and input the time series data into a time series prediction model to predict the interest values of the multiple user terminals in the multiple audio files at the current moment based on the time series prediction model; Calculate the interest weights at the current moment according to the interest values at the current moment.
3. The combined broadcast method according to claim 2, wherein The time series prediction model is an autoregressive integrated moving average model with exogenous inputs.
4. The combined broadcast method according to claim 1, wherein Calculating the push incentive for the target audio file includes: Determine the interest weight of the target user terminal in the target audio file at the current moment; Determine the broadcast duration of the target audio file on the target user terminal; Obtain the push incentive for the target audio file according to the product of the interest weight and the broadcast duration.
5. The combined broadcast method according to claim 4, characterized in that, Before constructing the action factors of the Markov decision, the method further includes: Determine the cost unit price for any user terminal to broadcast any audio file, and the push unit price for the platform to push any audio file to any user terminal; Multiply the broadcast duration of any of the above-mentioned audio on any of the above-mentioned user terminals by the push stream unit price to obtain a positive incentive; Multiply the broadcast duration of any of the above-mentioned audio on any of the above-mentioned user terminals by the cost unit price to obtain a negative incentive; Subtract the negative incentive from the positive incentive to obtain the user incentive for any of the above-mentioned user terminals; Use the user incentive as the influence factor.
6. The combined broadcast method according to claim 5, wherein The total incentive is expressed by the following function C: where r1 and r2 are weight values, V j is the streaming incentive corresponding to the j-th audio, is the streaming incentive of the j-th audio under the adjustment of the action factor a t , P j is the expected streaming incentive of the j-th audio, k(x, j) is the strategy selected by the x-th client in the joint broadcast strategy, k(x, j) takes values of 0 or 1, min is the function of taking the minimum value; n is the number of multiple audios, m is the number of multiple clients, s t is the state factor at the current time t, U i,j is the user incentive for the i-th client to broadcast the j-th audio.
7. The combined broadcast method according to claim 6, characterized in that, The reinforcement learning model is a deep Q network.
8. A combined broadcast system, characterized in that, Applied to a joint broadcast network, the joint broadcast network includes a platform and multiple user terminals. The platform is used to control the multiple user terminals to execute a joint broadcast strategy for multiple audio. The joint broadcast strategy includes: enabling each of the user terminals to broadcast one of the multiple audio, and any one of the audio can be broadcast on one or more of the user terminals. The system includes: A data statistics module, configured to count the push stream incentives of each of the multiple audio according to the joint broadcast strategy at the current moment; wherein, the push stream incentive of the target audio is related to the interest weight of the target user terminal in the target audio; the target audio is any one of the multiple audio, and the target user terminal is the user terminal that broadcasts the target audio among the multiple user terminals; the interest weight refers to the proportion of the interest value of the target user terminal in the target audio in the sum of the interest values of the target user terminal in the multiple audio, and the interest value of the user terminal in the audio changes over time; An action factor state factor construction module, configured to use the counted push stream incentives of each of the multiple audio as the state factor of the Markov decision at the current moment, and construct the action factor of the Markov decision, wherein the action factor includes an influence factor that can cause the user terminal to switch or stop the broadcast of the audio; An action factor adjustment module, configured to construct a reinforcement learning model, input the state factor at the current moment into the reinforcement learning model to obtain the action factor at the current moment, and based on the action factor at the current moment, obtain the state factor at the next moment based on a preset condition, and so on until the action factor at each moment is obtained. Based on the action factor at each moment, the platform adjusts the joint broadcast strategy of the multiple user terminals; the preset condition is to maximize the total incentive of the multiple audio, wherein the total incentive is related to the push stream incentives of each of the multiple audio.
9. An electronic device, characterized in that, Including: At least one control processor and a memory for communicating with the at least one control processor; The memory stores instructions executable by the at least one control processor. The instructions are executed by the at least one control processor to enable the at least one control processor to execute the joint broadcast method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the joint broadcast method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Individualized short video recommendation method and system based on reinforcement learning
CN113282787A
Data pushing method and system, electronic device and storage medium
WO2021169218A1