Decision model training method, live broadcast decision determination method and device
By constructing decision-making prompt texts and triplet clustering training samples, the adaptability and accuracy of the digital human live streaming decision-making model were improved, solving the problem of insufficient decision-making by existing models under diverse needs and improving the live streaming effect.
Patent Information
- Application Number
- CN202511841895.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-02-27
AI Technical Summary
Existing digital human live streaming decision-making models are unable to meet the diverse and complex needs of viewers, resulting in a lack of innovation and optimization in operational decision-making methods in rapidly changing live streaming scenarios.
By constructing decision prompt text, determining the operation based on the live room status, and constructing triples containing status, operation, and reward, clustering is performed to generate training samples to train the decision model and improve its adaptability and accuracy.
It improves the accuracy and stability of decision-making in digital human live streaming, enhances the practicality of operations, and improves live streaming effects such as product conversion rate and audience interaction experience.
Smart Images

Figure CN121585837A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer technology, and particularly relates to the technical fields of data processing, deep learning, digital person live broadcast, etc. BACKGROUND
[0002] In recent years, with the rapid development of the live broadcast industry, digital person live broadcast as a new force has emerged, and has attracted attention due to its advantages such as being not limited by time and space and being able to provide stable live broadcast services. In order to enhance the interactivity of live broadcast, current digital person live broadcast mostly sets a trigger point interaction operation based on preset rules, and the operation content is generated by a decision model. However, as the application scenarios of digital person live broadcast expand, the existing decision model is difficult to meet the innovation and optimization needs of digital person live broadcast. SUMMARY
[0003] The present disclosure provides a decision model training method, a live broadcast decision determination method and device.
[0004] According to an aspect of the present disclosure, a decision model training method is provided, comprising: constructing a decision prompt text based on a state of a live broadcast room; determining an operation for the state by using a decision model based on the decision prompt text, and executing the operation; determining a reward after executing the operation; constructing log sample data based on live broadcast data of the live broadcast room, the log sample data comprising a plurality of triplets, each triplet comprising a state, an operation and a reward having a corresponding relationship; clustering the plurality of triplets based on the state to obtain at least one class cluster; determining a training sample based on a plurality of triplets belonging to the same class cluster; training the decision model by using the training sample.
[0005] According to another aspect of the present disclosure, a live broadcast decision determination method is provided, comprising: determining a decision prompt text based on a state of a live broadcast room; determining an operation for the state by using a decision model based on the decision prompt text; The decision model is trained by using the training method provided by the present disclosure.
[0006] According to another aspect of the present disclosure, a decision model training device is provided, comprising: a first prompt construction module configured to construct a decision prompt text based on a state of a live broadcast room; a first operation determination module configured to determine an operation for the state by using a decision model based on the decision prompt text, and execute the operation; The first reward determining module is configured to determine a reward after the operation is performed; and construct log sample data based on the live streaming data of the live streaming room, the log sample data including a plurality of triplets, each triplet including a state, an operation and a reward having a corresponding relationship; The training sample determining module is configured to cluster the plurality of triplets based on the state to obtain at least one cluster; and determine a training sample based on the plurality of triplets belonging to the same cluster. The model training module is configured to train the decision model using the training sample.
[0007] According to another aspect of the present disclosure, a live streaming decision determining apparatus is provided, including: The second prompt constructing module is configured to determine a decision prompt text based on the state of the live streaming room. The second operation determining module is configured to determine an operation for the state using a decision model based on the decision prompt text. The decision model is trained by the training apparatus provided by the present disclosure.
[0008] According to another aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected with the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.
[0009] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method according to any of the embodiments of the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.
[0011] The present disclosure constructs decision prompt text based on the real-time state of a live broadcast room, can accurately capture key information in the live broadcast process, and provides targeted input for a decision model; then, the decision model is used to determine and execute operations according to the decision prompt text, realizes the connection of decision and execution, and effectively responds to the dynamic changes of the live broadcast; further, the rewards after the execution operations are determined, and a plurality of triplets containing the corresponding relationship of state, operation and reward are constructed, providing a data basis for the training of the decision model. By clustering the triplets based on the state and determining the training samples, the patterns and rules of the data can be mined, the model training is more representative and targeted, and then the training samples are used to train the decision model, which can make the decision model continuously learn and adapt to the state of the live broadcast room, and improve the accuracy, stability and generalization ability of the decision. On this basis, the operations output by the decision model in the live broadcast process can better meet the actual live broadcast demand, thereby improving the live broadcast effect.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them: Figure 1 is an application scenario diagram according to an embodiment of the present disclosure; Figure 2 is an implementation flowchart of a decision model training method according to an embodiment of the present disclosure; Figure 3 is a flowchart of training a decision model according to an embodiment of the present disclosure; Figure 4 is an implementation flowchart of a live broadcast decision determination method according to an embodiment of the present disclosure; Figure 5 is a flowchart of determining a live broadcast decision operation according to an embodiment of the present disclosure; Figure 6 is a structural diagram of a decision model training device 600 according to an embodiment of the present disclosure; Figure 7 is a structural diagram of a decision model training device 700 according to an embodiment of the present disclosure; Figure 8 is a structural diagram of a live broadcast decision determination device 800 according to an embodiment of the present disclosure; Figure 9 is a structural diagram of a live broadcast decision determination device 900 according to an embodiment of the present disclosure; Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0014] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0015] The term "and / or" in this disclosure indicates that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The term "at least one" in this document means any combination of at least two of a plurality of options, such as including at least one of A, B, and C, which can mean including any one or more elements selected from the set of A, B, and C. The terms "first" and "second" in this document refer to and distinguish multiple similar technical terms, and do not imply a specific order or a limitation to only two. For example, "first feature" and "second feature" refer to two types / two features; the first feature can be one or more, and the second feature can also be one or more.
[0016] In recent years, the live streaming industry has flourished, and digital human live streaming, as an emerging form of live streaming, is gradually gaining prominence and demonstrating enormous development potential. Digital human live streaming, with its unique advantages such as being unrestricted by time, space, and manpower, can provide stable and continuous live streaming services, bringing viewers a novel viewing experience.
[0017] In the operation of digital human live streaming, interactivity is a key factor in improving audience engagement and retention. To achieve effective interaction with the audience, current digital human live streaming primarily relies on preset rules to set specific trigger points for actions. This preset rule approach can, to some extent, improve the orderliness and controllability of action execution. Simultaneously, to make action execution more natural, richer, and personalized, actions can often be determined using a trained decision model. This decision model, through learning and analysis of large amounts of data, can determine appropriate actions based on different scenarios and audience feedback, thereby enhancing the attractiveness of digital human live streaming.
[0018] However, with the rapid development of digital human live streaming, its application scenarios continue to expand, and audience needs are increasingly diversified and complex. Digital human live streaming faces rapidly changing development trends, and higher performance requirements are placed on decision-making models. Under this application background, the existing digital human live streaming operation decision-making methods gradually expose potential challenges, to some extent, restricting the innovation and optimization of digital human live streaming.
[0019] To solve the above problems, the disclosure embodiment proposes a training method of a decision-making model. Figure 1 is an application scenario diagram according to the disclosure embodiment, as Figure 1 shown, the application scenario diagram of the disclosure embodiment can include but is not limited to a model training device 110 and a decision-making model 120, which can communicate through any type of wired or wireless network. Specifically, the decision-making model 120 can determine the operation to be performed in the digital human live streaming process according to the state of the live streaming room, and then determine the posterior data of performing the operation based on the execution result of the operation, wherein the posterior data can include the reward obtained by performing the operation in the digital human live streaming process. Further, based on the state of the live streaming room, the determined operation and the reward, training data is constructed. The model training device 110 can train the decision-making model 120 based on the training data.
[0020] In the disclosure embodiment, the model training device 110 can include a server for providing background management for the decision-making model 120. In addition, the disclosure embodiment does not specifically limit the number of model training devices 110, for example, the application scenario diagram of the disclosure embodiment can include one or more model training devices 110.
[0021] Figure 2 is an implementation flowchart of a training method of a decision-making model according to an embodiment of the disclosure, including: S210, constructing a decision-making prompt text based on the state of the live streaming room; S220, determining an operation for the state using a decision-making model based on the decision-making prompt text, and performing the operation; S230, determining the reward after performing the operation; S240, constructing log sample data based on live streaming data of the live streaming room, the log sample data including a plurality of triplets, each triplet including a state, an operation and a reward having a corresponding relationship; S250, clustering the plurality of triplets based on the state to obtain at least one class cluster; S260, determining a training sample based on a plurality of triplets belonging to the same class cluster; S270, training the decision-making model using the training sample.
[0022] In the embodiments of the present disclosure, the state of the live room can include features of the live room in the live process. In an example, the features of the live room can be aggregated to obtain the state of the live room. For example, the features of the live room can include the current number of online users in the live room, the number of orders in a fixed time period, the number of comments, the name of the product, the price of the product, and the like.
[0023] Further, the present disclosure can utilize the state of the live room to construct a decision prompt text. In the embodiments of the present disclosure, the decision prompt text input into the decision model can include a task background, for example, “you are an experienced host”, to clarify the role positioning of the decision model.
[0024] The task target in the decision prompt text can also be included, for example, “according to the state of the live room, rank the operations in the preset operation list from best to worst, and select the most suitable operation, and return the identification of the selected operation, the ranking of the operations, and the ranking basis”, which clarifies the specific goal to be achieved, that is, to evaluate and rank the operations in the preset operation list in combination with the state of the live room to find the optimal operation. It can be understood that the state of the live room and the preset operation list can be included in the decision prompt text.
[0025] The task requirements in the decision prompt text can also be included, for example, output format requirements (which can include ranking results, ranking basis, and optimal operation identification), ranking integrity requirements (that is, requiring evaluation and ranking of all operations), comprehensive consideration dimensions (such as requiring comprehensive consideration of the dimensions of discount strength, user distribution difference, user focus, user activity, and user inclination for evaluation), ranking basis explanation (requiring the ranking basis to be explained for the ranking of each operation), word limit (such as stipulating that the ranking basis cannot exceed 1000 tokens), and content limit (such as not outputting any other content in addition to the optimal operation identification, the ranking results of the operations, and the ranking basis).
[0026] In the embodiments of the present disclosure, the decision model can be a deep learning model with intelligent processing and reasoning capabilities, which can receive and understand the information conveyed by the decision prompt text. In an example, the decision model can deeply analyze the decision prompt text to determine the operation for the state of the live room. The related content of determining the operation will be described in detail later. In an example, the decision model can include a neural network model (such as a deep reinforcement learning (DRL) model containing a multi-layer perceptron), and can also include a large language model (LLM).
[0027] Further, after determining the operation, it needs to be conveyed to the relevant execution subject. In an example, in the digital human live broadcast scene, the execution subject can be a digital human anchor. For example, the operation is to introduce the efficacy and characteristics of the commodity, and the digital human anchor can introduce the relevant information of the commodity to the audience during the live broadcast.
[0028] In the embodiments of the present disclosure, for the live broadcast scene, the reward is a quantitative evaluation of the effect generated after performing a specific operation, and the purpose is to measure the contribution degree of the operation to the live broadcast. In an example, the factors for determining the reward can include sales indicators, user interaction indicators, user retention indicators, etc. For example, if the sales of the commodity in the live broadcast room is significantly improved after performing the operation, the reward value can be determined according to the growth rate of the sales; if the user interaction (such as the number of comments, the number of likes, the number of shares, etc.) is increased after performing the operation, it means that the operation effectively attracts the attention of the user and improves the participation of the user, and the reward can be determined according to the interaction; if the average length of stay of the user is prolonged or the retention rate is improved after performing the operation, it means that the operation has stronger attraction to the user, and the reward can be determined according to the changes of these indicators.
[0029] The present disclosure can collect the state of the live broadcast room, the performed operation, and the reward generated after performing the operation in real time. In an example, the data collection can be realized by using the data interface of the live broadcast platform, the background statistical data, or the self-defined data collection system.
[0030] Further, the present disclosure organizes the collected state of the live broadcast room, the performed operation, and the reward, so that each state is matched with the corresponding operation and reward. For example, when a certain operation is performed in a certain state, the reward obtained after performing the operation in the state is recorded, forming a complete triple. The multiple triplets after organization are stored as log sample data, which can be stored in the form of a database, a file system, etc. for subsequent model training and analysis.
[0031] In the embodiments of the present disclosure, a large number of triples containing states, operations, and rewards are collected, and these triples cover various live broadcast scenarios, but directly using these raw data for model training can face problems such as complex data distribution and difficult-to-catch patterns. Therefore, the present disclosure can first cluster the triples according to the state, and classify the triples with similar states into a class to form a class cluster, and then determine the training samples within each class cluster. Here, the specific clustering method will be described in detail in the subsequent content.
[0032] Further, the present disclosure can train the existing decision model using the training samples, so that the performance of the model is more optimal and the generalization ability is stronger.
[0033] By employing the above method, decision prompt text is constructed based on the real-time status of the live stream, accurately capturing key information during the live stream and providing targeted input for the decision-making model. Then, the decision-making model determines and executes operations based on the decision prompt text, achieving a seamless connection between decision-making and execution, effectively responding to the dynamic changes of the live stream. Furthermore, the reward after executing the operation is determined, and multiple triples containing the correspondence between status, operation, and reward are constructed, providing a data foundation for training the decision-making model. By clustering the triples based on status and determining training samples, patterns and regularities in the data can be mined, making model training more representative and targeted. Using these training samples to train the decision-making model allows it to continuously learn and adapt to the live stream's status, improving the accuracy, stability, and generalization ability of decisions. Based on this, the operations output by the decision-making model during the live stream are more in line with actual live stream needs, thereby improving the live stream effect. This improved effect can include increasing product conversion rates, enhancing audience interaction, and attracting and retaining viewers.
[0034] Figure 3 This is a flowchart illustrating the training of a decision model according to an embodiment of the present disclosure.
[0035] like Figure 3 As shown, in one example, the decision triggering module 310 can monitor the dynamic changes in user traffic in real time. Once a new user enters, it can initiate a decision-making process for the live stream. In another example, a timer can be set according to a preset time interval or a specific time point. Based on the timer, the decision triggering module 310 can initiate a decision-making process for the live stream. For example, the decision triggering module 310 can be controlled to trigger a decision request every 10 minutes; or, at a specific time point in the live stream (such as 5 minutes before a product is listed, or when the live stream is about to end), the decision triggering module 310 can be controlled to trigger a decision request through a timer.
[0036] In some implementations, it also includes: Obtain at least one feature of the live stream; Aggregate at least one feature of the live stream to obtain the state of the live stream.
[0037] In this embodiment of the disclosure, at least one feature of the live streaming room can be used to describe the attributes and conditions of the live streaming room in different aspects. These features can be obtained from multiple dimensions, such as the real-time data dimension of the live streaming room, including the current number of online viewers, viewer interaction rate (such as the number of likes, comments, and shares), the number of products sold, and the duration of the live stream; they can also be obtained from the live streaming content dimension, such as the type of product being explained (beauty, clothing, electronic products, etc.) and the theme and style of the live stream (entertainment, education, promotion, etc.). In one example, Table 1 shows multiple features of the live streaming room used in the process of determining the status of the live streaming room in this disclosure.
[0038] Table 1 Further, in the embodiments of the present disclosure, a single feature can only reflect a certain aspect of the live room, and the present disclosure aggregates multiple features (i.e., integrates these scattered features) to form a comprehensive and overall description, and obtains the state of the live room.
[0039] In the above manner, by obtaining at least one feature of the live room and further aggregating these features to obtain the state of the live room, the real-time running condition of the live room can be presented in a comprehensive and overall manner, and data support is provided for the construction of the decision prompt text.
[0040] In the embodiments of the present disclosure, the state of the live room can also include context information of the currently played script. For example, a script for selling a certain drink is currently being played, and the context of the script can be "Today's live room brings the treasure of drinks. Not every drink of the same type has this status. The opportunity is rare, the inventory is tight, and it is unknown how long the next opportunity will be missed." The context of the script can be "The essence of this drink comes from a piece of unique core brewing area. The local unique water source, carefully selected high-quality sorghum, and the ancient brewing process that has been passed down for generations - after a long production cycle, multiple steaming, fermentation, sampling, and several years of cellar deposit, every drop condenses the essence of time and craftsmanship."
[0041] Further, after the decision trigger module 310 triggers the decision request, the decision prompt text can be constructed based on the state of the live room at the time of triggering the decision request.
[0042] As shown in FIG. 3, in the embodiments of the present disclosure, the decision prompt text can be input into the decision model 320, and the decision model 320 can determine the operation for the state of the live room. Figure 3
[0043] In some embodiments, based on the decision prompt text, the operation for the state is determined by using the decision model, including: At least one feature of the live room contained in the decision prompt text is extracted by using the decision model, and the at least one feature includes numerical features and non-numerical features; The numerical features are discretized by using the decision model, and the input features are determined based on the non-numerical features and the discretized numerical features; Based on the input features, the selection probability of each candidate operation in the preset operation list is determined by using the decision model, and the preset operation list is contained in the decision prompt text; The candidate operation with the maximum selection probability is determined as the operation for the state.
[0044] In the embodiments of the present disclosure, the decision model 320 can extract at least one feature of the live room from the decision prompt text, which covers numerical features and non-numerical features. Among them, the numerical features can be quantitatively represented data such as the number of real-time online audience in the live room, the number of commodity sales, the number of likes, etc.; the non-numerical features can be descriptive information such as the theme type of the live room (makeup, food, technology, etc.), the style characteristics of the anchor (humorous and interesting, professional and rigorous, etc.), the current explanation of the commodity category, etc.
[0045] Further, the present disclosure utilizes the decision model 320 to discretize the extracted numerical features (or called bucketing processing). Here, the discretization processing can be to convert continuous numerical features into discrete intervals or categories, for example, the number of audience in the current live room can be divided into 0-100, 101-500, 501-1000 and above 1000 intervals. In this way, it can reduce the calculation complexity caused by too large numerical range, and also better capture the influence of different numerical intervals on decision-making.
[0046] The present disclosure can also utilize the decision model 320 to combine the non-numerical features and the discretized numerical features to determine the input features. In an example, the encoding layer or embedding layer in the decision model 320 can respectively encode or embed the non-numerical features and the discretized numerical features to obtain the feature vectors of the non-numerical features and the feature vectors of the discretized numerical features. Then, the decision model 320 can concatenate the feature vectors of the non-numerical features and the feature vectors of the discretized numerical features to obtain the input features.
[0047] Based on the above determined input features, the present disclosure utilizes the decision model 320 to calculate the selection probability of each candidate operation in the preset operation list. The preset operation list contains a series of selectable operations defined in advance, and the preset operation list is contained in the decision prompt text. In other words, the selectable operations for the state of the live room are explicitly defined in the decision prompt text. In an example, Table 2 shows the selectable operations contained in the preset operation list of the present disclosure.
[0048] Table 2 Further, from the calculated selection probabilities of each candidate operation, the candidate operation with the maximum selection probability is selected as the operation for the state of the live room.
[0049] In the foregoing manner, at least one feature of the live room containing numerical and non-numerical values is extracted from the decision prompt text, so as to comprehensively capture the key information of the live room and provide a data basis for subsequent operation determination. Then, the numerical features are discretized to reduce the complexity of the data, and the input features are determined in combination with the non-numerical features, so as to accurately reflect the actual state of the live room. Further, the selection probability of each candidate operation in the preset operation list is determined based on the input features, so as to improve the objectivity of operation determination, and finally the candidate operation with the maximum selection probability is determined as the operation for the current state. This operation determination method based on data driving and probability analysis evaluates each candidate operation from a quantitative perspective, so that the decision basis is more scientific and reliable, thereby improving the accuracy of the decision.
[0050] Further, the operation for the state of the live room can be input into the operation execution module 330 to execute the operation and obtain the reward after executing the operation.
[0051] In the embodiments of the present disclosure, the state of the live room contained in the decision prompt text, the operation for the state of the live room, and the reward after executing the operation contained in the live data of the live room can be integrated by the log construction module 340 to obtain a plurality of triplets, and then the log sample data is constructed based on the plurality of triplets. In an example, the live data of the live room can be the posterior data for the operation.
[0052] In an example, the presentation form of the triplet is <state, action, reward>.
[0053] Further, the log sample data can be sample extracted and clustered by the sample extraction module 350 to determine the training sample.
[0054] In some embodiments, based on the state, the plurality of triplets are clustered to obtain at least one cluster, including: determining at least one cluster center; determining the distance between each state and each cluster center based on at least one feature of the live room contained in each state; assigning the state to the cluster to which the nearest cluster center belongs to obtain at least one cluster.
[0055] In the embodiments of the present disclosure, the cluster center can be a core point used to represent a class cluster in cluster analysis. The present disclosure can randomly select at least one triplet from the log sample data as an initial cluster center. Here, the number of cluster centers can be a preset number. For example, if the state of the live room is to be divided into 3 class clusters, 3 triplets can be randomly selected from the log sample data as the initial cluster centers.
[0056] The state of the live room includes at least one feature of the live room, which can be a numerical type (such as the number of online people in the live room, the number of likes, etc.), or a non-numerical type (such as the live type, the commodity information, etc.). The present disclosure can vectorize the state of the live room, and determine the distance between the vectorized state of the live room and the cluster center.
[0057] In an example, after the state of the live room is vectorized, the state of the live room can be represented as , assuming that a cluster center is represented as , wherein, represents the 1st to nth features of the live room, represents the 1st to nth features of the state of the live room corresponding to the cluster center. The present disclosure can determine the distance between the features of the live room and the cluster center by using the Euclidean distance, and the formula can be represented as: (1) After the distance between the state of the live room and each cluster center is calculated, for each state, the nearest cluster center is found, and then the state is assigned to the class cluster to which the cluster center belongs. In this way, all states can be assigned to the corresponding class cluster, and at least one class cluster is formed.
[0058] For example, assuming that there are three cluster centers , and , and five states , , , and . By calculating the distance between each state and each cluster center, it is determined that is closest to , is closest to , and is closest to , and is assigned to the class cluster to which belongs, and is assigned to the class cluster to which The class cluster to which the state belongs, and is assigned to The class cluster to which the state belongs, and thus three class clusters are obtained. In each class cluster, the corresponding state is contained, and the operation and the reward having the corresponding relationship with the state are contained.
[0059] In the above manner, on the basis of the characteristics of the live broadcast room, the states of the live broadcast rooms having similar characteristics are classified into the same class cluster by means of the distance measurement, so that the potential relationship and the internal law between the states of the live broadcast rooms can be mined. In this way, the characteristics and the distribution of the states of different types of live broadcast rooms can be understood, and a data basis for subsequent construction of training samples is provided.
[0060] In some embodiments, the method further comprises: determining a feature mean value of each feature based on at least one feature of the live broadcast room contained in each state in the class cluster; updating the cluster center of the class cluster based on the feature mean value of each feature.
[0061] In the embodiments of the present disclosure, after the preliminary clustering of the states of the live broadcast rooms is completed, a plurality of class clusters are obtained, each class cluster contains at least one state of the live broadcast room, and each state has a plurality of features describing the live broadcast room, such as the number of online users, the gift income, and the interaction frequency. In an example, there is a class cluster C containing n states of live broadcast rooms Each state has m features For a feature , a feature mean value of all features in the class cluster on can be calculated, and the calculation formula is: (2) Wherein, represents the value of the feature of the state .
[0062] For example, there are 3 states in a class cluster, and the feature values of the number of online users in the states are 1000, 1200, and 800 respectively. Therefore, the feature mean value of the number of online users in the class cluster is (1000+1200+800) / 3=1000.
[0063] After obtaining the feature mean values of each feature in the class cluster, the cluster center of the class cluster is determined again by using the mean values.
[0064] For example, there are three features in the state: the number of online people, the gift income, and the interaction frequency; there are three states in a cluster, i.e., state 1, state 2, and state 3. Among them, the feature data corresponding to state 1 are (100, 500, 20) respectively; the feature data corresponding to state 2 are (150, 800, 30) respectively; and the feature data corresponding to state 3 are (120, 600, 25) respectively. Further, based on the feature data contained in each state in the cluster, the coordinates of the new cluster center are determined as [(100+150+120) / 3, (500+800+600) / 3, (20+30+25 / 3)], i.e., (370 / 3, 1900 / 3, 75 / 3), and then the cluster center of the cluster to which the state belongs is updated by using the coordinates of the new cluster center.
[0065] In the above manner, the cluster center of the cluster is updated by determining the feature mean of each feature in the state, which can more accurately capture the core features of the cluster, make the cluster center closer to the typical feature distribution of the states in the cluster, improve the accuracy of subsequent assignment of the state to the cluster, reduce the misclassification, and provide a data basis for the construction of the training sample.
[0066] Further, in some embodiments, the training sample is determined based on a plurality of triplets belonging to the same cluster, including: The plurality of triplets belonging to the same cluster are arranged in descending order of the reward to obtain a triplet sequence; The training sample is determined based on a preset number of head triplets located at the head of the triplet sequence and a preset number of tail triplets located at the tail of the triplet sequence.
[0067] In the embodiments of the present disclosure, the triplets can be sorted from high to low according to the reward in the triplets. In an example, the reward can be quantified by using a model that estimates the reward value corresponding to the reward. Further, the triplet with the highest reward value is arranged at the front of the triplet sequence, and the triplet with the lowest reward value is arranged at the rear of the triplet sequence.
[0068] In the embodiments of the present disclosure, a specific numerical value (i.e., a preset number) can be preset, which determines the number of head triplets and tail triplets selected from the sorted triplet sequence. For example, 5 head triplets and 5 tail triplets are selected. It should be noted that in the present disclosure, the number of head triplets and tail triplets is the same.
[0069] In an example, the preset number can be represented by a proportion. For example, the head triplets can be the first k% triplets in the triplet sequence, and the tail triplets can be the last k% (i.e., the first 1-k%) triplets in the triplet sequence.
[0070] Further, the training sample is determined based on the head triplets and the tail triplets.
[0071] In the above manner, the head triplets of the same cluster are arranged in descending order of rewards, and the preset number of head triplets at the head of the sequence usually correspond to high-reward states and operations, thereby improving the quality of the training sample. Further, the training sample is determined based on the head triplets and the tail triplets, so that the decision model can make reasonable predictions for operations when facing new states, thereby improving the generalization ability of the decision model.
[0072] In some embodiments, the training sample is determined based on the preset number of head triplets at the head of the sequence of triplets and the preset number of tail triplets at the tail of the sequence of triplets, and includes: replacing the operation and the reward in the tail triplets with the operation and the reward in the head triplets to obtain reorganized triplets; determining the training sample based on the reorganized triplets, the head triplets, and other triplets except the tail triplets and the head triplets.
[0073] In the embodiments of the present disclosure, for the tail triplets , the operation and the reward are replaced with the operation and the reward of the head triplets, and the obtained reorganized triplets can be expressed as . It should be noted that the replaced operation and reward can come from the same head triplet.
[0074] In some embodiments, the replacement manner includes sequential coverage replacement, reverse sequential coverage replacement, or random sampling replacement.
[0075] In the embodiments of the present disclosure, the sequential coverage replacement refers to sequentially replacing the operation and the reward in the head triplets into the tail triplets according to the original order of the head triplets in the sorted sequence. Specifically, the first tail triplet is replaced with the operation and the reward of the first head triplet, the second tail triplet is replaced with the operation and the reward of the second head triplet, and so on until all tail triplets are replaced.
[0076] In an example, it is assumed that in the sorted sequence of triplets, there are 3 head triplets , and , and 3 tail triplets , and . After the sequential coverage replacement, the reorganized triplets are , and .
[0077] Reverse order coverage replacement is opposite to the order coverage replacement. In this way, the operation and the reward in the head triplets are replaced into the tail triplets in the reverse order in the sorted sequence. That is, the first tail triplet is replaced by the operation and the reward of the last head triplet, the second tail triplet is replaced by the operation and the reward of the second last head triplet, and so on.
[0078] Taking the above examples of the head triplets and the tail triplets, after using the reverse order coverage replacement, the recombined triplets are , and .
[0079] Random sampling replacement means that a triplet is randomly sampled from the head triplets, and the operation and the reward thereof are replaced into the tail triplets. In each replacement, the probability of each head triplet being selected is equal, and each sampling is independent. That is, the same head triplet can be selected for replacing different tail triplets for multiple times, or can not be selected for once.
[0080] Taking the above examples of the head triplets and the tail triplets, the first random sampling is H2 , then T1 is replaced by ; the second random sampling is H1 , then T2 is replaced by ; the third random sampling is H3 , then T3 is replaced by .
[0081] Further, the disclosure combines the recombined triplets, the head triplets, and other triplets (i.e., the triplets in the middle part) other than the tail triplets and the head triplets together as training samples.
[0082] Using the above method, by constructing the recombined triplets, not only the number of high-reward operation samples is increased, but also the training samples are formed together with the head triplets and other middle triplets, so as to provide the decision model (i.e., the decision model) with more abundant state-operation-reward scenarios. Such diversified training samples can enable the decision model to learn general decision rules, and thus to make operations closer to the posterior data, thereby improving the generalization ability and decision accuracy of the decision model.
[0083] Further, as shown in Figure 3 , the disclosure can train the decision model 320 by using the training samples.
[0084] In some embodiments, the decision model is trained using the training samples, including: inputting the state contained in the triple in the training sample into the decision model to determine a predicted operation corresponding to the state; performing the predicted operation and determining an actual reward after performing the predicted operation; determining a loss function based on at least one of the predicted operation, the actual reward, and the operation and the reward corresponding to the state; training the decision model using the loss function.
[0085] In the embodiments of the present disclosure, the state of the triple in the training sample is input into the decision model 320, and the decision model 320 can output a predicted operation based on the current state. Here, the predicted operation can be determined according to the selection probability of each candidate operation in the preset operation list.
[0086] Further, in the live broadcast scenario, the predicted operation output by the decision model 320 is performed, and the live broadcast scenario returns an actual reward (such as user stay time, conversion rate improvement, etc.) for the predicted operation according to the operation execution result. Here, the actual reward can be considered as a quantitative feedback of the predicted operation.
[0087] In the embodiments of the present disclosure, the operation corresponding to the state is contained in the training sample, and the present disclosure can calculate a loss function using the predicted operation and the operation. Here, the loss function can be a mean square error. The calculation formula can be: (3) wherein, represents the predicted operation, represents the operation corresponding to the state, and n is the number of training samples. In an example, using the loss function (i.e. ) can directly reduce the numerical gap between the predicted operation and the operation corresponding to the state. For example, by updating the parameters of the decision model through gradient descent, the gap between the predicted operation and the real operation (i.e., the operation corresponding to the state) is minimized.
[0088] In the embodiments of the present disclosure, the reward after performing the operation under the current state is contained in the training sample, and the present disclosure can calculate a loss function using the actual reward after performing the predicted operation and the reward after performing the operation. Here, the loss function can be a mean square error, and the calculation method is similar to formula (3), and finally is obtained. Wherein, can be considered as a loss function for training the decision model. In an example, using the loss function (i.e. The gap between the reward after performing the operation and the actual reward after performing the predicted operation can be directly reduced, for example, by back-propagating the gradient of a loss function through a gradient descent algorithm to update the parameters of the decision model, so that the predicted operation determined by the decision model in the subsequent iteration is more inclined to obtain an actual reward close to the historical optimal reward (i.e., the reward after performing the operation), thereby minimizing the gap between the predicted operation and the operation corresponding to the state.
[0089] The present disclosure can also simultaneously utilize the operation loss function ( ) and the reward loss function ( ) to calculate the loss function (which can be referred to as ) for training the decision model. In an example, the corresponding weight coefficients can be set for the operation loss function and the reward loss function, respectively, and then the operation loss function and the reward loss function are weighted by using the weight coefficients to obtain the loss function for training the decision model, and the calculation formula can be: (4) wherein, β can be the weight coefficient of the operation loss function, and β can be the weight coefficient of the reward loss function, and The sum of α and β is 1.
[0090] In the above manner, by introducing the predicted operation and the actual reward, the decision model can capture the changes of the live broadcast, and the parameters of the decision model can be adjusted in real time through the loss function back propagation, thereby improving the adaptability of the decision model to the live broadcast scene and improving the generalization ability of the decision model.
[0091] The present disclosure also proposes a live broadcast decision determination method, Figure 4 is an implementation flowchart of a live broadcast decision determination method according to an embodiment of the present disclosure, comprising: S410, determining a decision prompt text based on the state of the live broadcast room; S420, determining an operation for the state by using a decision model based on the decision prompt text.
[0092] The decision model can be trained by using the training method mentioned above.
[0093] In this embodiment, the state of the live stream can be converted into structured text information, i.e., decision prompt text. This disclosure can collect multi-dimensional state data (i.e., the state of the live stream) during the live stream. In one example, the state of the live stream can include the characteristics of the live stream during the live stream. This disclosure obtains the state of the live stream by aggregating these characteristics. For example, the characteristics of the live stream can include the current number of online users, the number of orders placed within a fixed time period, the number of comments, product names, and product prices. Furthermore, the live stream state can also include the context information of the script currently being played in the live stream.
[0094] In one example, the decision prompt text may also include at least one of the following: task context, task objectives, and task requirements.
[0095] In this embodiment of the disclosure, the decision model can be a deep learning model with intelligent processing and reasoning capabilities. In one example, the decision model can perform in-depth analysis of the decision prompt text to determine the appropriate action for the live stream status.
[0096] By adopting the above method, the state of the live broadcast room is transformed into decision prompt text, which realizes the expression of multi-dimensional information and makes the information input to the decision model interpretable and context-aware. Furthermore, this disclosure uses the decision model to determine the operation for the current state of the live broadcast room. This operation can reflect the real-time needs during the live broadcast, and thus the live broadcast effect can be improved based on this operation.
[0097] Figure 5 This is a schematic diagram of the process for determining a live broadcast decision according to an embodiment of the present disclosure.
[0098] like Figure 5 As shown, there exists a timeline on which the live streaming time of the digital human anchor is recorded. For example, the live stream starts at time t1 and ends at time t2. The state of the live stream room at time t1 is S0. Based on S0, the corresponding operation A0 can be determined. After executing A0, the reward R0 can be obtained. This reward can be the duration of the live stream watched by the audience (i.e., the viewing time).
[0099] In this timeline, the state of the live broadcast room at time t3 is St. Based on St, the decision operation for the state of the live broadcast room can be determined, including the following steps.
[0100] S501, Issue a decision request.
[0101] In the embodiments of the present disclosure, the decision request can be issued by monitoring the dynamic change of user traffic or setting a timer. In an example, the timer can be set according to a preset time interval or a specific time point (such as 5 minutes before the goods are put on the shelf, or the live broadcast is about to end, etc.), and then the decision request is issued when the time requirement of the timer is met.
[0102] S502, obtain a preset operation list.
[0103] S503, aggregate at least one feature of the live broadcast room to obtain a state of the live broadcast room.
[0104] In some embodiments, further comprising: obtaining at least one feature of the live broadcast room; aggregating the at least one feature of the live broadcast room to obtain a state of the live broadcast room.
[0105] In the embodiments of the present disclosure, at least one feature of the live broadcast room can be obtained from the self-built digital human feature library and the live broadcast room real-time feature library. In an example, the present disclosure can save the features of the live broadcast room generated during the live broadcast into the self-built digital human feature library and the live broadcast room real-time feature library through a data transmission interface.
[0106] In the embodiments of the present disclosure, the features of the live broadcast room can include user behavior features, product attributes, etc., and the specific content can refer to Table 1.
[0107] Further, the present disclosure can aggregate the obtained features of the live broadcast room for processing to generate a structured state of the live broadcast room, in other words, the present disclosure can integrate the scattered features into a unified description that can reflect the comprehensive situation of the live broadcast room. In an example, the present disclosure can use a splicing method to combine numerical features and category features (i.e. non-numerical features) into the state of the live broadcast room.
[0108] In the above manner, by obtaining at least one feature of the live broadcast room, the core variables affecting the live broadcast effect can be covered, and the decision bias caused by information missing can be reduced; further, the scattered features are aggregated into a unified state representation, which improves the storage efficiency and transmission efficiency of the state information, and provides data support for the determination of the subsequent decision prompt text.
[0109] S504, determine a decision prompt text based on the state of the live broadcast room.
[0110] In the embodiments of the present disclosure, the state of the live room and the obtained preset operation list can be fused to determine a decision prompt text. In an example, the decision prompt text can include a task background and a task target, wherein the task target clearly indicates a specific target to be achieved, that is, the operations in the preset operation list are evaluated and sorted according to the state of the live room to determine the optimal operation.
[0111] S505, feature extraction of the decision prompt text is performed by using the decision model, and an operation is determined.
[0112] In some embodiments, based on the decision prompt text, the operation for the state is determined by using the decision model, including: At least one feature of the live room included in the decision prompt text is extracted by using the decision model; The numerical features in the at least one feature of the live room are discretized by using the decision model, and input features are determined based on the discretized numerical features and non-numerical features; Based on the input features, the selection probability of each candidate operation in the preset operation list is determined by using the decision model; The candidate operation with the maximum selection probability is determined as the operation for the state.
[0113] In the embodiments of the present disclosure, the decision model can identify key entities (such as the number of audience, interaction rate, etc.) and their corresponding numerical values in the decision prompt text through natural language understanding technology, and extract non-numerical features (such as product name, product selling point, etc.).
[0114] Further, the decision model of the present disclosure can perform discretization processing on the extracted numerical features, that is, continuous numerical values are divided into discrete intervals, and then the encoding layer or embedding layer of the decision model is used to encode or embed the discretized numerical features to obtain the feature vector of the discretized numerical features.
[0115] The encoding layer or embedding layer of the decision model of the present disclosure can also be used to encode or embed the non-numerical features to obtain the feature vector of the non-numerical features. Further, the feature vector of the discretized numerical features and the feature vector of the non-numerical features are spliced to obtain the input features.
[0116] Based on the determined input features, the decision model of the present disclosure can calculate the selection probability of each candidate operation in the preset operation list. Wherein, the preset operation list can refer to Table 2. In an example, the decision model can calculate the selection probability of each candidate operation according to the association rules between the input features and the operation.
[0117] Further, the disclosure can determine the candidate operation with the maximum selection probability as the operation of the current live room state.
[0118] In the above manner, the live room features are accurately extracted from the decision prompt text by using the decision model, and the numerical features and non-numerical features are fused for input, so that the decision model can dynamically evaluate the matching degree of each candidate operation in the preset operation list with the current state, and output the quantitative selection probability, reducing the subjectivity and hysteresis of manual decision. Further, based on the operation for the state, the live streaming process can be optimized, and the effect of live streaming can be improved.
[0119] In some embodiments, further comprising: performing the operation and determining the reward after performing the operation; based on the live streaming data of the live room, constructing log sample data, the log sample data including a plurality of triplets, each triplet including a state, an operation and a reward having a corresponding relationship; saving the log sample data.
[0120] As shown in Figure 5 The disclosure can feed back the operation At to the digital anchor in the live room, and the digital anchor can perform the operation At. Assuming that the live streaming ends at t4, the disclosure can obtain the reward Rt after performing the operation At. Here, the reward Rt can be the viewing duration of the user watching the live streaming. In an example, t1 and t3 can be considered as the live streaming start time, and t2 and t4 can be considered as the live streaming end time, so that t1-t2 and t3-t4 can be considered as one live streaming cycle respectively.
[0121] Further, the disclosure can organize the state of the live room, the operation determined for the state and the reward after performing the operation, so that each state is matched with the corresponding operation and reward, and a complete triplet is formed, and the multiple triplets after organization are stored as log sample data.
[0122] Finally, the disclosure can write the log sample data into a database or a file system, and store it according to the writing time or the live room identifier.
[0123] In the above manner, by performing the operation corresponding to the state, the reward after performing the operation can be determined according to the business indicators generated after performing the operation, reducing the subjectivity of manual reward evaluation; at the same time, the complete link of each decision is structured as log sample data, forming a reusable decision knowledge base, providing data support for the training of the decision model.
[0124] The disclosure also proposes a training device of a decision model, Figure 6is a structural schematic diagram of a training apparatus 600 of a decision model according to an embodiment of the present disclosure, comprising: A first prompt construction module 610 is configured to construct a decision prompt text based on a state of a live broadcast room. A first operation determination module 620 is configured to determine an operation for the state by using the decision model based on the decision prompt text, and execute the operation. A first reward determination module 630 is configured to determine a reward after the operation is executed; construct log sample data based on live broadcast data of the live broadcast room, the log sample data comprising a plurality of triplets, each of the triplets comprising the state, the operation, and the reward having a corresponding relationship; A training sample determination module 640 is configured to cluster the plurality of triplets based on the state to obtain at least one cluster; and determine a training sample based on the plurality of triplets belonging to the same cluster. A model training module 650 is configured to train the decision model by using the training sample.
[0125] In some embodiments, the training sample determination module 640 is configured to: arrange the plurality of triplets belonging to the same cluster in descending order of the reward to obtain a triplet sequence; determine the training sample based on a preset number of head triplets located at a head of the triplet sequence and a preset number of tail triplets located at a tail of the triplet sequence.
[0126] In some embodiments, the training sample determination module 640 is configured to: replace the operation and the reward in the tail triplets with the operation and the reward in the head triplets to obtain reorganized triplets; determine the training sample based on the reorganized triplets, the head triplets, and other triplets except the tail triplets and the head triplets.
[0127] In some embodiments, the replacement manner comprises sequential cover replacement, reverse cover replacement, or random sampling replacement.
[0128] In some embodiments, the model training module 650 is configured to: input the state contained in the triplet in the training sample into the decision model to determine a predicted operation corresponding to the state; execute the predicted operation and determine an actual reward after the predicted operation is executed; determine a loss function based on at least one of the predicted operation, the actual reward, and the operation and the reward corresponding to the state; train the decision model by using the loss function.
[0129] In some embodiments, the disclosure also provides a training device of a decision model, Figure 7 FIG. 7 is a structural schematic diagram of a training device 700 of a decision model according to an embodiment of the disclosure, comprising: The first state determining module 760 is configured to obtain at least one feature of the live room, and aggregate the at least one feature of the live room to obtain a state of the live room.
[0130] In some embodiments, the first operation determining module 620 is configured to: extract, by using the decision model, at least one feature of the live room contained in the decision prompt text, the at least one feature comprising numerical features and non-numerical features; discretize, by using the decision model, the numerical features, and determine input features based on the non-numerical features and the discretized numerical features; determine, by using the decision model, selection probabilities of each candidate operation in a preset operation list based on the input features, wherein the preset operation list is contained in the decision prompt text; determine, as the operation for the state, the candidate operation with the largest selection probability.
[0131] In some embodiments, the training sample determining module 640 is configured to: determine at least one cluster center; determine distances between each state and each cluster center based on at least one feature of the live room contained in each state; assign the state to a cluster to which the cluster center with the closest distance belongs, to obtain at least one class cluster.
[0132] In some embodiments, the training sample determining module 640 is further configured to: determine a feature mean of each feature based on at least one feature of the live room contained in each state in the class cluster; update the cluster center of the class cluster based on the feature mean of each feature.
[0133] The disclosure also provides a live decision determining device, Figure 8 FIG. 8 is a structural schematic diagram of a live decision determining device 800 according to an embodiment of the disclosure, comprising: The second prompt constructing module 810 is configured to determine a decision prompt text based on a state of a live room; The second operation determining module 820 is configured to determine an operation for the state by using a decision model based on the decision prompt text; The decision model can be trained by using the training device mentioned above.
[0134] In some embodiments, the present disclosure also provides a live decision determination apparatus, Figure 9 FIG. 9 is a structural schematic diagram of a live decision determination apparatus 900 according to an embodiment of the present disclosure, comprising: The second state determination module 930 is configured to acquire at least one feature of the live room, and aggregate the at least one feature of the live room to obtain a state of the live room.
[0135] In some embodiments, the second operation determination module 820 is configured to: extract, by using the decision model, at least one feature of the live room contained in the decision prompt text; discretize, by using the decision model, a numerical feature in the at least one feature of the live room, and determine an input feature based on the discretized numerical feature and a non-numerical feature; determine, by using the decision model, a selection probability of each candidate operation in the preset operation list based on the input feature; determine, as the operation for the state, the candidate operation with the maximum selection probability.
[0136] In some embodiments, the present disclosure further comprises: The second reward determination module 940 is configured to perform the operation and determine a reward after the operation is performed, construct log sample data based on live data of the live room, the log sample data comprising a plurality of triplets, each triplet comprising a state, an operation and a reward having a corresponding relationship, and save the log sample data.
[0137] The specific functions and examples of each module and sub-module of the apparatus according to the embodiments of the present disclosure are described above in the corresponding steps of the method embodiments, and will not be described here.
[0138] In the technical solutions of the present disclosure, the acquisition, storage and application of the user's personal information all comply with relevant laws and regulations and do not violate public order and good customs.
[0139] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0140] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0141] As shown in Figure 10 The device 1000 includes a computing unit 1001 that can perform various appropriate operations and processes in accordance with a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0142] Various components in the device 1000 are connected to the I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; the storage unit 1008, such as a magnetic disk, a magneto-optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0143] The computing unit 1001 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, central processing units (CPUs), graphics processing units (GPUs), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the training method of the decision model and the live decision determination method. For example, in some embodiments, the training method of the decision model and the live decision determination method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded onto the RAM 1003 and executed by the computing unit 1001, one or more steps of the training method of the decision model and the live decision determination method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the training method of the decision model and the live decision determination method by any other suitable means, such as by means of firmware.
[0144] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0145] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0146] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0147] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0148] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0149] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0150] It should be understood that the various forms of flow shown above can be re-ordered, added to, or have steps deleted, using the steps described above. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure can be achieved, which is not limited herein.
[0151] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for training a decision model, comprising: constructing a decision prompt text based on a state of a live room; determining an operation for the state by using a decision model based on the decision prompt text, and performing the operation; determining a reward after performing the operation; constructing log sample data based on live data of the live room, the log sample data comprising a plurality of triplets, each of the triplets comprising the state, the operation and the reward having a corresponding relationship; clustering the plurality of triplets based on the state to obtain at least one cluster; determining a training sample based on a plurality of triplets belonging to a same cluster; training the decision model by using the training sample.
2. The method of claim 1, wherein, The determining of the training sample based on the plurality of triplets belonging to the same cluster comprises: arranging the plurality of triplets belonging to the same cluster in descending order of the reward to obtain a triplet sequence; determining the training sample based on a preset number of head triplets located at a head of the triplet sequence and a preset number of tail triplets located at a tail of the triplet sequence.
3. The method of claim 2, wherein, The determining of the training sample based on the preset number of head triplets located at the head of the triplet sequence and the preset number of tail triplets located at the tail of the triplet sequence comprises: replacing the operation and the reward in the tail triplets with the operation and the reward in the head triplets to obtain reorganized triplets; determining the training sample based on the reorganized triplets, the head triplets, and other triplets except the tail triplets and the head triplets.
4. The method of claim 3, wherein, The replacement manner comprises sequential cover replacement, reverse cover replacement or random sampling replacement.
5. The method of any one of claims 1-4, wherein, The training of the decision model by using the training sample comprises: inputting a state contained in a triplet in the training sample into the decision model to determine a predicted operation corresponding to the state; performing the predicted operation and determining an actual reward after performing the predicted operation; determining a loss function based on at least one of the predicted operation, the actual reward, and the operation and the reward corresponding to the state; training the decision model by using the loss function.
6. The method of claim 1, further comprising: obtaining at least one feature of the live room; aggregating the at least one feature of the live room to obtain the state of the live room.
7. The method of claim 6, wherein, The determining of the operation for the state by using the decision model based on the decision prompt text comprises: extracting, by using the decision model, at least one feature of the live room contained in the decision prompt text, the at least one feature comprising a numerical feature and a non-numerical feature; discretizing, by using the decision model, the numerical feature, and determining an input feature based on the non-numerical feature and the discretized numerical feature; determining, by using the decision model, a selection probability of each candidate operation in a preset operation list based on the input feature; wherein the preset operation list is contained in the decision prompt text; determining the operation for the state as a candidate operation having the maximum selection probability.
8. The method of claim 6, wherein, The clustering of the plurality of triplets based on the states comprises: determining at least one cluster center; determining a distance between each of the states and each of the cluster centers based on at least one feature of the live room included in each of the states; assigning each of the states to a cluster to which the cluster center with the closest distance belongs, to obtain the at least one cluster.
9. The method of claim 8, further comprising: determining a feature mean value of each of the features based on at least one feature of the live room included in each of the states in the cluster; updating the cluster center of the cluster based on the feature mean value of each of the features.
10. A live decision determination method, comprising: determining a decision prompt text based on a state of a live room; determining an operation for the state by using a decision model based on the decision prompt text; wherein the decision model is trained by using the training method of any one of claims 1-9.
11. The method of claim 10, further comprising: obtaining at least one feature of a live room; aggregating the at least one feature of the live room to obtain the state of the live room.
12. The method of claim 11, wherein the determining of the operation for the state by using the decision model based on the decision prompt text comprises: extracting at least one feature of the live room included in the decision prompt text by using the decision model; discretizing a numerical feature in the at least one feature of the live room by using the decision model, and determining an input feature based on the discretized numerical feature and a non-numerical feature; determining a selection probability of each of the candidate operations in a preset operation list based on the input feature by using the decision model; determining the operation for the state as the candidate operation with the largest selection probability.
13. The method of claim 10 or 12, further comprising: executing the operation and determining a reward after the execution of the operation; constructing log sample data based on live data of the live room, the log sample data comprising a plurality of triplets, each of the triplets comprising the state, the operation, and the reward having a corresponding relationship; saving the log sample data.
14. A training device of a decision model, comprising: a first prompt construction module configured to construct a decision prompt text based on a state of a live room; a first operation determination module configured to determine an operation for the state by using a decision model based on the decision prompt text, and execute the operation; a first reward determination module configured to determine a reward after the execution of the operation; constructing log sample data based on live data of the live room, the log sample data comprising a plurality of triplets, each of the triplets comprising the state, the operation, and the reward having a corresponding relationship; a training sample determination module configured to cluster the plurality of triplets based on the states to obtain at least one cluster; determining a training sample based on a plurality of triplets belonging to the same cluster. A model training module configured to train the decision model using the training samples.
15. A live decision determination apparatus, comprising: A second prompt construction module configured to determine a decision prompt text based on a state of a live room; A second operation determination module configured to determine an operation for the state based on the decision prompt text using a decision model; The decision model is trained by the training apparatus of claim 14.
16. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.
17. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-13.
18. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-13.
Citation Information
Patent Citations
Generation method and device for decision network model of vehicle automatic driving
CN107169567A
Taxi scheduling method and system based on deep reinforcement learning
CN111862579A
Multi-task prompt decision converter construction method and device, equipment and storage medium
CN118456423A
Reader emotion prediction method and system based on text emotion behavior knowledge
CN118503349A
Fine tuning method of large language model, live broadcast processing method, device and equipment
CN118555414A