Live script evaluation model training method and device, equipment and storage medium
By using fixed-length text features and multi-source data training samples, combined with embedding layers and attention modules, the problem of randomness in user behavior and degree of interest in the live script evaluation model is solved, achieving more accurate live script evaluation and optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
- Filing Date
- 2025-10-15
- Publication Date
- 2026-07-24
AI Technical Summary
When constructing a live streaming script evaluation model, the randomness and uncertainty of user behavior make it difficult to directly evaluate the quality of live streaming scripts, and the characteristic information of users' interest in live streaming scripts may affect the reliability of the model and the accuracy of the evaluation.
Training samples are constructed using fixed-length text features, and duration-assisted targets are introduced. Training samples and labels are determined by integrating multi-source data. The model is trained using an embedding layer, an attention module, and an evaluation module to reduce interference from interest level and improve evaluation accuracy.
It improves the reliability and accuracy of the live streaming script evaluation model, enabling targeted optimization of live streaming scripts to enhance live streaming quality and conversion rates.
Smart Images

Figure CN121280862B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to the fields of data processing, deep learning, and large models. Background Technology
[0002] In recent years, live streaming with digital humans has shown a booming development trend. As the core guide for live streaming with digital humans, the live streaming script covers information such as complete language expression, content logic, and pacing during the live stream. Building a model to evaluate the live streaming script helps to optimize the live streaming content in a targeted manner, enhance interaction with users, and thus improve the live streaming effect. Summary of the Invention
[0003] This disclosure provides training methods, apparatus, devices, and storage media for live script evaluation models.
[0004] According to one aspect of this disclosure, a method for training a live streaming script evaluation model is provided, comprising: Obtain the live stream script, user dwell time and behavior in the live stream; Based on the live stream script, user dwell time and behavior in the live stream, training samples and corresponding labels are determined. The training samples contain text features, user features and live stream features; among them, text features include script fragments of fixed length. The live streaming script evaluation model was trained using training samples and labels.
[0005] According to another aspect of this disclosure, a training apparatus for a live script evaluation model is provided, comprising: The data acquisition module is used to acquire the live stream script, user dwell time and behavior in the live stream; The sample determination module is used to determine training samples and their corresponding labels based on the live stream script, user dwell time and behavior in the live stream. The training samples contain text features, user features and live stream features; among them, text features include script fragments of fixed length. The training module is used to train the live script evaluation model using training samples and labels.
[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0009] This disclosure determines training samples and labels by integrating multi-source data, providing rich and comprehensive information for training the live streaming script evaluation model. This enables the model to learn the relationship between live streaming scripts and user behavior from multiple perspectives, reducing evaluation bias caused by limited data. Furthermore, determining the labels corresponding to the training samples through the acquired data reflects the effectiveness of the live streaming scripts in practical applications, providing the model with clear learning objectives. This allows for continuous parameter adjustment during training, improving the accuracy and reliability of the model's evaluation of live streaming scripts.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this disclosure; Figure 2 This is a flowchart illustrating the implementation of a training method for a live script evaluation model according to an embodiment of the present disclosure. Figure 3 This is a schematic diagram illustrating the relationship between the live streaming time and the time a user enters the live streaming room according to an embodiment of the present disclosure. Figure 4 This is a schematic diagram illustrating the determination of labels for training samples according to an embodiment of the present disclosure; Figure 5 This is a flowchart of the training process of a live script evaluation model according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of an attention module according to an embodiment of the present disclosure; Figure 7 This is a schematic diagram of the structure of a training device 700 for a live script evaluation model according to an embodiment of the present disclosure; Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] The term "and / or" in this disclosure indicates that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The term "at least one" in this document means any combination of at least two of a plurality of options, such as including at least one of A, B, and C, which can mean including any one or more elements selected from the set of A, B, and C. The terms "first" and "second" in this document refer to and distinguish multiple similar technical terms, and do not imply a specific order or a limitation to only two. For example, "first feature" and "second feature" refer to two types / two features; the first feature can be one or more, and the second feature can also be one or more.
[0014] In recent years, live streaming with digital humans has developed rapidly, and the quality control and optimization of live streaming scripts have become a focus of industry attention. As the core content framework of live streaming with digital humans, the live streaming script comprehensively covers information such as language expression and pacing. Accurately evaluating the live streaming script is a key means to enhance the attractiveness of the live stream, improve user interaction, and increase the conversion rate.
[0015] However, the construction of a live streaming script evaluation model faces several pressing issues. Firstly, in real-world live streaming scenarios, user behavior is highly random and uncertain, making it nearly impossible for users to listen to the entire live streaming script in the training samples. This makes it difficult to directly evaluate the quality of a complete live streaming script based on collected samples. Secondly, the user's dwell time in the live streaming room is a crucial factor that cannot be ignored when constructing the textual features of the live streaming script. If the training samples are directly constructed using the complete live streaming script actually heard by the user, the training samples may provide the live streaming script evaluation model with feature information about the user's level of interest in the script (e.g., the longer the script a user actually hears, the higher their level of interest is likely to be). Consequently, the trained live streaming script evaluation model will become dependent on this feature information, reducing the model's reliability and evaluation accuracy.
[0016] To address the aforementioned issues, embodiments of this disclosure employ fixed-length text features to construct training samples, eliminating feature information indicating user interest in the livestream script. These training samples are then used to train a livestream script evaluation model. Furthermore, embodiments of this disclosure may introduce a duration-assisted objective. Here, the duration-assisted objective, while maintaining a fixed text feature length, can estimate the number of words a user might hear, and this estimation result can be applied to subsequent target tasks related to livestream script evaluation (such as predicting product card click-through rates and conversion rates).
[0017] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this disclosure, such as... Figure 1 As shown in the illustration, the application scenario diagram of this disclosure may include, but is not limited to, a model training device 110 and a live script evaluation model 120. The model training device 110 and the live script evaluation model 120 can communicate via any type of wired or wireless network. Specifically, the model training device 110 can train the model based on training samples to obtain the live script evaluation model 120. Here, the training samples may include multiple live scripts and user behavior data corresponding to each live script. The model training device 110 may include a server for providing backend management for the live script evaluation model 120. Furthermore, this disclosure does not impose a specific limitation on the number of model training devices 110; for example, the application scenario diagram of this disclosure may include one or more model training devices 110.
[0018] Figure 2 This is a flowchart illustrating the implementation of a training method for a live script evaluation model according to an embodiment of the present disclosure, including: S210, Obtain the live stream script, user dwell time and behavior in the live stream; S220. Based on the live stream script, user dwell time and behavior in the live stream, determine the training samples and the corresponding labels for the training samples. The training samples contain text features, user features and live stream features; among them, the text features include script fragments of fixed length. S230. The live script evaluation model is trained using training samples and labels.
[0019] In this embodiment, the live stream script (also known as the live stream plan) is the core planning document for the live stream, recording what the host will say, the product information to be displayed, the interactive segments to be arranged, and other content. It can include aspects such as language expression, content logic, and pacing. The method for obtaining the live stream script will be described in detail later.
[0020] In this embodiment of the disclosure, the time a user spends in the live stream reflects their level of attention and interest in the live stream content. In one example, users with longer dwell times may be more interested in the live stream content, while users with shorter dwell times may be less interested. By collecting data on user dwell time in the live stream, it is possible to clarify the user's acceptance level of different live stream content, which indirectly clarifies the user's interest level in the live stream script contained within the live stream content, providing a reference for evaluating the quality of the live stream script.
[0021] In this embodiment of the disclosure, user behavior in the live stream may include actions such as commenting, liking, sharing, clicking on product cards, and making conversions. These behaviors directly reflect the user's interaction with the live stream content and their level of participation. Collecting user behavior in the live stream helps to evaluate the effectiveness of the live stream script in attracting user interaction and promoting product conversions.
[0022] Furthermore, based on the obtained live streaming script, the user's dwell time and behavior in the live streaming room, training samples for training the live streaming script evaluation model, as well as the labels corresponding to the training samples, can be extracted.
[0023] In this embodiment of the disclosure, the training samples may include text features, user features, and live stream features. In one example, the text features may be extracted from the live stream script and divided into script segments of fixed length. For example, the script segments may include key statements from the live stream process, product introductions, interactive elements, and other core information that reflects the live stream script.
[0024] In one example, user features can be extracted based on the user's dwell time and behavior in the live stream. For example, the user's dwell time, comment frequency, and product conversion history can all be used as user features.
[0025] In one example, live stream characteristics may include the type of live stream, the live stream time, and the streamer's information. Live stream characteristics can reflect the external environment and background factors of the live stream, and will also affect the effectiveness of the live stream script. For example, different types of live streams have different user groups and needs, and the same script will have different effects in different types of live streams.
[0026] The process of determining the training samples and their corresponding labels will be described in detail later.
[0027] Furthermore, in this embodiment, training samples can be used as input data and fed into the live streaming script evaluation model to obtain the evaluation result of the live streaming script evaluation model on the training samples. By comparing the evaluation result with the labels corresponding to the training samples, a loss function is determined, and the model parameters are continuously adjusted according to the loss function, enabling the model to learn the relationship between text features, user features, live streaming features, and labels in the training samples. During the training process, the live streaming script evaluation model can be adjusted to the optimal parameter combination, minimizing the error between the model's evaluation result on the training samples and the labels.
[0028] The live streaming script evaluation model is trained using a large number of training samples and labels, enabling it to assess the quality of live streaming scripts. After training, the model can evaluate the potential performance of a new live streaming script based on the input, such as the click-through rate and conversion rate of product cards (cards displaying product information in e-commerce live streaming scenarios).
[0029] By employing the above method and integrating multi-source data to determine training samples and labels, rich and comprehensive information is provided for training the live streaming script evaluation model. This enables the model to learn the relationship between live streaming scripts and user behavior from multiple perspectives, reducing evaluation bias caused by limited data. Furthermore, determining the labels corresponding to the training samples using the acquired data reflects the effectiveness of the live streaming scripts in practical applications, providing the model with clear learning objectives. This allows for continuous parameter adjustment during training, improving the accuracy and reliability of the model's evaluation of live streaming scripts.
[0030] By evaluating the live stream script, live stream planners can optimize and adjust it in a targeted manner, thereby improving the quality and effectiveness of the live stream, increasing conversion rates, and enhancing its commercial value.
[0031] The following content details the process of determining training samples and labels.
[0032] Figure 3 This is a schematic diagram illustrating the relationship between the live streaming time and the time a user enters the live streaming room, according to an embodiment of this disclosure.
[0033] In the embodiments disclosed herein, such as Figure 3As shown, the live stream script can be rendered in segments (i.e., the complete live stream script is divided into multiple script segments), and these segments are sent to the client. Because the time a user enters and exits the live stream is uncertain, the user may not hear the entire segment. Assume the start and end times of the live stream correspond to times t1 and t6, respectively. A user enters the live stream at time t2 and leaves at time t3; the user re-enters the live stream at time t4 and leaves at time t5. In this case, the user hears portions of multiple live stream script segments.
[0034] In one example, after a script segment is played in the live stream, the script segment can be stored in an external script database. The script database can also store the live stream identifier that played the script segment and the time information when the script segment was played.
[0035] This disclosure can obtain multiple script segments, and sort and aggregate them according to the relationship between the script segments and the live room to obtain the complete live room script.
[0036] In some implementations, obtaining the live stream script includes: Retrieve multiple first scripts, each corresponding to a live stream identifier; The first scripts with the same live stream identifier are sorted and aggregated according to the time sequence of their playback to obtain the live stream scripts.
[0037] In this embodiment of the disclosure, the first script can be a segment of the script that has been saved to the script database. Each first script is associated with a live room identifier (such as roomid), thus binding each first script to a specific live room. For example, first script 1 is associated with live room A, first script 2 is associated with live room B, first script 3 is associated with live room C, first script 4 is associated with live room A, and first script 5 is associated with live room A.
[0038] Furthermore, in this embodiment of the disclosure, multiple first scripts can be grouped according to the live room identifier. For example, if the live room identifiers corresponding to first script 1, first script 4, and first script 5 are all live room A, then these three first scripts can be grouped into one group, and all the first scripts in this group are scripts played in live room A.
[0039] For multiple first scripts in the same live stream, they are sorted in ascending order based on their actual playback time (e.g., the start time of playback), with the first script played earlier appearing first and the last script played later appearing last. Then, the sorted first scripts are aggregated to form the complete script sequence for this live stream, i.e., the live stream script for that live stream.
[0040] By using the above method, based on the unique association of the live room identifier and the time sequence of the first script playback, the first script is sorted and aggregated, which can transform the scattered first script into a logically coherent whole, ensuring the integrity and traceability of the live room script, and thus providing a data foundation for subsequent sample construction.
[0041] In some implementations, training samples are determined based on the live stream script, user dwell time in the live stream, and user behavior, including: Based on the live stream script and the user's dwell time in the live stream, determine the script file that the user listens to in the live stream; Divide the script file into N fixed-length script segments, where N is a positive integer; N training samples are determined, and each training sample corresponds one-to-one with a script fragment. Each training sample contains a script fragment (i.e., text features), user features, and live room features.
[0042] In this embodiment, if a user's stay in the live stream is too short, it may indicate that the user entered the live stream by chance and did not actually pay attention to the live stream content. The data generated by this user may not accurately reflect their reception and response to the live stream script. This disclosure can filter out candidate samples (i.e., the script content played in the live stream during the user's stay) from the live stream script that correspond to a user's stay time in the live stream being greater than or equal to a preset time threshold. In one example, the preset time threshold can be 3 seconds, thus filtering out candidate samples corresponding to a user's stay time in the live stream being greater than or equal to 3 seconds. For example, if user A enters the live stream at 10:00:00 and leaves at 10:00:02, the stay time is only 2 seconds, less than the preset time threshold. Therefore, this record is not used for model training in the live stream script. If user B enters the live stream at 10:05:00 and leaves at 10:10:00, the stay time is 5 minutes, greater than the preset time threshold. Therefore, this record can be used as valid data in the subsequent training sample construction process.
[0043] Live stream scripts typically have a defined time schedule, which may include the start time and duration of each content segment. For example, a live stream script might specify that product introductions will take place from 10:00:00 to 10:05:00, and promotional activities will be explained from 10:05:00 to 10:10:00. This disclosure can construct a timeline based on the script's preset times to clearly define the script content corresponding to each moment.
[0044] Furthermore, for each valid data point corresponding to a user, the script file the user listened to in the live stream can be determined based on the time the user entered the live stream, combined with the timing of the live stream script and the host's speaking speed. In one example, a live stream starts at 10:00:00, and the host explains the script at a speed of 240 words per minute (4 words per second). User C enters the live stream at 10:03:30. Based on the host's speaking speed, approximately 840 words have been explained between the start of the live stream and user C's entry. By querying the live stream script, the specific content of these 840 words can be determined, thus pinpointing the location where user C first heard the script.
[0045] After determining the initial position of the script the user heard, and combining this with the user's position within the live stream, we can determine the complete script file the user listened to. Continuing with the example of user C, let's assume user C stayed in the live stream for 5 minutes. Based on the host's speaking speed, approximately 1200 words of content could be explained within those 5 minutes. Therefore, starting from the previously located script position, selecting the 1200 words corresponding to that position constitutes the script file user C listened to in the live stream.
[0046] Furthermore, this disclosure can divide the determined script file into N fixed-length script segments (N is a positive integer). In one example, this disclosure can use 400 words as the fixed length, which can be determined by statistical data on the general speaking speed of a digital human anchor and the actual listening time of users in the live broadcast sample; or, it can be determined according to the text features related to the product, such as 400 words being sufficient to fully introduce a product under normal circumstances.
[0047] In this embodiment of the disclosure, for each user's script file, 400 words can be selected sequentially from the start time of the script file (i.e., the start time when the user starts listening to the script content) as a script segment.
[0048] Taking the script file containing 1200 characters as an example, based on the fixed length, the script file can be split into 3 script segments (N is 3 in this case). That is, the first script segment includes characters 1-400, the second script segment includes characters 401-800, and the third script segment includes characters 801-1200.
[0049] Furthermore, this disclosure addresses the aforementioned division into N fixed-length script segments, allowing for the determination of corresponding training samples based on each script segment. That is, the number of training samples N is exactly equal to the number of script segments N. For example, if a script file is divided into 3 script segments, then 3 training samples can be generated based on these 3 script segments.
[0050] The training samples can include script fragments, user features, and live stream features. Here, the aforementioned fixed-length script fragments are used as text features in the training samples, reflecting the specific information content transmitted during the live stream; user features can include basic information about the user corresponding to the script fragment (such as age, gender, region, etc.), historical behavioral data (such as the user's dwell time and behavior in the live stream, etc.); live stream features can include the type of live stream, the live stream time, and information about the host in the live stream.
[0051] By employing the above method, the script file listened to by the user is determined based on the live stream script and the user's dwell time, enabling a relatively accurate location of the content the user actually encounters. By dividing the script file into N fixed-length script segments, and then using these segments to determine training samples, the training samples do not reflect the user's level of interest in the live stream script. Using these training samples to train the live stream script evaluation model allows the model to focus on the substantive content of the live stream script, unaffected by user interest levels, thus improving the reliability and accuracy of the live stream script evaluation model.
[0052] In one example, this disclosure can randomly generate N non-repeating starting points in a script file and divide the script file into segments of a fixed length. During this process, this disclosure allows for a certain degree of content overlap or content interval between script segments. In other words, among the multiple script segments divided by this disclosure, there may be a situation where the ending point of the previous script segment is not equal to the starting point of the next script segment.
[0053] In some implementations, the labels corresponding to the training samples are determined based on the live stream script, the user's dwell time and behavior in the live stream, including: Determine the first time step corresponding to the script segment of the training sample; Determine the second moment corresponding to the user's behavior in the live stream; Determine the time interval between the first and second moments; The labels corresponding to the training samples are determined based on the time interval and the user's behavior in the live broadcast room.
[0054] In this embodiment of the disclosure, for a script segment corresponding to a training sample, it is necessary to determine the first moment corresponding to the script segment. Here, the first moment may include the start moment, end moment, or any moment within the time range corresponding to the script segment. In one example, the time range corresponding to a live broadcast segment may be 10:00:00-10:01:40, then the first moment may be 10:00:00, 10:01:40, or any moment within 10:00:00-10:01:40. It is understood that the time range and moment can be represented using common date and time formats or using timestamps.
[0055] This disclosure utilizes the data collection system of a live streaming platform to record user behavior in the live stream and the second moment in which it occurs. In one example, if a user performs an action (such as a click or conversion) in the live stream at 10:03:30, then the second moment corresponding to the user's action in the live stream can be 10:03:30.
[0056] Furthermore, the time interval is obtained by calculating the difference between the second time point and the first time point.
[0057] Figure 4 This is a schematic diagram illustrating the determination of labels for training samples according to an embodiment of the present disclosure.
[0058] like Figure 4 As shown, the time range for the script file listened to by the user in the live broadcast room is 10:00:00-10:05:00 (i.e., the user enters the live broadcast room at 10:00:00 and leaves at 10:05:00). Assuming the host plays the script file at a speed of 4 words per second, 1200 words can be played in 5 minutes. This disclosure uses 400 words as a fixed length, so the above 1200 words can be divided into 3 script segments, such as... Figure 4 The first, second, and third script snippets in the text.
[0059] The first script segment corresponds to a time range of 10:00:00-10:01:40, and this first script segment contains the content from the first character to the 400th character in the script file. The second script segment corresponds to a time range of 10:01:40-10:03:20, and this second script segment contains the content from the 401st character to the 800th character in the script file. The third script segment corresponds to a time range of 10:03:20-10:05:00, and this third script file contains the content from the 801st character to the 1200th character in the script segment. Based on the time ranges corresponding to each script segment, the start time of each segment can be taken as the first moment of each script segment, that is, the first moment of the first script segment is 10:00:00, the first moment of the second script segment is 10:01:40, and the first moment of the third script segment is 10:03:20, with a time interval of 100 seconds between two adjacent first moments (i.e., the time to play a script segment of fixed length).
[0060] like Figure 4 As shown, if a user performs an action in the live stream at 10:03:30, then the second time corresponding to the user's action in the live stream can be 10:03:30. Furthermore, the time interval between the second time interval and the first time interval of the first script segment is 210 seconds, the time interval between the second time interval and the first time interval of the second script segment is 110 seconds, and the time interval between the second time interval and the first time interval of the third script segment is 10 seconds.
[0061] Then, based on the time interval between the second time point and the first time point of each script segment and the user's behavior in the live broadcast room, the labels of the training samples corresponding to each script segment are determined.
[0062] By employing the above method, determining the first moment of the live stream script and the second moment of the user's action within the live stream, and calculating the time interval between the two, the labels corresponding to the training samples are determined based on this time interval and the user's behavior in the live stream. This provides rich and semantically clear training data for the live stream script evaluation model, thereby improving the model's evaluation accuracy and generalization ability. Furthermore, this disclosure can determine diverse labels for the training samples, enabling refined model training and allowing the model to accurately predict live stream scripts for different user groups and behavioral scenarios.
[0063] In some implementations, the labels corresponding to the training samples are determined based on time intervals and user behavior in the live stream, including: Based on user behavior in the live stream, determine the basic labels for the training samples; The attenuation coefficient is determined based on the time interval, and the value of the attenuation coefficient is negatively correlated with the length of the time interval. Based on the base label and decay coefficient, the label corresponding to the training sample is determined.
[0064] In this embodiment of the disclosure, the script segment heard by the user in the live stream at the second moment of the user's behavior in the live stream is determined. For example... Figure 4 As shown, at 10:03:30, the user performed an action in the live stream. At the second moment, the user was listening to the third script segment.
[0065] In this disclosed embodiment, it can be determined Figure 4 In this embodiment of the disclosure, the base label of the training samples corresponding to the first, second, and third script fragments can be set to 1.
[0066] In this embodiment of the disclosure, a tag attribution window can be set, for example, the tag attribution window starts from a predefined start time (such as the time when the user begins to hear the live script) and ends at the time when the target behavior (such as the user's behavior in the live room) actually occurs.
[0067] In one example, if the script file a user listens to in a live stream is divided into multiple script segments, then the training samples determined by the script segments before the user's action in the live stream can all be considered positive samples. In this example, the attenuation coefficient can be determined based on the time interval between the second moment corresponding to the user's action in the live stream and the first moment corresponding to the script segment. Then, based on the base label and attenuation coefficient of each script segment, the label of the training sample corresponding to each script segment can be determined. The value of this attenuation coefficient is negatively correlated with the length of the time interval; that is, the longer the time interval, the smaller the attenuation coefficient (in other words, the greater the attenuation, the smaller the corresponding label).
[0068] In one example, the label can be calculated based on the nonlinear relationship between the base label and the attenuation coefficient, as shown in the following formula: (1) in, The label represents the training sample corresponding to the script fragment; The basic tags that represent script fragments; Indicates the attenuation coefficient; This is a parameter used to calculate the attenuation coefficient, and its value is greater than 0 and less than 1, for example, a value of 0.8; M is another parameter used to calculate the attenuation coefficient. Its value is positively correlated with the time interval between the first and second moments. For example, the formula for calculating M can be: (2) in, This represents the time interval between the first and second moments. Indicates the length of the script segment; This indicates rounding down to the nearest integer.
[0069] like Figure 4 As shown, the time interval between the first moment corresponding to the first script snippet and the second moment corresponding to the user's behavior in the live stream is 210 seconds; the time interval between the first moment corresponding to the second moment corresponding to the user's behavior in the live stream is 110 seconds; and the time interval between the first moment corresponding to the third script snippet and the second moment corresponding to the user's behavior in the live stream is 10 seconds. The length of each script snippet is 100 seconds.
[0070] In this example, using formulas (1) and (2), the attenuation coefficient of the first script segment is calculated as follows: M= Therefore, the attenuation coefficient of the first script segment is (i.e., 0.64). Furthermore, based on the attenuation coefficient of the first script segment and the base label (with a value of 1), the label of the first script segment can be determined to be 0.64.
[0071] Similarly, the label of the second script fragment is calculated to be 0.8, and the label of the third script fragment is 1.
[0072] In another example, the label can be calculated based on the linear relationship between the base label and the attenuation coefficient, and the calculation formula can be: (3) in, The label represents the training sample corresponding to the script fragment; The basic tags that represent script fragments; ( () represents the attenuation coefficient; This is a parameter used to calculate the attenuation coefficient, and its value is greater than 0 and less than 1, for example, a value of 0.1; M is another parameter used to calculate the attenuation coefficient. Its value is positively correlated with the time interval between the first and second moments. For example, it can be calculated using formula (2).
[0073] In this example, using formulas (2) and (3), the attenuation coefficient of the first script segment is: M= Therefore, the attenuation coefficient of the first script segment is 1-2 =0.8. Furthermore, based on the attenuation coefficient of the first script fragment and the base label (with a value of 1), the label of the first script fragment can be determined to be 0.8.
[0074] Similarly, the label of the second script fragment is calculated to be 0.9, and the label of the third script fragment is 1.
[0075] In some implementations, user behavior in the live stream includes at least one of product card clicks and conversions.
[0076] In this embodiment of the disclosure, clicking on a product card can be an action in which a user directly interacts with a product card in a live streaming scenario. This usually manifests as the user actively clicking on a product card displayed in the live streaming room (such as a thumbnail, price tag, promotional logo, etc.), reflecting the user's immediate interest in the product card.
[0077] Product conversion (or simply conversion) generally refers to the behavior of users transforming potential demand for goods into actual value through methods such as payment and placing an order. Typically, in e-commerce digital human live streaming scenarios, conversion behavior can include placing an order and making a payment; essentially, it's the user paying the actual cost for the product.
[0078] In this embodiment, the base label can be determined based on the user's actual behavior in the live stream. For example, if a user engages in behavior while listening to a script file in the live stream, that script file can be considered a positive sample. The script file is divided into N script segments, and the base label for each of these N script segments can be 1. Furthermore, based on the time interval between the first and second moments, a decay coefficient is set for the base label of each script segment. The label of the training sample corresponding to each script segment, determined based on the base label and the decay coefficient, represents the degree of influence of each script segment on the user's behavior in the live stream within the label attribution window.
[0079] By adopting the above method and introducing a decay coefficient that is negatively correlated with the time interval, the labels of the training samples can be dynamically decayed as the time of the behavior occurs. This allows the model to focus on the user's recent active behavior in the live broadcast room based on the labels and training samples, thereby improving the training effect of the live broadcast script evaluation model.
[0080] The following content details the training process of the live script evaluation model.
[0081] Figure 5 This is a flowchart of the training process of a live script evaluation model according to an embodiment of the present disclosure.
[0082] In some implementations, the live streaming script evaluation model includes an embedding layer, an attention module, and a first evaluation module; wherein, The embedding layer is used to determine the feature vectors of text features, user features, and live room features. The attention module is used to weight the feature vectors of text features based on the number of words in the live script heard by the user, and then input the weighted features into the first evaluation module. The first evaluation module is used to evaluate the live streaming script based on the feature vectors of user features, the feature vectors of live streaming room features, and the weighted features.
[0083] Figure 5 The specific architecture of the live streaming script evaluation model is shown, such as... Figure 5 As shown, the live script evaluation model may include an embedding layer 510, an attention module 520, and a first evaluation module 530.
[0084] In this embodiment of the disclosure, the embedding layer 510 may include a live room feature encoding module 511, a user feature encoding module 512, and a text feature encoding module 513.
[0085] In one example, based on a defined training sample, this disclosure can use the live room feature encoding module 511 to determine the feature vector of the live room feature, the user feature encoding module 512 to determine the feature vector of the user feature, and the text feature encoding module 513 to determine the feature vector of the text feature. Here, the live room feature may include basic information about the live room, real-time live data, etc.; the user feature may include user attributes, user behavior information in the live room, etc.; and the text feature may include script fragments from the training sample.
[0086] In this embodiment, the attention module 520 can weight the feature vector of the text features based on the number of words heard by the user to determine the weighted features, and then input the weighted features into the first evaluation module 530. In one example, the attention module 520 can be sequence attention (seq attention). Details regarding the attention module 520 will be provided later.
[0087] In this embodiment of the disclosure, the first evaluation module 530 may include a first transformation block 531 and a tower module 532. The first evaluation module 530 is capable of evaluating script fragments (i.e., part or all of the live script content) of the training samples based on the weighted features input from the attention module 520, as well as the feature vectors of the user features and the feature vectors of the live room features.
[0088] By employing the above approach, the embedding layer unifies the vectorization of text features, user features, and live stream features, reducing the cost of manual feature engineering and improving the accuracy of feature extraction. The attention module weights the feature vectors of text features based on the number of words in the live stream script heard by the user, suppressing noise and reinforcing key content, making the evaluation results more closely reflect the user's actual perception. The first evaluation module integrates the feature vectors of user features, live stream features, and the weighted feature vectors of text features, enabling accurate evaluation of the live stream script and providing efficient and scalable technical support for live stream script optimization.
[0089] In some implementations, training samples and labels are used to train the live script evaluation model, including: Construct a second evaluation module, which has the same structure as the first evaluation module; Based on the training samples, a second evaluation module is used to predict the number of words in the live stream script that the user hears; Based on the number of words in the live stream script heard by the user, the feature vector of the text features, the feature vector of the user features, the feature vector of the live stream features, and the tags, the first evaluation module and the attention module are adjusted.
[0090] like Figure 5 As shown, a second evaluation module 540 is constructed in this embodiment to assist in the training of the live script evaluation model. Here, the second evaluation module 540 has the same structure as the first evaluation module 530. In one example, the second evaluation module 540 may consist of a second conversion block 541 and an auxiliary tower module 542. The second conversion block 541 may have the same neural network structure as the first conversion block 531, and the auxiliary tower module 542 may have the same neural network structure as the main tower module 532. Furthermore, the connection relationship between the second conversion block 541 and the auxiliary tower module 542 is the same as the connection relationship between the first conversion block 531 and the main tower module 532.
[0091] In this embodiment of the disclosure, the main tower module 532 and the auxiliary tower module 542 can constitute a dual-tower model.
[0092] After constructing the second evaluation module 540, the training samples can be used to predict the number of words in the live stream script that users hear. The process of predicting the number of words in the live stream script will be explained in detail later.
[0093] Furthermore, based on the predicted number of words in the live stream script heard by users, the feature vectors of text features, the feature vectors of user features, the feature vectors of live stream features, and the labels corresponding to the training samples, the attention module 520 and the first evaluation module 530 can be adjusted. The adjustment process for the attention module 520 and the first evaluation module 530 will be explained in detail later.
[0094] By constructing a second evaluation module with the same structure as the first evaluation module, the number of words in the live stream script heard by the user can be predicted based on multiple features, providing data support for the subsequent live stream script evaluation model. Furthermore, by jointly predicting the number of words and evaluating the quality of the live stream script, the live stream script evaluation model can utilize the gradient information from the word count prediction to adjust the first evaluation module and the attention module, thereby improving the accuracy of evaluating the live stream script.
[0095] In some implementations, based on training samples, a second evaluation module is used to predict the number of words in the live stream script that the user hears, including: The training samples are input into the embedding layer, which determines the feature vectors of text features, user features, and live room features. The feature vectors of text features, user features, and live room features are then input into the second evaluation module. The second evaluation module is used to predict the number of words in the live stream script that users hear.
[0096] like Figure 5 As shown, this disclosure can input training samples into the embedding layer 510, use the live room feature encoding module 511 of the embedding layer 510 to obtain the feature vector of the live room feature; use the user feature encoding module 512 to obtain the feature vector of the user feature; and use the text feature encoding module 513 to obtain the feature vector of the text feature.
[0097] Further, the feature vectors of text features, user features, and live stream features are input into the second evaluation module. Specifically, this disclosure allows the feature vectors of text features, user features, and live stream features to be input into the auxiliary tower module 542 of the second evaluation module 540, respectively. Furthermore, the feature vectors of text features are input into the second conversion block 541 of the second evaluation module 540, whereby the second conversion block 541 outputs the processing result of the text feature vectors, and then inputs the processing result into the auxiliary tower module 542. Here, the processing result of the text feature vectors by the second conversion block 541 may include context-aware feature vectors. In this embodiment, the second conversion block 541 can capture the dependencies between the various feature vectors in the text feature vectors, and enhance the expressive power of the text feature vectors based on these dependencies.
[0098] Furthermore, the auxiliary tower module 542 can predict the number of words in the live script that the user hears based on the feature vectors of the received text features, the feature vectors of the user features, the feature vectors of the live room features, and the processing results of the feature vectors of the text features by the second conversion block 541.
[0099] By using the above method, the feature vector structure of the training samples is constructed through the embedding layer, which can accurately extract the feature vectors of user features, live room features, and text features. Furthermore, the second evaluation module can predict the number of words in the live script listened to by the user based on these feature vectors, providing data support for the training of the subsequent live script evaluation model.
[0100] In some implementations, the first evaluation module and the attention module are adjusted based on the number of words in the live stream script heard by the user, the feature vector of the text features, the feature vector of the user features, the feature vector of the live stream features, and the tags, including: The feature vectors of user features and live room features are input into the first evaluation module, and the feature vectors of text features are input into the attention module. An attention module is used to weight the feature vectors of the text features based on the number of words in the live script heard by the user, and the weighted features are then input into the first evaluation module. The first evaluation module evaluates the live streaming script based on feature vectors of user characteristics, feature vectors of live streaming room characteristics, and weighted features. The first assessment module and the attention module were adjusted based on the assessment results and labels.
[0101] like Figure 5As shown, the feature vectors of the live room features and the feature vectors of the user features are input into the main tower module 532 of the first evaluation module 530, and the feature vectors of the text features are input into the attention module 520.
[0102] Furthermore, an attention module 520 is employed to weight the feature vector of the text features based on the number of words in the live script heard by the user, as predicted by the second evaluation module 540. In one example, weights are assigned to different segments in the feature vector of the text features according to the number of words in the live script heard by the user (i.e., the feature vector of the text features is weighted) to obtain the weighted features.
[0103] In the embodiments disclosed herein, such as Figure 5 As shown, the weighted features can be input into the first conversion block 531 and the main tower module 532 respectively, and the first conversion block 531 can further process the input weighted features, and then input the processing result of the first conversion block 531 into the main tower module 532.
[0104] Furthermore, the main tower module 532 of the first evaluation module 530 can evaluate the live stream segments (i.e., part or all of the live stream script content) contained in the training samples based on the above input content (i.e., the feature vector of the live stream features, the feature vector of the user features, the feature vector of the text features after weighting, and the result of processing the weighted feature using the first transformation block 531). And based on the evaluation results and the labels corresponding to the training samples, it can adjust the attention module 520 and the first evaluation module 530.
[0105] In one example, this disclosure can calculate an evaluation loss function based on the difference between the evaluation result of the live script and the labels of the training samples corresponding to the live script. The evaluation loss function may include mean squared error, weighted cross-entropy (WCE) loss function, etc. Furthermore, using the evaluation loss function, the gradient is backpropagated to the adjustable parameters of the attention module 520 and the first evaluation module 530 via a backpropagation algorithm, thereby enabling the adjustment of the attention module 520 and the first evaluation module 530.
[0106] By adopting the above approach and introducing an attention mechanism based on the number of words listened to by the user, the feature vectors of text features can be weighted. As a result, the live script evaluation model can focus on the key content of the live script based on the weighted features, thereby improving the training effect of the model.
[0107] In some implementations, training the live script evaluation model further includes: The second evaluation module is adjusted based on the number of words in the live stream script heard by the user and the number of words in the live stream script actually heard by the user.
[0108] In this embodiment of the disclosure, a word count prediction loss function can be determined based on the difference between the number of words in the live script heard by the user and the actual number of words in the live script listened to by the user. The word count prediction loss function may include mean squared error, WCE loss function, etc. Furthermore, using the word count prediction loss function, the gradient is backpropagated to the adjustable parameters of the second evaluation module 540 via a backpropagation algorithm, thereby adjusting the second evaluation module.
[0109] In some implementations, the number of words in the live stream script that a user actually listens to is related to the number of script files the user listens to in the live stream room.
[0110] In this embodiment of the disclosure, the number of words in the live script that the user actually listens to can be determined based on the actual number of words in the script fragments in the training samples. These script fragments can be obtained by dividing the script file that the user hears.
[0111] By using the above method, the second evaluation module is adjusted based on the number of words in the live script predicted by the user and the actual number of words in the live script listened to by the user. This can improve the accuracy of predicting the number of words in the live script listened to by the user. Furthermore, by inputting the number of words in the live script listened to by the user into the attention mechanism, the training effect of the live script evaluation model can be improved.
[0112] Figure 6 This is a schematic diagram of the structure of an attention module according to an embodiment of the present disclosure.
[0113] In some implementations, the attention module includes a multilayer perceptron and a feature weighting module.
[0114] like Figure 6 As shown, the feature vectors of the text features are input into the Multilayer Perceptron (MLP) 610 and the feature weighting module 620, respectively.
[0115] In this embodiment of the disclosure, the MLP 610 first integrates the feature vectors of the text features and the number of words in the live script heard by the user into a joint feature vector through a concatenation operation. Furthermore, the MLP 610 can map the concatenated joint feature vector to a latent space (e.g., 512-dimensional).
[0116] Based on the joint feature vector, MLP 610 generates weight coefficients through independent branching structures. In one example, MLP 610 can use an activation function (such as the sigmoid activation function) to map the number of words in the live script being listened to by the user to a range of 0 to 1, generating a dynamic gating value. This dynamic gating value can suppress text features that do not match the live script being listened to by the current user, thereby reducing the weight coefficients of text features that do not match the live script being listened to by the current user.
[0117] Furthermore, after calculating the weight coefficients of the feature vectors of the text features, the feature weighting module 620 can be used to fuse the weight coefficients with the feature vectors of the text features. In one example, the feature weighting module 620 may include a multiplier that can perform multiplication operations on the weight coefficients and the feature vectors of the text features, that is, multiply each weight coefficient of the feature vector of the text features with the corresponding feature vector, thereby obtaining the weighted features (i.e., the weighted feature vectors of the text features).
[0118] Using the above method, this disclosure fuses the number of words in the live script heard by the user with the feature vector of the text features, enabling the MLP of the attention module to calculate the weight coefficients of the feature vector of the text features. Then, the feature weighting module is used to determine the weighted features, providing data support for training the live script evaluation model and thus improving the model training effect.
[0119] In the data evaluation phase after training the live script evaluation model, this disclosure designs a dual evaluation mechanism based on training samples and new samples.
[0120] The evaluation of the live streaming script evaluation model based on the training samples can include the following steps: First, directly infer the training samples, aggregate the script segments according to the live streaming room identifier, and calculate the average product card click-through rate and conversion rate of each group of training samples; Second, under the same live streaming room dimension, compare the performance differences of different live streaming scripts and generate paired samples for performance analysis.
[0121] For evaluating new samples, two matching strategies can be used: In the time-matching strategy, newly heard live stream scripts within a specific time period are selected. The scaling ratio is calculated based on the length of the original live stream script and the length of the new live stream script. The starting character of the new live stream script is located by combining the character interval of the original live stream script. A new script segment of fixed length is extracted to replace the script segment in the training sample. Evaluation samples are determined based on the new script segment. The evaluation samples are then input into the trained live stream script evaluation model. The trained live stream script evaluation model infers the click-through rate and conversion rate of the evaluation samples and constructs paired samples for effect comparison.
[0122] In the content relevance matching strategy, this disclosure also replaces the training samples based on the new live script. Furthermore, it calculates the similarity between the new script segment and the content actually heard by the user through a sliding window, selects the most relevant window as the starting character, extracts a new script segment of fixed length, and inputs the new script segment into the trained live script evaluation model. The trained live script evaluation model infers the click-through rate and conversion rate of the new script segment, and constructs paired samples for effect comparison.
[0123] This disclosure also proposes a training device for a live script evaluation model. Figure 7 This is a schematic diagram of the structure of a training device 700 for a live script evaluation model according to an embodiment of the present disclosure, comprising: The data acquisition module 710 is used to acquire the live stream script, the user's dwell time and behavior in the live stream; The sample determination module 720 is used to determine training samples and corresponding labels based on the live stream script, the user's dwell time and behavior in the live stream. The training samples include text features, user features and live stream features; among them, the text features include script fragments of fixed length. Training module 730 is used to train the live script evaluation model using training samples and labels.
[0124] In some implementations, the sample determination module 720 is used for: Based on the live stream script and the user's dwell time in the live stream, determine the script file that the user listens to in the live stream; Divide the script file into N fixed-length script segments, where N is a positive integer; N training samples are determined, and each training sample corresponds one-to-one with a script fragment. Each training sample contains a script fragment, user features, and live room features.
[0125] In some implementations, the sample determination module 720 is used for: Determine the first time step corresponding to the script segment of the training sample; Determine the second moment corresponding to the user's behavior in the live stream; Determine the time interval between the first and second moments; The labels corresponding to the training samples are determined based on the time interval and the user's behavior in the live broadcast room.
[0126] In some implementations, the sample determination module 720 is used for: Based on user behavior in the live stream, determine the basic labels for the training samples; The attenuation coefficient is determined based on the time interval, and the value of the attenuation coefficient is negatively correlated with the length of the time interval. Based on the base label and decay coefficient, the label corresponding to the training sample is determined.
[0127] In some implementations, user behavior in the live stream includes at least one of product card clicks and conversions.
[0128] In some implementations, the live streaming script evaluation model includes an embedding layer, an attention module, and a first evaluation module; wherein, The embedding layer is used to determine the feature vectors of text features, user features, and live room features. The attention module is used to weight the feature vector of the text features based on the number of words in the live script heard by the user, and input the weighted features into the first evaluation module. The first evaluation module is used to evaluate the live streaming script based on the feature vectors of user features, the feature vectors of live streaming room features, and the weighted features.
[0129] In some implementations, the training module 730 is used for: Construct a second evaluation module that has the same structure as the first evaluation module; Based on the training samples, a second evaluation module is used to predict the number of words in the live stream script that the user hears; Based on the number of words in the live stream script heard by the user, the feature vector of the text features, the feature vector of the user features, the feature vector of the live stream features, and the tags, the first evaluation module and the attention module are adjusted.
[0130] In some implementations, the training module 730 is used for: The training samples are input into the embedding layer, which determines the feature vectors of text features, user features, and live room features. The feature vectors of text features, user features, and live room features are then input into the second evaluation module. The second evaluation module is used to predict the number of words in the live stream script that users hear.
[0131] In some implementations, the training module 730 is used for: The feature vectors of user features and live room features are input into the first evaluation module, and the feature vectors of text features are input into the attention module. An attention module is used to weight the feature vectors of the text features based on the number of words in the live script heard by the user, and the weighted features are then input into the first evaluation module. The first evaluation module evaluates the live streaming script based on feature vectors of user characteristics, feature vectors of live streaming room characteristics, and weighted features. Adjust the first assessment module and the attention module using the assessment result labels.
[0132] In some implementations, the training module 730 is used for: The second evaluation module is adjusted based on the number of words in the live stream script heard by the user and the number of words in the live stream script actually heard by the user.
[0133] In some implementations, the number of words in the live stream script that a user actually listens to is related to the number of script files the user listens to in the live stream room.
[0134] In some implementations, the data acquisition module 710 is used for: Retrieve multiple first scripts, each corresponding to a live stream identifier; The first scripts with the same live stream identifier are sorted and aggregated according to the time sequence of their playback to obtain the live stream scripts.
[0135] In some implementations, the attention module includes a multilayer perceptron and a feature weighting module.
[0136] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0137] The acquisition, storage, and application of personal information by users involved in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.
[0138] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0139] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0140] like Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0141] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0142] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the training method for a live script evaluation model. For example, in some embodiments, the training method for a live script evaluation model can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the training method for a live script evaluation model described above can be performed. Alternatively, in other embodiments, computing unit 801 may be configured by any other suitable means (e.g., by means of firmware) to execute a training method for evaluating a live script model.
[0143] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0144] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0146] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0147] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0148] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0149] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0150] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A training method for a live streaming script evaluation model, comprising: Obtain the live stream script, user dwell time and behavior in the live stream; Based on the live stream script, the user's dwell time and behavior in the live stream, training samples are determined. These training samples include text features, user features, and live stream features. The text features include script fragments of fixed length. A first time point corresponding to the script fragment of the training sample is determined. A second time point corresponding to the user's behavior in the live stream is determined. The time interval between the first and second time points is determined. Based on the user's behavior in the live stream, basic labels for the training samples are determined. A decay coefficient is determined based on the time interval, the value of which is negatively correlated with the length of the time interval. Based on the basic labels and the decay coefficient, labels corresponding to the training samples are determined. The live streaming script evaluation model is trained using the training samples and labels.
2. The method according to claim 1, wherein, The process of determining training samples based on the live stream script, the user's dwell time and behavior in the live stream, includes: Based on the live stream script and the user's dwell time in the live stream, determine the script file that the user is listening to in the live stream. The script file is divided into N fixed-length script segments, where N is a positive integer; N training samples are determined, and each of the N training samples corresponds one-to-one with N script fragments. Each training sample contains one of the script fragments, the user features, and the live broadcast room features.
3. The method according to claim 1 or 2, wherein, The user's behavior in the live stream includes at least one of product card clicks and conversions.
4. The method according to claim 1 or 2, wherein, The live streaming script evaluation model includes an embedding layer, an attention module, and a first evaluation module; wherein... The embedding layer is used to determine the feature vectors of the text features, the user features, and the live streaming room features; The attention module is used to weight the feature vector of the text feature based on the number of words in the live script heard by the user, and input the weighted feature into the first evaluation module. The first evaluation module is used to evaluate the live script based on the feature vector of the user features, the feature vector of the live room features, and the weighted features.
5. The method according to claim 4, wherein, The step of training the live script evaluation model using the training samples and labels includes: Construct a second evaluation module, which has the same structure as the first evaluation module; Based on the training samples, the second evaluation module is used to predict the number of words in the live stream script that the user hears. The first evaluation module and the attention module are adjusted based on the number of words in the live script heard by the user, the feature vector of the text feature, the feature vector of the user feature, the feature vector of the live room feature, and the tag.
6. The method according to claim 5, wherein, The step of predicting the number of words in the live stream script heard by the user, based on the training samples and using the second evaluation module, includes: The training samples are input into the embedding layer, which determines the feature vectors of the text features, the user features, and the live room features. The feature vectors of the text features, the user features, and the live room features are then input into the second evaluation module. The second evaluation module is used to predict the number of words in the live stream script heard by the user.
7. The method according to claim 5, wherein, The adjustment of the first evaluation module and the attention module based on the number of words in the live stream script heard by the user, the feature vector of the text features, the feature vector of the user features, the feature vector of the live stream features, and the tags includes: The feature vectors of the user features and the feature vectors of the live room features are input into the first evaluation module, and the feature vectors of the text features are input into the attention module; The attention module performs weighted processing on the feature vector of the text features based on the number of words in the live script heard by the user, and inputs the weighted features into the first evaluation module. The live script is evaluated using the feature vectors of the user features, the feature vectors of the live room features, and the weighted features from the first evaluation module. The first evaluation module and the attention module are adjusted based on the evaluation results and the labels.
8. The method according to claim 5, wherein, Training the live script evaluation model further includes: The second evaluation module is adjusted by comparing the number of words in the live stream script heard by the user with the actual number of words in the live stream script listened to by the user.
9. The method according to claim 8, wherein, The number of words in the live stream script that the user actually listens to is related to the script file that the user listens to in the live stream room.
10. The method according to claim 1 or 2, wherein, The process of obtaining the live stream script includes: Obtain multiple first scripts, each corresponding to a live stream identifier; The first scripts with the same live stream identifier are sorted and aggregated according to the time sequence of their playback to obtain the live stream scripts.
11. The method according to claim 4, wherein, The attention module includes a multilayer perceptron and a feature weighting module.
12. A training device for a live streaming script evaluation model, comprising: The data acquisition module is used to acquire the live stream script, user dwell time and behavior in the live stream; A sample determination module is used to determine training samples based on the live stream script, the user's dwell time and behavior in the live stream, and the training samples containing text features, user features, and live stream features; wherein, the text features include script fragments of fixed length; determine the first moment corresponding to the script fragment of the training sample; determine the second moment corresponding to the user's behavior in the live stream; determine the time interval between the first moment and the second moment; determine the basic label of the training sample based on the user's behavior in the live stream; determine a decay coefficient according to the time interval, the value of which is negatively correlated with the length of the time interval; and determine the label corresponding to the training sample based on the basic label and the decay coefficient. The training module is used to train the live script evaluation model using the training samples and labels.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.
Citation Information
Patent Citations
Model training and quality evaluation method and device, electronic equipment and storage medium
CN115099350A
Model training method and device, live broadcast script generation method and device and electronic equipment
CN117313715A