Live broadcast data processing method and device, computer readable medium and program product
By performing multimodal time-series alignment and fusion on live broadcast data, identifying live broadcast topics and user group behavior characteristics, the one-sidedness of single-modal analysis is solved, and a comprehensive and accurate evaluation of live broadcast content and user behavior is achieved.
Patent Information
- Application Number
- CN202510941104.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-10
AI Technical Summary
Existing live streaming platforms mainly focus on single-modal data in live streaming content analysis, which leads to one-sided analysis results and is unable to fully portray live streaming content and user behavior.
By aligning the time series of multimodal live broadcast data, integrating image, text and voice data, using neural network models to identify live broadcast topics and user group behavior characteristics, and combining the health determination model, accurate data basis is provided.
It achieves a comprehensive portrayal of live broadcast content and user behavior, provides accurate assessment of the health of live broadcast rooms, and improves the accuracy and comprehensiveness of analysis.
Smart Images

Figure CN120769074A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a live broadcast data processing method, device, computer-readable medium, and program product. Background Art
[0002] In recent years, live streaming platforms have become an important channel for content creators to interact with their audiences. Mainstream live streaming platforms primarily analyze live content by focusing on simple statistics of single-modal data, such as voice keywords and bullet comment keywords.
[0003] This section is intended to provide a background or context to the embodiments of the present application that are recited in the claims. Nothing herein is admitted to be prior art by virtue of its inclusion in this section. Summary of the Invention
[0004] Multiple aspects of the present application provide a live broadcast data processing method, device, computer-readable storage medium and program product, which are used to combine live broadcast data of multiple modalities to accurately and comprehensively characterize the live broadcast content and various types of users participating in the live broadcast process, and provide accurate and comprehensive data basis for downstream tasks.
[0005] In one aspect of the present application, a live broadcast data processing method is provided, comprising: performing time alignment on live broadcast data of multiple modes within a time period to obtain multimodal alignment data; determining a live broadcast topic within the time period based on the multimodal alignment data; determining group behavior characteristic data of at least one type of user participating in the live broadcast topic based on the multimodal alignment data; and determining the health of the live broadcast process based on the group behavior characteristic data under the live broadcast topic of each of the multiple time periods during the live broadcast process.
[0006] Another aspect of the present application provides a live broadcast data processing device, including: a data alignment unit, configured to perform time alignment on live broadcast data of multiple modes within a time period to obtain multimodal alignment data; a topic determination unit, configured to determine the live broadcast topic within the time period based on the multimodal alignment data; the data determination unit, configured to determine the number of group behavior characteristics of at least one type of user participating in the live broadcast topic based on the multimodal alignment data; and a health determination unit, configured to determine the health of the live broadcast process based on the group behavior characteristic data under the live broadcast topic of each of the multiple time periods during the live broadcast process.
[0007] Another aspect of the present application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the live data processing method shown above.
[0008] In another aspect of the present application, a computer readable storage medium is provided, which stores computer program instructions executable by a processor to implement the live data processing method shown above.
[0009] In another aspect of the present application, a computer program product is provided, which comprises a computer program executable by a processor to implement the live data processing method shown above.
[0010] In the scheme provided by the embodiments of the present application, the multi-modal alignment data is obtained by time sequence alignment of the live data of multiple modalities in a time period; the live topic in the time period is determined according to the multi-modal alignment data; the group behavior feature data of at least one type of user participating in the live topic is determined according to the multi-modal alignment data; and the health degree of the live process is determined according to the group behavior feature data under the live topic of each time period in the live process, so as to accurately and comprehensively depict the live content and each type of user participating in the live process by combining the live data of multiple modalities, to provide accurate and comprehensive data basis for downstream tasks, and to accurately determine the health degree of the live room. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0012] Other features, objects and advantages of the present application will become more apparent through reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings: Figure 1 Flowchart of a live data processing method provided by an embodiment of the present application; Figure 2 Flowchart of a live data processing method provided by another embodiment of the present application; Figure 3 Structure diagram of a live data processing device provided by an embodiment of the present application; Figure 4 Structure diagram of a device suitable for implementing the scheme in the embodiments of the present application The same or similar reference signs in the drawings represent the same or similar components. DETAILED DESCRIPTION
[0013] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0014] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPU, Central Processing Unit), input / output interfaces, network interfaces and memory.
[0015] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0016] Computer-readable media include permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. Information can be computer program instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0017] The embodiment of the application provides a live data processing method, which comprises the following steps: performing time sequence alignment on live data of multiple modes in a time period to obtain multi-modal alignment data; determining a live topic in the time period according to the multi-modal alignment data; determining group behavior feature data of at least one type of user participating in the live topic according to the multi-modal alignment data; and determining the health degree of a live process according to the group behavior feature data of the live topic in each time period in the live process, so as to accurately and comprehensively depict the live content and each type of user participating in the live process by combining the live data of multiple modes, to provide accurate and comprehensive data basis for downstream tasks, and to accurately determine the health degree of the live room.
[0018] In actual scenarios, the execution subject of the method can be a user device, or a device formed by integrating a user device and a network device through a network, or an application program running on the device. The user device includes, but is not limited to, computers, smart phones, tablet computers, and various terminal devices. The network device includes, but is not limited to, network hosts, single network servers, multiple network server sets, or computer sets based on cloud computing. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing, wherein cloud computing is a kind of distributed computing, and a virtual computer composed of a group of loosely coupled computer sets.
[0019] It can be understood that in the specific embodiments of the application, the collection, use or processing of data is involved, and when the embodiments of the application are applied to specific products or technical implementations, the permission or consent of the data subject needs to be obtained, and the collection, use or processing of related data needs to comply with relevant laws, regulations and standards of the data source and implementation location, and through desensitization technology, it is ensured that the finally used is the desensitized data after safe processing, and the rights and interests and data security of the data subject are protected.
[0020] Figure 1 A processing flow 100 of a live data processing method provided by the embodiment of the application is shown, and the live data processing method at least includes the following processing steps: Step 101, performing time sequence alignment on live data of multiple modes in a time period to obtain multi-modal alignment data.
[0021] In the embodiment, the execution subject can obtain the live data of multiple modes in the time period from a remote place or a local place through a wired network connection mode or a wireless network connection mode, and perform time sequence alignment on the live data of multiple modes in the time period to obtain multi-modal alignment data.
[0022] A live broadcast process can be a complete live broadcast by a live broadcast user (anchor) or a portion of a complete live broadcast. By dividing the live broadcast process according to a preset duration, multiple time periods can be obtained. The preset duration can be flexibly set based on actual circumstances (e.g., the live broadcast scenario and the data processing capabilities of the executing entity) and is not limited here.
[0023] Live streaming data in multiple modalities includes, but is not limited to, image (or video) data, voice data, and text data. Image data primarily includes the live streaming room's display interface data, voice data primarily includes the voice data of the live streaming user and some viewers, and text data primarily includes comments, comments, likes, and other data from viewers.
[0024] As an example, timing alignment can be performed based on the respective timestamps of live broadcast data of multiple modalities. First, during the data acquisition phase, ensure that the acquisition equipment or software for image data, text data, and voice data can provide accurate timestamp records. For example, using professional live broadcast recording equipment or software, they can usually capture video and audio at the same time and add accurate timestamps to each frame of image, each audio clip, and text message (such as barrage, comments). Then, based on the timestamps, the data of different modalities are arranged and matched in chronological order. For image data, text data, and voice data at the same time point or in a similar time interval, they are considered to correspond to each other, thereby achieving timing alignment.
[0025] As another example, key frames and key events can be used to achieve temporal alignment between live broadcast data of multiple modalities. First, for image data, key frames in the video are extracted using a key frame extraction algorithm (such as those based on inter-frame differences, motion detection, etc.). At the same time, key events are annotated for text and voice data, such as identifying semantic turning points in text, specific keywords or semantic units in voice. Then, data from different modalities are matched and aligned using key frames and key events as anchors. For example, the time point corresponding to the image key frame is found, and then the key event closest to that time point is found in the text and voice data, and they are associated to achieve alignment.
[0026] Combining live streaming data from multiple modalities, such as images, text, and audio, can address the one-sidedness of single-modal data and complement each other's shortcomings. For example, text data cannot capture the host's expressions and movements, but image data can compensate for this deficiency by analyzing the host's facial expressions, gestures, and other body language, providing a more comprehensive understanding of the live broadcast content. Image data cannot directly convey the emotional and semantic information of language, while text and audio data can provide details such as the verbal communication, tone, and emotion between the host and the audience. After fusion, the various modalities complement each other, making the data processing results more accurate and complete. Especially for complex live broadcast scenarios, multimodal data fusion can integrate multiple information and avoid the one-sidedness of single-modal analysis.
[0027] Furthermore, live streaming data in multiple modalities, such as images, text, and audio, can experience timing mismatches and real-time conflicts due to factors like network latency and varying update frequencies. Accurately aligning the timing of multimodal data can accurately reflect audience feedback and improve the accuracy of live content analysis. This is crucial for subsequent live topic identification and analysis of group behavior characteristics.
[0028] Step 102: Determine the live broadcast topic within the time period based on the multimodal alignment data.
[0029] In this embodiment, a live broadcast topic refers to the core content of the live broadcast user or audience user's explanation, discussion, presentation, or interaction around a specific content, theme, or field during the live broadcast. It is the main thread throughout the live broadcast process during the time period, guiding the content direction of the live broadcast and attracting audience attention and participation.
[0030] For example, in a beauty livestream, the topic is "A must-have light, breathable foundation tutorial for summer." The host showcases foundations and concealers from various brands, explains how to choose the right product based on skin type, demonstrates application steps and techniques, and interacts with viewers, answering questions about foundation. This topic clearly identifies the focus of the livestream on summer foundation, making it highly targeted and practical, attracting viewers with relevant needs. In a knowledge-sharing livestream, the topic is "The Application and Prospects of Artificial Intelligence in Healthcare." The host invites experts in related fields to discuss how AI can assist with medical diagnosis and treatment, sharing the latest research findings and case studies and analyzing future development trends. This focus on specialized knowledge is appealing to those interested in AI and healthcare, sparking in-depth discussion and reflection.
[0031] As an example, first, feature extraction and fusion are performed on the multimodal aligned data to obtain fused features. The fused features are then input into a topic recognition model, which outputs the live broadcast topic. The topic recognition model characterizes the correspondence between the fused features and the live broadcast topic. For example, a neural network model, such as a recurrent neural network model or a residual network model, is trained using a supervised machine learning algorithm.
[0032] The fusion operation can be performed at different times. For example, the timing of feature fusion can be: Early fusion: At the data level, features from different modalities are concatenated to form a unified feature vector, which is then fed into a single model for processing. For example, the feature vectors of images, audio, and text can be directly concatenated and fed into a fully connected neural network.
[0033] Intermediate fusion: Interaction and fusion of features from different modalities in the model's intermediate layers. Architectures such as multimodal transformers can be used to allow information interaction and fusion between features from different modalities in the model's hidden layers.
[0034] Late fusion: Data from different modalities are analyzed and features extracted independently to obtain their respective feature representations, which are then fused at the decision-making level.
[0035] In order to further improve the effectiveness and accuracy of multimodal feature fusion, the above execution entity can also perform feature extraction and fusion processes in the following ways: First, modal feature extraction and preliminary quantification are performed on the aligned data of each modality in the multimodal aligned data.
[0036] For image data, we use deep convolutional neural networks (CNNs) to extract multiple layers of features, including appearance features (such as the livestreamer's clothing, the shape and color of props), and motion features (such as the optical flow characteristics of the livestreamer's gestures and facial expressions). These features are then quantized into vector representations, such as feature vectors extracted using a ResNet (Residual Network) model.
[0037] For audio data, MFCC (Mel-Frequency Cepstrum) and speech emotion feature extraction tools (such as OpenSMILE) are used to extract the acoustic and emotional features of the audio, such as pitch, timbre, speaking rate, volume, and emotional tendency (positive, negative, or neutral), and also convert them into vector form.
[0038] For text data, we use NLP (Natural Language Processing) techniques, such as pre-trained Transformer-based language models, to extract semantic feature vectors from the text. For viewer comments, we analyze their sentiment, keywords (such as "like," "great," "bad," etc.), and topics (such as inquiries about product prices and comments about product quality). An example of a pre-trained language model is BERT (Bidirectional Encoder Representations from Transformers).
[0039] Then, for each modal feature, a dynamic modal weight is calculated. A modal weight calculation module based on the attention mechanism is established. This module analyzes the modal features in a fixed time window (for example, every 5 seconds).
[0040] Within the time window, the correlation between each modal feature and other modal features is calculated, as well as the correlation between each modal feature and the current analysis task (such as interaction analysis). For example, for interaction analysis tasks, mutual information can be used to measure the correlation between image features (such as a live broadcast user's smile) and text features (such as positive emojis in the comments of viewer users), as well as their correlation with the viewer user's like behavior.
[0041] Based on the correlation calculation results, the weight of each modality in the fusion process is dynamically adjusted. The weight value ranges from 0 to 1, and the sum of all modal weights is 1. When a modality contributes significantly to the analysis task within the current time window, its weight is increased accordingly. For example, when a live streamer showcases a product and explains its features in detail, the weight of audio and text features may be increased, because the audience's comments and the live streamer's explanation are more critical to understanding the interaction intent.
[0042] Finally, optimize the feature fusion process based on downstream tasks. Determine the specific objective function for the analysis task. For example, in a sentiment analysis task, the objective function might be the accuracy of predicting the viewer's emotional polarity; in a retention prediction task, the objective function might be the accuracy of predicting whether a viewer will continue to watch the live broadcast.
[0043] The dynamically weighted multimodal features are input into a fusion model, which can be an MLP (Multilayer Perceptron) or LSTM (Long Short-Term Memory). Using optimization algorithms such as gradient descent, the parameters of the fusion model are adjusted according to the objective function to ensure that the fused features can better serve the specific analysis task.
[0044] During training, the parameters of the modal weight calculation module are continuously updated to better adapt to the changes in modal importance across different tasks and at different times. For example, when training a sentiment analysis model, the parameters of the attention mechanism module are also adjusted to improve the ability to capture the modal features of sentiment-related information.
[0045] Specifically, the modal weight calculation module based on the attention mechanism can be implemented as follows: First, feature vectors extracted from various modalities (image, audio, text) within different time windows are used as input. These feature vectors can be viewed as elements in a sequence, similar to how words are mapped to vectors to obtain an input sequence when processing natural language text in a Transformer. However, the sequence elements here are multimodal features.
[0046] Next, multiple attention heads are created, each with a set of learnable parameters (query matrix, key matrix, and value matrix). Different heads have different parameters, allowing them to capture feature relationships from different perspectives. For each multimodal feature sequence, a query matrix, key matrix, and value matrix are constructed. For example, for a sequence of image feature vectors, a linear transformation (through a fully connected layer) is performed to obtain the query matrix, key matrix, and value matrix corresponding to the image modality. Similarly, the same operation is performed for audio and text feature vector sequences to obtain their corresponding query matrix, key matrix, and value matrix.
[0047] Then, within each attention head, a dot product operation is performed between the query matrix and the key matrix to calculate the similarity between the feature elements of each modality. This is similar to measuring the relevance of different modal features to the target task within the current time window. For example, for an interactive analysis task, calculating the similarity between image and text feature elements within a certain time window can be seen as determining whether there is a correlation between the host's expression in the image and the audience's reaction in the text comments, thereby assessing their potential joint contribution to the interactive analysis task.
[0048] Then, the calculated similarity value is divided by the square root of the key matrix dimension for scaling to prevent the gradient from vanishing or exploding. The softmax function is then applied to the scaled value to obtain the attention weight distribution of each modal feature under the current attention head. The sum of these weight values is 1, which reflects the proportion of each modal feature in the fusion focus under the perspective of this head.
[0049] Finally, the attention weights calculated by multiple attention heads are applied to the corresponding value matrix to obtain the weighted feature representation of each head. These weighted feature representations are then concatenated or averaged to obtain a comprehensive multimodal feature representation. This process integrates the inter-modal relationships and key information captured by each attention head from different angles, providing more comprehensive and targeted feature input for subsequent task-oriented optimization fusion.
[0050] As another example, first, for the alignment data of each modality in the multimodal alignment data, semantic understanding is performed based on the alignment data of the modality to determine the sub-topic recognition result corresponding to the modality; then, based on the temporal relationship and voice correlation between the multimodal alignment data, the sub-topic recognition results corresponding to each of the multiple modalities are combined to determine the live broadcast topic.
[0051] The number of live broadcast topics within the target time period can be one or more. When there is one live broadcast topic, it can be the dominant topic within the time period. When there are multiple live broadcast topics, they can include, for example, at least one live broadcast user topic initiated by the live broadcast user and at least one audience user topic initiated by the audience user.
[0052] In some optional implementations of this embodiment, the execution entity may perform step 102 as follows: The first step is to determine the popularity of multiple candidate topics within a time period based on multimodal alignment data.
[0053] As an example, first, through the above-mentioned feature extraction and fusion methods, fused features of the multimodal aligned data are obtained. Then, the fused features are input into a popularity calculation model, which outputs the popularity of multiple candidate topics within a time period. The popularity calculation model is used to characterize the correspondence between the fused features of the multimodal aligned data and the popularity of multiple candidate topics within a time period. This can be achieved by training a neural network model using a machine learning algorithm.
[0054] As another example, first, based on the multimodal alignment data, a preset calculation indicator related to each candidate topic is determined; then, for each candidate topic among the multiple candidate topics, the popularity of the candidate topic within the time period is determined in combination with the multiple preset calculation indicators corresponding to the candidate topic. The multiple preset calculation indicators corresponding to a candidate topic may include, for example, the number of comments and comments related to the candidate topic, the number of users participating in the candidate topic, the emotional intensity of users participating in the candidate topic, and the proportion of users who follow the candidate topic among all users participating in the candidate topic.
[0055] The multiple candidate topics may be all topics in a candidate topic set, including all topics that may be involved in a live broadcast scenario. To improve the efficiency of determining candidate topics, corresponding candidate topic sets may be set for different types of live broadcast users or different types of live broadcast rooms, including topics that may be involved in the corresponding types of live broadcast users or live broadcast rooms.
[0056] The second step is to determine the live broadcast topic from multiple candidate topics based on popularity.
[0057] First, multiple candidate topics are sorted in descending order of popularity to obtain a topic sequence; then, a live broadcast topic is determined from the candidate topic sequence.
[0058] For example, a preset number of candidate topics ranked first may be used as live broadcast topics; for another example, candidate topics whose popularity is greater than a preset popularity threshold may be used as live broadcast topics.
[0059] In this implementation, based on the popularity of multiple candidate topics within a time period, a live broadcast topic is determined from multiple candidate topics. Under the restriction of multiple candidate topics, it helps to improve the accuracy and standardization of the determination of the live broadcast topic.
[0060] In some optional implementations of this embodiment, before executing the first step, the execution subject may also determine multiple candidate topics in the following manner: First, the feature data of the multimodal alignment data in multiple time periods during the live broadcast process are clustered to obtain multiple topic clusters; then, multiple candidate topics corresponding to the multiple topic clusters are determined.
[0061] As an example, after obtaining the fused features of multimodal aligned data using the above method, first select an appropriate clustering algorithm, such as K-Means. Next, determine the number of clusters, K, using either the elbow rule (calculating the sum of squared clustering errors for different K values and selecting the K value at the inflection point of the error) or the silhouette coefficient method (evaluating the clustering effect and selecting the K value with the largest silhouette coefficient). The fused features are then fed into the clustering algorithm, which clusters the fused features corresponding to multiple time periods during the live broadcast to generate multiple topic clusters. The data within each cluster has similar characteristics and represents a specific topic.
[0062] Finally, for each topic cluster, analyze some of the features of the multimodal aligned data it contains. For example, find out the image elements (such as pictures of a certain product), text keywords (such as "game guide" and "promotion") and key words in speech (such as "this product" and "precautions") that appear frequently in the cluster. Count the frequency and importance of these features within the cluster. Based on the results of the feature analysis within the cluster, potential candidate topics are mined. Text mining algorithms (such as TF-IDF to extract keywords and then generate topic phrases through semantic combination) or manual labeling can be used to correspond each topic cluster to a specific candidate topic. For example, if the features of a topic cluster are concentrated on related words and image elements such as "electronic products", "parameter introduction" and "user experience", it can be named the candidate topic "Electronic Product Evaluation and Recommendation".
[0063] In some implementations, the emotional tendencies of live broadcast users and audience users when participating in live broadcast topics can also be determined based on multimodal alignment data, and candidate topics can be generated by combining the emotional tendency data. As an example, for the partial alignment data of each modality in a topic cluster, the feature data representing the emotions is extracted, and the emotion recognition sub-result corresponding to the modality is determined based on the feature data; the emotional tendency sub-results corresponding to each of the multiple modalities are weighted and fused to obtain the emotional tendencies of the users under the candidate topics corresponding to the topic cluster. The emotional tendencies obtained by comprehensive calculation are annotated on the candidate topics to intuitively display the user's emotional reactions triggered by the live broadcast topic.
[0064] For a live broadcast process, all topics involved therein are taken as candidate topics to determine the live broadcast topic for each time period during the live broadcast process.
[0065] In this implementation, the feature data of the multimodal aligned data in multiple time periods during the live broadcast are clustered to obtain multiple topic clusters, and then multiple candidate topics are obtained, which further narrows the scope of candidate topics and helps to further improve the accuracy and standardization of the determination of live broadcast topics.
[0066] Step 103: Determine group behavior characteristic data of at least one type of users participating in the live broadcast topic based on the multimodal alignment data.
[0067] In this embodiment, user types include, but are not limited to, temporary viewers, first-time visitors, and following users among live broadcast viewers. Temporary viewers are generally those who enter the live broadcast room temporarily and have not established a strong connection with the live broadcast user or the live broadcast room. They are generally referred to as "passers-by"; first-time visitors are those who enter the live broadcast room or the live broadcast platform for the first time. They are generally referred to as "new users"; following users are viewers who follow the live broadcast user and have a deep understanding of the live broadcast user and the live broadcast content. They are generally referred to as "fans."
[0068] Group behavior characteristic data is a comprehensive collection of data describing the behavior and characteristics of audience groups during live broadcasts. This data analyzes the behavioral characteristics of different types of users when viewing live broadcasts, including viewing duration, viewing frequency, and interactive behavior. For example, following users may have a longer average viewing duration and a higher viewing frequency, watching live broadcasts frequently each week. Casual viewers may have shorter viewing durations and lower viewing frequency, perhaps only occasionally visiting the live broadcast room. First-time users have a viewing duration and frequency that fall somewhere in between, indicating they are in the exploratory phase.
[0069] In this embodiment, other group-related data of at least one type of user can also be determined. Group-related data, for example, is a comprehensive description of the multi-dimensional characteristics of the user type. It comprehensively presents the overall characteristics and behavior patterns of the group through quantification and classification. For example, in addition to group behavior characteristic data, group-related data also covers data characteristics of the basic information, device and platform preferences, interest preferences, social relationships, and feedback evaluation of the type of users participating in the live broadcast topic, which is used to accurately portray the common characteristics of the user group.
[0070] As an example, first, extract the user ID (Identity document) from the bullet comments and comments in the text data, call the user data system, and obtain user-related data for each user. User-related data includes, for example, whether the user is a follower, whether it is a first-time visitor, account level, and information such as the bullet comments and comments posted by the user. Based on the user-related data, determine the user type label. Then, associate all bullet comments and comments participating in the live broadcast topic (obtained by semantic clustering) with the user type label, and determine the users included in at least one type of user. Finally, for each type of user, the user-related data of all users under that type of user is counted to obtain group-related data.
[0071] In terms of interest preferences, keyword extraction and topic model analysis from text data can be used to determine the interests of different types of users in specific subtopics within a livestream topic. For example, in a livestream topic featuring multiple beauty product introductions, following users may be more interested in high-end brands, while casual viewers may be more engaged in discussions about affordable products. First-time users may also be more interested in explanations of entry-level products. Furthermore, combined with image and audio feature analysis, it can be seen that following users tend to be more interested in images showcasing high-end brand products and are more active in audio interactions (such as asking questions) when listening to the livestreamer explain high-end brand products.
[0072] It should be noted that the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals. For example, before collecting information from a user, a data collection request is issued to the user, and the information collection operation is performed with the user's authorization.
[0073] In some optional implementations of this embodiment, the execution entity may perform step 103 as follows: The first step is to determine the group participation data of at least one type of user based on the multimodal alignment data, wherein the group participation data represents the participation of the type of user in the live broadcast topic.
[0074] The group participation data for a certain type of user includes, but is not limited to, the percentage of users of that type participating in the live broadcast topic, the density of comments made by that type of user under the live broadcast topic, the length of their stay, and other data. Comment density represents the number of comments per unit time.
[0075] For multimodal aligned data, especially text data, statistics are performed on multiple dimensions such as the number of at least one type of users, the number of comments, and the length of stay, to determine the group engagement data of at least one type of user.
[0076] The second step is to determine the difference data between the group participation data of at least one type of user and the historical group participation data corresponding to the historical live broadcast process.
[0077] For each data in the group participation data of each type of user, determine the difference between the data of this type of user and the historical data corresponding to this data during the historical live broadcast process, and combine the differences corresponding to each data under this type of user to obtain difference data.
[0078] The historical live broadcast process may be, for example, a plurality of live broadcasts preceding the current live broadcast. The historical group participation data of a type of user may be the average of the group participation data corresponding to the type of user in the plurality of historical live broadcasts.
[0079] Taking the percentage of following users as an example, the percentage of following users in the current live broadcast process is subtracted from the average percentage of following users in the historical live broadcast process to obtain the difference data corresponding to the percentage of following users.
[0080] The third step is to obtain group behavior feature data of at least one type of user by combining the group participation data and difference data of the type of user.
[0081] The group behavior feature data of a type of user mainly includes two types of data, one of which is group engagement data of the type of user, and the other is difference data of the type of user.
[0082] In the implementation, a specific determination manner of the group behavior feature data is provided, the richness and comprehensiveness of the group behavior feature data are improved in combination with the group engagement data and the difference data, and the accuracy of a downstream task based on the group behavior feature data is further improved.
[0083] In step 104, the health degree of the live broadcast process is determined according to the group behavior feature data of each live broadcast topic in each time period in the live broadcast process.
[0084] The health degree of the live broadcast process refers to the degree of good running state of the live broadcast activity reflected by various elements in the live broadcast process, and is a quantitative result obtained by monitoring, evaluating and analyzing a plurality of key indicators related to the live broadcast.
[0085] As an example, the health degree of the live broadcast process is determined by inputting the live broadcast topics in each time period in the live broadcast process and the group behavior feature data under the live broadcast topics into a health degree determination model. The health degree determination model represents the corresponding relationship between the live broadcast topics in the live broadcast process, the group behavior feature data under the live broadcast topics and the health degree of the live broadcast process, and can be obtained by training a neural network model such as a recurrent neural network or a long short-term memory model using a machine learning algorithm.
[0086] As another example, the health degree of the live broadcast process is determined by inputting the live broadcast topics in each time period in the live broadcast process and the group behavior feature data under the live broadcast topics into a large language model based on the powerful natural language understanding and logical analysis capabilities of the large language model. In order to further improve the determination result of the health degree, the large language model can be fine-tuned in the live broadcast scene to adapt the determination task of the health degree of the live broadcast process.
[0087] In some optional implementations of the embodiment, the above-mentioned execution subject can also perform the following operation: determining the guidance degree of at least one type of user to the live broadcast topic according to the group engagement data and the difference data of each type of user.
[0088] For each type of user, the guidance degree of the type of user to the live broadcast topic is determined according to the group engagement data and the difference data of the type of user.
[0089] As an example, the execution subject can input the group participation data and the difference data of each type of user as prompt data into the large language model. Based on the powerful natural language understanding capability and logical reasoning capability, the large language model determines the guidance degree of each type of user to the live topic.
[0090] As another example, the guidance score representing the guidance degree can be determined by using the following calculation formula: Guidance score = W1 Proportion data of type user + W2 Speech density of type user + W3 Stay duration of type user - W4 Difference data.
[0091] Wherein, W1, W2, W3, W4 respectively represent the weight of the proportion data of the type user, the speech density of the type user, the stay duration of the type user, and the difference data.
[0092] The guidance degree can represent the dominant type of the live topic, which generally includes: Focus user dominant type: high proportion of focus users, stable speech, and positive content; Temporary access user dominant type: temporary audience users surge, topic controversial, and low participation of focus users; Suspected manipulation type: first-time access users speak collectively, with radical emotions and screen brushing behavior; Mixed type: all types of users participate, and the dominant group is not obvious.
[0093] In the implementation mode, the execution subject can determine the group behavior characteristic data of the type user by combining the group participation data, the difference data, and the guidance degree of the type user.
[0094] For each type of user, the group behavior characteristic data of the type of user is determined by combining the group participation data, the difference data, and the guidance degree of the type of user.
[0095] In the implementation mode, the specific determination method of the group behavior characteristic data is provided, which further improves the richness and comprehensiveness of the group behavior characteristic data by combining the group participation data, the difference data, and the guidance degree, and helps to further improve the accuracy of the downstream task based on the group behavior characteristic data.
[0096] In some optional implementation modes of the embodiment, the execution subject can further perform the following operations: The first step is to use preset anomaly determination rules to determine abnormal time periods and the anomaly types corresponding to the abnormal time periods from multiple time periods based on the live broadcast topics and group behavior characteristic data corresponding to each of the multiple time periods during the live broadcast process.
[0097] Based on pre-set anomaly detection rules, we analyze multiple time periods during the live broadcast. For each time period, we examine multiple indicators from the behavioral characteristics of various user groups participating in the live broadcast topic during that time period, including their proportion of all users, comment density, and duration of stay. Furthermore, we examine the degree of variation across these indicators. Based on these data characteristics, we determine whether an anomaly occurred during that time period and, if so, what type of anomaly it was.
[0098] For example, if the comment density of a certain type of user drops sharply within a certain time period, and there is a significant negative deviation from the historical average, and the length of stay is also significantly shortened, then this period can be determined to be an abnormal time period, and its abnormal type can be determined to be "a sudden drop in user interaction."
[0099] For example, if the number of comments from temporary visiting users increases suddenly (for example, the growth rate exceeds a preset speed threshold), and the repetition of the comment content is higher than the preset repetition threshold, then the abnormal type corresponding to the abnormal time period can be determined to be "online hired opinion manipulators flooding the screen"; if the output content of live broadcast users or audience users includes sensitive keywords, and the participation of followed users is lower than the preset participation threshold, then the abnormal type corresponding to the abnormal time period can be determined to be "induced promotion"; if the number of users' likes and comments increases but the number of interactive users grows slowly, then the abnormal type corresponding to the abnormal time period can be determined to be "false interaction".
[0100] Preset exception determination rules can be set based on the live broadcast platform's historical exception situations to cover various exception types encountered in the past. The classification of exception types can be flexibly set according to the needs of the live broadcast platform itself.
[0101] In the second step, according to the abnormality type, target data is determined from the group behavior characteristic data corresponding to the abnormal time period for clustering to obtain abnormal clustering results.
[0102] After identifying the anomalous time period, for each anomaly type, we precisely extract target data representing the anomaly type from the corresponding group behavior characteristic data. For example, if the anomaly type is "a sudden drop in user interaction," the target data includes the proportion of each user type, comment density, duration of stay, and the difference from the historical mean during that time period. This target data is then aggregated into sample vectors, and clustering analysis is performed on these sample vectors using a clustering algorithm, such as K-Means or a hierarchical clustering algorithm. Clustering algorithms group similar sample vectors together based on inherent data similarities. This algorithmic calculation yields several clusters, each representing a group of samples with similar data characteristics from the anomalous time period, forming the anomaly clustering results.
[0103] The third step is to determine the abnormal pattern of the live broadcast process based on the abnormal clustering results.
[0104] We can further analyze the clustering results and explore the data characteristic patterns shared by each cluster during the abnormal time period. For example, if we find that within a cluster, the abnormal time period shows a consistent trend of a decrease in the proportion of a certain type of user, a decrease in comment density, and a decrease in duration of stay, we can define this common characteristic as an abnormal pattern and name it "Specific User Group Interaction Decline Pattern."
[0105] For example, in a certain cluster, if it is found that the semantic feature content is sparse, the interactive content density is low, and the fan participation is low, the abnormal pattern is marked as hanging broadcast.
[0106] By analyzing each cluster one by one, we can sort out the various abnormal patterns that may occur during the live broadcast, accurately identify and classify live broadcast abnormalities, and assist in the subsequent live broadcast optimization and adjustment strategy formulation.
[0107] In this implementation, a method for analyzing abnormal patterns in a live broadcast process is provided, which combines two progressive abnormality determination methods, namely, preset abnormality judgment rules and clustering, to improve the accuracy of abnormal patterns.
[0108] In some optional implementations of this embodiment, the execution subject may further perform the following operation: determining the correlation between the multiple abnormal time periods based on the live broadcast topics and group behavior characteristic data corresponding to the multiple abnormal time periods.
[0109] As an example, first, we build a correlation analysis model. Specifically, we convert the live broadcast topic and group behavior feature data for each abnormal time period into a transactional data format suitable for the association rule mining algorithm, with each feature as an item. We then set support and confidence thresholds based on the specific data analysis requirements. Support indicates the frequency of an association rule in the dataset, while confidence indicates its reliability. We then run the association rule mining algorithm to identify association rules that meet the support and confidence thresholds. These rules form the correlation analysis model, revealing the characteristic associations between different abnormal time periods.
[0110] Then, we conduct correlation analysis based on time series. Specifically, we plot a time series graph of the abnormal time periods during the live broadcast to visually demonstrate the distribution and order of these abnormal time periods. We calculate the time intervals between adjacent abnormal time periods and analyze whether there are specific time interval patterns. For example, we examine whether certain abnormality types recur at regular intervals. We also construct a Markov chain model to describe the transition probabilities between abnormal time periods. We treat each abnormal time period as a state and calculate the probability of transitioning from one abnormal time period to another, thereby revealing the dynamic correlation between abnormal time periods.
[0111] Finally, by combining the results of association rule mining and time series analysis, we identify correlation patterns between abnormal time periods. Specifically, we identify common correlation patterns. For example, periods of abnormal "sudden drops in user engagement" are often followed by periods of abnormal "new user churn." We then interpret these identified correlation patterns and analyze their possible causes and mechanisms. For example, a sudden drop in user engagement could cause new users to lose interest in live streaming content, leading to "new user churn."
[0112] In this implementation, the execution subject may perform the third step to determine the abnormal pattern of the live broadcast process in the following manner: combining the abnormal clustering results and the correlation to determine the abnormal pattern of the live broadcast process.
[0113] As an example, the execution entity can combine the anomaly patterns obtained through clustering with the correlations between time periods obtained through correlation analysis to modify and improve the anomaly patterns. For example, if a cluster in the anomaly clustering results contains abnormal time periods that often experience a sudden drop in engagement, but correlation analysis reveals that these time periods often occur after abnormal time periods with significantly different live content, these two anomaly types can be combined into a new anomaly pattern, representing an abnormal pattern where a sudden drop in engagement occurs due to significantly different content.
[0114] In this implementation, based on the abnormal clustering results, combined with the correlation between abnormal time periods, the abnormal pattern of the live broadcast process is determined, which further improves the accuracy of the abnormal pattern.
[0115] In some optional implementations of this embodiment, the above-mentioned execution subject may further perform the following operation: determining an operation strategy for the live broadcast user during the live broadcast process according to the abnormal pattern.
[0116] Identified abnormal patterns are categorized and organized to form an abnormal pattern library. The characteristics of each abnormal pattern are recorded in detail, including the relevant abnormality type, its performance in different live broadcast topics, the types of users involved, and behavioral changes. For example, the "engagement drop - content deviation - audience loss" pattern is characterized by a decrease in engagement and shortened duration of stay among various user types within a specific live broadcast topic due to content deviation from the main theme, ultimately leading to a significant loss of viewers.
[0117] We thoroughly analyze the root causes of each abnormal pattern, focusing on multiple factors, including live broadcast content, host performance, and technical issues. For example, the "Technical Failure - Viewer Loss" pattern occurs when technical issues such as freezes and blurry images during the live broadcast lead to a poor viewer experience and subsequent viewer loss. The "Competitive Product Attraction - Viewer Diversion" pattern occurs when competing live broadcasts, with more compelling content, attract viewers to the current broadcast room.
[0118] Based on the causes and characteristics of the abnormal patterns, formulate corresponding operation strategies for each abnormal pattern. For example, the operation strategies include the following: 1. Content Optimization Strategy To address the "content deviation-engagement drop" pattern: Strengthen livestream training to ensure they stick to the main theme and plan content structure in advance. For example, for beauty livestreams, introduce products sequentially according to skincare steps to maintain content coherence. Set keyword reminders to prompt livestreamers to return to the main point if they stray from the main theme.
[0119] Addressing the "outdated content – audience churn" model: Regularly update livestream topics and incorporate new elements and perspectives. For example, for technology livestreams, include cutting-edge product experiences and expert interviews to enhance content novelty.
[0120] 2. Interaction Enhancement Strategy
[0121] Addressing the "single interactive form - low engagement" model: Diversify interactive forms, incorporating raffles, Q&A sessions, voting, etc. For example, include a Q&A session during a knowledge livestream where viewers can answer questions and win rewards, thus increasing engagement.
[0122] To address the "delayed feedback and cold atmosphere" model: deploy a professional director team to shorten the barrage feedback time; add real-time interactive topics to guide audience discussion. For example, set up a "match result prediction" topic in sports live broadcasts to enhance the timeliness of interaction.
[0123] 3. User retention strategy
[0124] To address the "new user churn - low familiarity" model: Optimize the onboarding process for new users, highlighting the welcome section and introducing livestream highlights and benefits. For example, for game livestreams, a guide page will be displayed for new users, explaining the game rules, benefits, and providing dedicated personnel to answer any questions.
[0125] To address the "churn of old users – shifting interests" model: establish a user interest feedback mechanism, regularly collect feedback, and adjust content accordingly. For example, for live broadcasts on mothers and babies, add child education psychology lectures based on user feedback, broaden the content scope, and rekindle the interest of old users.
[0126] It should be noted that the operation strategy may also include a penalty strategy for live broadcast users, such as using manual review to determine the live broadcast user's live broadcast authority before the live broadcast user's next live broadcast, and demoting the live broadcast user.
[0127] In this implementation, the operational strategy for live broadcast users can be determined in a targeted manner based on abnormal patterns, which can assist in the identification and prevention of abnormal patterns and help improve the live broadcast effect and the stickiness of audience users.
[0128] In some optional implementations of this embodiment, the above-mentioned execution entity can perform the above-mentioned step 104 in the following manner: first, based on the group behavior characteristic data under the live broadcast topics of each of the multiple time periods during the live broadcast process, determine the first evaluation indicator, the second evaluation indicator, the third evaluation indicator and the fourth evaluation indicator, wherein the first evaluation indicator represents the number of newly added followers of the live broadcast user when the increase rate of temporary audience users during the live broadcast exceeds the preset growth rate threshold, the second evaluation indicator represents the comparison result between the proportion data of the followers and the historical proportion average data, the third evaluation indicator represents the interaction time of the first-time visiting user, and the fourth evaluation indicator represents the concentration degree of text data in the multimodal alignment data on the preset sensitive topic.
[0129] The health of the live streaming process is then determined by combining the first, second, third, and fourth evaluation indicators. Live streaming health is a comprehensive indicator derived from a quantitative assessment of multiple key indicators during the live streaming process, including overall quality, user engagement, interactive effects, and content security. It is used to comprehensively measure the health and sustainability of the live streaming process.
[0130] As an example, first, the above-mentioned evaluation indicators are calculated as follows.
[0131] 1. The first evaluation indicator: Calculate the rate of change in the number of casual viewers within each time period. This is done by subtracting the number of casual viewers from the previous time period from the current time period, then dividing the result by the previous time period's number of casual viewers to arrive at the growth rate. This growth rate is compared with a preset growth rate threshold. If it exceeds the threshold, count the number of newly added followers during this time period and use this as the value of the first evaluation indicator. The preset growth rate threshold can be set based on historical live broadcast data or industry experience, for example, 20%.
[0132] Assume that during a certain time period of a live broadcast, the number of temporary audience users increased by 30% compared with the previous time period, exceeding the preset growth rate threshold of 20%. The number of newly added followers during this time period is 500, then the first evaluation indicator is 500.
[0133] 2. Second evaluation indicator: For each time period, calculate the percentage of following users, which is the number of following users divided by the total number of users in the live broadcast room. Collect the percentage of following users in each time period of multiple historical live broadcasts and calculate the average as the historical percentage average data.
[0134] Compare the percentage of users following the live broadcast in each time period with the historical average percentage data, and calculate the difference between the two as the second evaluation indicator. For example, if the historical average percentage is 40%, and the percentage of users following the live broadcast in a certain time period is 45%, the second evaluation indicator is +5%; if the current percentage is 35%, the second evaluation indicator is -5%.
[0135] 3. The third evaluation indicator: Identify first-time visiting users in each time period and record their interaction time in the live broadcast room, including the time from entering the live broadcast room to leaving, as well as the length of time they stay in different links.
[0136] Calculate the average interaction duration of first-time users as the third evaluation metric. For example, if there are 100 first-time users in a certain time period and their total interaction duration is 2000 minutes, then the average interaction duration is 20 minutes, which means the third evaluation metric is 20 minutes.
[0137] 4. The fourth evaluation indicator: Analyze the text data in the multimodal alignment data, which may include text content such as bullet comments and comments. Use natural language processing technology to identify words and sentences involving pre-set sensitive topics.
[0138] Count the proportion of text data related to sensitive topics in all text data. For example, if 10% of the content in the barrage and comments in a certain time period involves preset sensitive topics, the fourth evaluation indicator is 10%.
[0139] Then, we assign weights to each evaluation metric based on its importance to the health of the live broadcast. For example, the first evaluation metric has a weight of 30%, the second has a weight of 25%, the third has a weight of 20%, and the fourth has a weight of 25%.
[0140] Normalize the values of each evaluation indicator so that they are within the same dimensional range, such as between 0 and 1. For example, for the first evaluation indicator, you can set a maximum and minimum value and linearly map it to the range of 0-1.
[0141] Calculate the weighted sum: multiply the standardized values of each evaluation indicator by its corresponding weight, and then add them together to get the overall score of live broadcast health. For example, if the first evaluation indicator is standardized to 0.8, the second evaluation indicator is standardized to 0.6, the third evaluation indicator is standardized to 0.5, and the fourth evaluation indicator is standardized to 0.4, the overall score is 0.8×0.3 +0.6×0.25 + 0.5×0.2 + 0.4×0.25 = 0.55.
[0142] The health level is set based on the comprehensive score, such as 0.8-1 for excellent, 0.6-0.8 for good, 0.4-0.6 for fair, and 0-0.4 for poor. This determines the health level of the current live broadcast process. For example, if the score is 0.55, it is considered to be at a fair health level.
[0143] In this implementation, a specific calculation method for the health of the live broadcast process is provided, which combines the evaluation indicators of various dimensions to improve the accuracy of the determined health and the comprehensiveness of the characterization of the live broadcast process.
[0144] In some optional implementations of this embodiment, the above-mentioned execution subject may further perform the following operation: determining an operation strategy for the live broadcast user during the live broadcast process according to the health level.
[0145] First, determine the health range to which the health degree belongs and determine the health type; then, based on the health type, determine the operation strategy for the live broadcast user during the live broadcast process.
[0146] As an example, 1. When the health level is excellent (0.8-1), the following operation strategy is adopted: 1.1 User Maintenance Strategy: Feedback for Followers: Regularly launch exclusive promotions for those who follow us. For example, hold a monthly raffle for those who follow us, with prizes like high-quality merchandise related to the livestream, such as autographed posters or customized livestream prop models. During the livestream, provide followers with more opportunities for interaction, such as dedicated Q&A sessions, collecting their questions in advance and answering them during the livestream, enhancing their sense of participation and belonging.
[0147] Community Building: Encourage users to establish their own communities or fan groups. Organize various topic discussions and fan creation activities within the community, such as fan painting competitions and copywriting competitions. With users as the core, a positive community atmosphere is formed to further enhance their loyalty and activity.
[0148] 1.2 New User Development Strategy: Word-of-mouth marketing: Leverage the high satisfaction of engaged users to spread word-of-mouth. During the livestream, encourage followers to share the link to their social circles and offer incentives, such as limited-time membership privileges or exclusive virtual gifts. Furthermore, collaborate with bloggers and opinion leaders in related fields to encourage them to recommend the livestream, thereby expanding its influence and visibility and attracting more potential new users.
[0149] High-quality content sharing: Create a collection of highlights from your livestream and publish them on major video platforms and social media platforms. These clips can include funny moments, exciting performances, valuable knowledge sharing, and more, attracting new users to watch the entire livestream. Include a schedule and summary of the livestream to encourage new users to follow and participate.
[0150] 2. When the live broadcast health is good (0.6-0.8), adopt the following operation strategy: 2.1 User stickiness improvement strategy: Content optimization: Fine-tune livestream content. By analyzing user feedback and interaction data, understand their interest in different types of content. For example, if you find that users are highly engaged with a specific gameplay segment in a certain game livestream, increase the weight of that segment and, before the livestream, use pre-release announcements and social media notifications to encourage users to watch on time.
[0151] Personalized recommendations: We provide personalized content recommendations for different users based on their viewing history and interaction preferences. For new users, we recommend other live streaming rooms or related content they may be interested in based on their first-time live streaming behavior data, such as the live streaming sections they spent most time on and the types of hosts they interacted with. This helps increase user engagement and retention on the platform.
[0152] 2.2 User conversion strategy: Follow-up: During live streams, add prompts and incentives to convert casual viewers into followers. For example, during key moments in the live stream, such as explaining key knowledge points or showcasing exciting products, prompts will pop up to remind viewers to follow the livestream for more benefits and subsequent updates. Also, set up follow-up rewards, such as small gifts like virtual props or coupons, to attract casual viewers.
[0153] Interactive event design: Organize various interactive activities, such as guessing games, voting, and mini-games, to encourage participation from casual viewers and new users. Establish a reward mechanism within the event, giving participants the opportunity to win generous prizes, such as physical goods and platform virtual currency, to increase their enthusiasm for participation. During the event, guide them to pay attention to the live broadcast room so that they can continue to participate in similar activities.
[0154] 3. When the live broadcast health is average (0.4-0.6), adopt the following operation strategy: 3.1 User Loss Recovery Strategy: Lost User Analysis: Conduct a detailed analysis of recently lost followers to understand the reasons for their churn. Gather information through questionnaires and user feedback channels, such as whether the livestream content didn't meet expectations, the livestream timing wasn't appropriate, or there was competition with other platforms. Based on the analysis, develop targeted strategies for retaining them.
[0155] Exclusive Retrieval Campaigns: Design exclusive retrieval campaigns for lapsed users. For example, send personalized invitation emails or text messages informing them of new changes to the livestream content and new promotions. Offer retrieval rewards, such as exclusive membership upgrades or limited-edition virtual gifts, to entice lapsed users back to the livestream.
[0156] 3.2 Content and interaction optimization strategy: Content structure adjustment: Comprehensively organize and optimize live broadcast content. Reduce content sections that are inconsistent with user interests or are of low quality, and increase popular, valuable, and engaging content. At the same time, rationally schedule live broadcasts to ensure they align with the active hours of the target user group, increasing user viewing time and engagement.
[0157] Increase the frequency of interaction: Increase the frequency and quality of interaction with users. During live broadcasts, hosts should proactively engage with viewers, responding to comments and barrages to build positive relationships. Set up multiple interactive sessions, such as a raffle or Q&A every half hour, to keep viewers engaged and focused, reducing the possibility of user churn.
[0158] 4. When the live broadcast health is poor (0-0.4), the following operation strategy is adopted: 4.1 Emergency Problem Solving Strategy: Content Safety Rectification: If the fourth assessment indicator performs poorly and involves a high number of sensitive topics, the livestream content will be subject to immediate and rigorous review and rectification. Training for streamers and users will be strengthened, with clearer content standards and sensitive topics defined on the platform. A real-time monitoring mechanism will be established to promptly address any sensitive topics identified to avoid negative impacts and user loss.
[0159] User Complaint Handling: We prioritize user complaints and negative feedback, handling and resolving them promptly. We establish a dedicated user complaint handling team to record, categorize, and address each complaint in detail, providing users with a satisfactory response within the specified timeframe. Through a proactive approach and effective solutions, we aim to restore user trust and satisfaction and mitigate the spread of negative word-of-mouth.
[0160] 4.2 Comprehensive optimization strategy: Live Streaming Team Adjustments: Evaluate and adjust the live streaming team, including anchors, planners, technicians, and other positions. If any issues are found in team members' work ability or sense of responsibility, timely training or personnel adjustments will be implemented to ensure the team can collaborate efficiently and provide users with better services and content.
[0161] Reposition the direction of live streaming: Reposition the direction of live streaming based on market research and user data analysis. This may require adjusting the theme, style, target audience, and other aspects of the live streaming to adapt to market changes and user needs. For example, if the original live streaming content was sharing expertise in a niche field but user engagement was low, consider expanding to related but more popular areas, or changing the sharing method to make it more interesting and practical to attract more user attention and participation.
[0162] It should be noted that the operation strategy may also include a penalty strategy for live broadcast users, such as using manual review to determine the live broadcast user's live broadcast authority before the live broadcast user's next live broadcast, and demoting the live broadcast user.
[0163] In this implementation, the operational strategy for live broadcast users is determined based on the health of the live broadcast process, which can accurately improve user stickiness, retention and conversion, optimize content quality and security, enhance interactive effects, adapt to changes in user needs, and promote the long-term and stable development of the live broadcast room.
[0164] Figure 2 The following is a processing flow of a live data processing method provided by an embodiment of the present application. The live data processing method includes at least the following processing steps: Step 201: Time-series alignment is performed on live broadcast data of multiple modes within a time period to obtain multi-modal aligned data.
[0165] Step 202: Determine the popularity of multiple candidate topics within a time period based on the multimodal alignment data.
[0166] Step 203: Determine a live broadcast topic from multiple candidate topics based on popularity.
[0167] Step 204 : Determine group participation data of at least one type of user based on the multimodal alignment data.
[0168] Among them, group participation data represents the participation of different types of users in live broadcast topics.
[0169] Step 205 : Determine difference data between the group participation data of at least one type of user and the historical group participation data corresponding to the historical live broadcast process.
[0170] Step 206 : Determine the degree of guidance of the live broadcast topic by each of the at least one type of users based on the group participation data and difference data of each of the at least one type of users.
[0171] Step 207 : Determine group behavior characteristic data of type users by combining the group participation data, difference data, and guidance degree of type users.
[0172] Step 208 , using preset abnormality determination rules, according to the live broadcast topics and group behavior characteristic data corresponding to the multiple time periods during the live broadcast process, determines the abnormal time period and the abnormality type corresponding to the abnormal time period from the multiple time periods.
[0173] Step 209 : According to the abnormality type, target data is determined from the group behavior characteristic data corresponding to the abnormal time period for clustering to obtain abnormal clustering results.
[0174] Step 210 : determining the correlation between the multiple abnormal time periods based on the live broadcast topics and group behavior characteristic data corresponding to the multiple abnormal time periods.
[0175] Step 211: Determine the abnormal pattern of the live broadcast process based on the abnormal clustering result.
[0176] Step 212: Determine an operation strategy for the live broadcast user during the live broadcast process based on the abnormal pattern.
[0177] It can be seen from this embodiment that Figure 1Compared with the corresponding embodiments, the live broadcast data processing method in this embodiment specifically explains the process of determining the live broadcast topic, the process of determining the group behavior characteristic data, the process of determining the abnormal pattern, and the process of determining the operation strategy, which improves the accuracy of the determined abnormal pattern and thereby improves the matching degree between the operation strategy and the live broadcast users.
[0178] In addition, the embodiment of the present application also provides a live data processing device, the structure of which is as follows: Figure 3 shown.
[0179] A live broadcast data processing device 300 includes: a data alignment unit 301, configured to perform time alignment on live broadcast data of multiple modes within a time period to obtain multimodal alignment data; a topic determination unit 302, configured to determine the live broadcast topic within the time period based on the multimodal alignment data; a data determination unit 303, configured to determine group behavior characteristic data of at least one type of user participating in the live broadcast topic based on the multimodal alignment data; and a health determination unit 304, configured to determine the health of the live broadcast process based on the group behavior characteristic data under the live broadcast topic of each of the multiple time periods during the live broadcast process.
[0180] In some optional implementations of this embodiment, the topic determination unit 302 is further configured to: determine the popularity of multiple candidate topics within a time period based on the multimodal alignment data; and determine a live broadcast topic from the multiple candidate topics based on the popularity.
[0181] In some optional implementations of this embodiment, the above-mentioned device also includes: a topic generation unit (not shown in the figure), which is configured to: cluster the feature data of the multimodal alignment data in multiple time periods during the live broadcast process to obtain multiple topic clusters; and determine multiple candidate topics corresponding one-to-one to the multiple topic clusters.
[0182] In some optional implementations of this embodiment, the data determination unit 303 is further configured to: determine the group participation data of at least one type of user based on the multimodal alignment data, wherein the group participation data characterizes the participation of the type of user in the live broadcast topic; determine the difference data between the group participation data of at least one type of user and the historical group participation data corresponding to the historical live broadcast process; for at least one type of user, combine the group participation data and the difference data of the type of user to obtain the group behavior characteristic data of the type of user.
[0183] In some optional implementations of this embodiment, the above-mentioned device also includes: a degree determination unit (not shown in the figure), which is configured to determine the degree of guidance of at least one type of user on the live broadcast topic based on the group participation data and difference data of at least one type of user; and the data determination unit 303 is further configured to: determine the group behavior characteristic data of the type of user in combination with the group participation data, difference data and guidance degree of the type of user.
[0184] In some optional implementations of this embodiment, the above-mentioned device also includes: a pattern determination unit (not shown in the figure), which is configured to: adopt preset abnormality judgment rules, and determine abnormal time periods and abnormality types corresponding to the abnormal time periods from multiple time periods based on the live broadcast topics and group behavior characteristic data corresponding to each of the multiple time periods in the live broadcast process; according to the abnormality type, determine target data from the group behavior characteristic data corresponding to the abnormal time period for clustering to obtain abnormal clustering results; and determine the abnormal pattern of the live broadcast process based on the abnormal clustering results.
[0185] In some optional implementations of this embodiment, the above-mentioned device also includes: a correlation determination unit (not shown in the figure), which is configured to: determine the correlation between multiple abnormal time periods based on the live broadcast topics and group behavior characteristic data corresponding to each of the multiple abnormal time periods; and the pattern determination unit is further configured to: determine the abnormal pattern of the live broadcast process in combination with the abnormal clustering results and the correlation.
[0186] In some optional implementations of this embodiment, the apparatus further includes: a strategy determination unit (not shown in the figure), configured to determine an operation strategy for live broadcast users during the live broadcast process according to the abnormal pattern.
[0187] In some optional implementations of this embodiment, at least one type of user includes temporary audience users, first-time visiting users and followed users, and the health determination unit 304 is further configured to determine a first evaluation indicator, a second evaluation indicator, a third evaluation indicator and a fourth evaluation indicator based on the group behavior characteristic data under the live broadcast topics of each of the multiple time periods during the live broadcast process, wherein the first evaluation indicator represents the number of newly added followed users of the live broadcast users when the increase rate of temporary audience users during the live broadcast process exceeds a preset growth rate threshold, the second evaluation indicator represents the comparison result between the proportion data of followed users and the historical proportion average data, the third evaluation indicator represents the interaction time of first-time visiting users, and the fourth evaluation indicator represents the concentration degree of text data in the multimodal alignment data on preset sensitive topics; the health of the live broadcast process is determined in combination with the first evaluation indicator, the second evaluation indicator, the third evaluation indicator and the fourth evaluation indicator.
[0188] In some optional implementations of this embodiment, the strategy determination unit is further configured to determine an operation strategy for the live broadcast user during the live broadcast process according to the healthiness.
[0189] In the live broadcast data processing device provided in the embodiment of the present application, the data alignment unit performs time alignment on the live broadcast data of multiple modes within a time period to obtain multimodal alignment data. The topic determination unit determines the live broadcast topic within the time period based on the multimodal alignment data. The data determination unit determines the group behavior characteristic data of at least one type of user participating in the live broadcast topic based on the multimodal alignment data, thereby combining the live broadcast data of multiple modes to accurately and comprehensively characterize the live broadcast content and the various types of users participating in the live broadcast process, providing accurate and comprehensive data basis for downstream tasks, and being able to accurately determine the health of the live broadcast room.
[0190] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application. The method corresponding to the electronic device may be the live broadcast data processing method in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.
[0191] The electronic device can be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on such a device. The user device includes, but is not limited to, computers, mobile phones, tablets, smart watches, wristbands, and other terminal devices. The network device includes, but is not limited to, network hosts, single network servers, multiple network servers, or a collection of cloud computing-based computers, and can be used to implement some of the processing functions required for setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computers.
[0192] Figure 4The structure of a device suitable for implementing the methods and / or technical solutions in the embodiments of the present application is shown. The device 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 402 or programs loaded from a storage unit 408 into a random access memory (RAM) 403. RAM 403 also stores various programs and data required for system operation. CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.
[0193] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, touch screen, microphone, infrared sensor, and the like; an output section 407 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), LED display, OLED display, and speakers; a storage section 408 including one or more computer-readable media such as a hard disk, optical disk, magnetic disk, and semiconductor memory; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card or a modem. Communication section 409 performs communication processing via a network such as the Internet.
[0194] In particular, the methods and / or embodiments of the present application can be implemented as computer software programs. For example, the embodiments disclosed herein include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method illustrated in the flowchart. When the computer program is executed by the central processing unit (CPU) 401, the aforementioned functions defined in the method of the present application are performed.
[0195] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon. The computer program instructions can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application.
[0196] Specifically, this embodiment may employ any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0197] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0198] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0199] The computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0200] The flow diagrams and block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of devices, methods and computer program products according to various embodiments disclosed. In this regard, each block in the flow diagrams and block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and
[0201] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.
[0202] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or page components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0203] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0204] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.
[0205] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some of the steps of the method described in various embodiments of this application. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc., various media capable of storing program code.
[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
[0207] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.
Claims
1. A live broadcast data processing method, comprising: Perform time sequence alignment on live broadcast data of multiple modes within a time period to obtain multi-modal aligned data; Determining a live broadcast topic within the time period based on the multimodal alignment data; Determining group behavior characteristic data of at least one type of users participating in the live broadcast topic based on the multimodal alignment data; The healthiness of the live broadcast process is determined based on the group behavior characteristic data under the live broadcast topics of each of the multiple time periods during the live broadcast process.
2. The method according to claim 1, wherein Determining the live broadcast topic within the time period according to the multimodal alignment data includes: Determining the popularity of multiple candidate topics within the time period based on the multimodal alignment data; According to the popularity, the live broadcast topic is determined from the plurality of candidate topics.
3. The method according to claim 1, wherein Before determining the popularity of multiple candidate topics within the time period based on the multimodal alignment data, the method further includes: Clustering feature data of the multimodal alignment data within multiple time periods during the live broadcast process to obtain multiple topic clusters; A plurality of candidate topics corresponding one-to-one to the plurality of topic clusters are determined.
4. The method according to claim 1, wherein Determining group behavior characteristic data of at least one type of users participating in the live broadcast topic based on the multimodal alignment data includes: Determining group participation data of at least one type of user according to the multimodal alignment data, wherein the group participation data represents the participation of the type of user in the live broadcast topic; Determining difference data between group engagement data of at least one type of user and historical group engagement data corresponding to a historical live broadcast process; For at least one type of user, group behavior feature data of the type of user is obtained by combining the group participation data and difference data of the type of user.
5. The method according to claim 4, wherein Also includes: Determining the degree to which at least one type of user guides the live broadcast topic based on the group participation data and difference data of at least one type of user; as well as Combining the group participation data and difference data of the type of users to obtain group behavior feature data of the type of users includes: The group behavior characteristic data of the users of the type are determined by combining the group participation data, difference data and guidance degree of the users of the type.
6. The method according to any one of claims 1 to 5, wherein Also includes: Adopting a preset abnormality determination rule, based on the live broadcast topics and group behavior characteristic data corresponding to each of the multiple time periods during the live broadcast process, determining an abnormal time period and an abnormality type corresponding to the abnormal time period from the multiple time periods; According to the abnormality type, target data is determined from the group behavior characteristic data corresponding to the abnormal time period for clustering to obtain abnormal clustering results; According to the abnormal clustering result, an abnormal pattern of the live broadcast process is determined.
7. The method according to claim 6, wherein Also includes: Determining the correlation between the multiple abnormal time periods based on the live broadcast topics and group behavior characteristic data corresponding to the multiple abnormal time periods; as well as Determining the abnormal pattern of the live broadcast process according to the abnormal clustering result includes: The abnormal pattern of the live broadcast process is determined by combining the abnormal clustering result and the correlation.
8. The method according to claim 6, wherein: Also includes: Based on the abnormal pattern, an operation strategy for the live broadcast user during the live broadcast process is determined.
9. The method according to any one of claims 1 to 5, wherein At least one type of user includes a temporary viewer user, a first-time visitor user, and a follower user, and Determining the health of the live broadcast process based on group behavior characteristic data under the live broadcast topics in each of the multiple time periods during the live broadcast process includes: Determine a first evaluation indicator, a second evaluation indicator, a third evaluation indicator, and a fourth evaluation indicator based on the group behavior characteristic data under the live broadcast topics of each of the multiple time periods during the live broadcast process, wherein the first evaluation indicator represents the number of newly added followers of the live broadcast user when the increase rate of the temporary audience users during the live broadcast exceeds a preset growth rate threshold; the second evaluation indicator represents the comparison result between the proportion data of the followers and the historical average proportion data; the third evaluation indicator represents the interaction duration of the first-time visiting users; and the fourth evaluation indicator represents the concentration degree of the text data in the multimodal alignment data on the preset sensitive topic; The health of the live broadcast process is determined by combining the first evaluation indicator, the second evaluation indicator, the third evaluation indicator and the fourth evaluation indicator.
10. The method according to claim 9, wherein: Also includes: An operation strategy for the live broadcast user during the live broadcast process is determined based on the health level.
11. A live broadcast data processing device, comprising: A data alignment unit is configured to perform time sequence alignment on live broadcast data of multiple modes within a time period to obtain multi-modal aligned data; a topic determination unit, configured to determine a live broadcast topic within the time period based on the multimodal alignment data; a data determining unit configured to determine group behavior characteristic data of at least one type of users participating in the live broadcast topic based on the multimodal alignment data; The healthiness determination unit is configured to determine the healthiness of the live broadcast process according to group behavior characteristic data under the live broadcast topics of each of the multiple time periods during the live broadcast process.
12. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.
13. A computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions are executable by a processor to implement the method according to any one of claims 1 to 10.
14. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Cited By
Cross-border e-commerce live broadcast interaction analysis and optimization system based on multi-modal big data
CN122053872A