Live broadcast content adjustment method and system based on user interaction behavior analysis
By processing the natural language of live comments and improving the BIRCH clustering algorithm, combined with a multi-dimensional scoring model, the problem of hosts having difficulty identifying audience needs is solved, efficient and real-time adjustments to live content are achieved, and the audience's interactive experience is enhanced.
Patent Information
- Application Number
- CN202510876369.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In live broadcast scenarios, it is difficult for hosts to quickly identify the common needs of the majority of viewers, resulting in low efficiency of the interactive model that relies on manual intervention and inability to effectively adjust the live broadcast content.
Natural language processing technology is used to filter out invalid information in the barrage, combined with standardized expressions of the domain knowledge base, and an improved BIRCH clustering algorithm is used to cluster demands. The TOP3 demands are output based on a multi-dimensional scoring model, and the algorithm parameters are dynamically adjusted through reinforcement learning. The host's response feedback is monitored to optimize content adjustments.
It enables anchors to intuitively obtain the common needs of the majority of viewers, improves the efficiency of live content adjustment, reduces reliance on manual voting, and ensures the real-time and diversity of content response and audience interaction.
Smart Images

Figure CN120602712A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image communication, and in particular relates to a method and system for adjusting live broadcast content based on user interactive behavior analysis. Background Art
[0002] In live broadcast scenarios, anchors often dynamically adjust content through user interactive behaviors (such as gift notes, voting, etc.) to maintain audience retention.
[0003] However, when faced with a massive amount of textual requests, anchors struggle to quickly identify the common needs of the majority of viewers, often requiring manual voting to aggregate opinions. This interactive model, which relies on manual intervention, is inefficient and urgently requires automated or intelligent methods to optimize the decision-making process. Summary of the Invention
[0004] Based on this, it is necessary to provide a live broadcast content adjustment method and system based on user interactive behavior analysis to address the above problems.
[0005] The embodiment of the present invention is implemented as follows: a method for adjusting live content based on user interactive behavior analysis, comprising the following steps:
[0006] For real-time bullet comments, NLP (natural language processing) technology is used to filter out invalid information (such as emoticons and advertisements) to obtain the first processed bullet comment. This is then standardized (correcting typos and adding or removing words to make the expression coherent) based on the domain knowledge base to obtain the second processed bullet comment. Key demand words are extracted from the second processed bullet comment to obtain the third processed bullet comment (e.g., extracting "waterproof" from "Is XX waterproof?").
[0007] Using an improved BIRCH clustering algorithm (an algorithm), the third-party comments are combined based on semantic similarity (for example, "waterproof" + "fall-proof" are grouped as "product durability verification"), and cluster popularity is calculated in real time: popularity = frequency of occurrence × user value weight (high-paying users have higher value weights, which can effectively filter out low-value requests and avoid wasting resources processing invalid information later, such as 10 free accounts flooding the screen with "lucky draw").
[0008] A multi-dimensional scoring model is established based on basic popularity (60%), user group value (30%, weighted by high-paying user demand), and timeliness coefficient (10%, such as demand for limited-time events). The top three demands and confidence scores are output through the multi-dimensional scoring model.
[0009] Monitor user feedback on the host's response to adjusting the live broadcast content according to demand (such as the subsequent interaction volume), and dynamically adjust the parameters of the improved BIRCH clustering algorithm through reinforcement learning.
[0010] In one embodiment, the present invention provides a method for adjusting live content based on user interactive behavior analysis, wherein the processing of real-time bullet comments, using NLP technology to filter invalid information to obtain a first processed bullet comment, combining the first processed bullet comment with a domain knowledge base, standardizing the expression to obtain a second processed bullet comment, extracting key demand words in the second processed bullet comment, and obtaining a third processed bullet comment, specifically includes:
[0011] Using predefined rules (such as regular expressions) and NLP tools (such as TextBlob), invalid information is filtered in real time. Invalid information includes emoticons, URLs (resource locators), and ad keywords. Short text (less than 3 characters) and repeated comments are also removed, retaining valid semantic fragments to obtain the first processed comments.
[0012] Combined with domain knowledge bases (such as e-commerce product dictionaries), text error correction models (such as the BERT-based spellcorrector) are used to correct typos. Dependency analysis (a key technique for parsing sentence grammatical structure in natural language processing) is used to complete missing subjects and predicates (e.g., "Is it waterproof?" → "Is this phone waterproof?"), generating grammatically coherent second-level comments.
[0013] Using named entity recognition (NER) and dependency analysis, verb-noun phrases (such as "waterproof performance" and "is it drop-resistant") are extracted from the second-processed barrage, normalized into core demand words (such as "waterproof" and "drop-resistant"), and the third-processed barrage is obtained.
[0014] In one embodiment, the present invention provides a method for adjusting live content based on user interactive behavior analysis, which uses an improved BIRCH clustering algorithm to process barrage merging requirements from the third processing based on semantic similarity and calculate cluster popularity in real time: popularity = frequency of occurrence × speaking user value weight step, specifically including:
[0015] The lightweight Sentence-BERT model is used to encode the demand words into 768-dimensional vectors (high-dimensional numerical arrays used to represent text or semantic features). Semantic relevance is calculated using cosine similarity (for example, the similarity between "waterproof" and "moisture-proof" is greater than 0.8). The demand words are then merged from the third-processed bullet comments.
[0016] The improved BIRCH clustering algorithm is used to incorporate user value weights (payment levels are mapped to weight coefficients, such as VIP users = 5.0, free users = 0.2) into cluster radius calculation. The demand frequency × weight sum within the cluster is updated in real time as the initial heat, and the threshold for merging similar clusters is set to 0.75.
[0017] Scan clusters (sets formed by data grouping) every 5 seconds, sort them according to the formula heat = Σ(demand frequency × user weight), and set a heat decay mechanism (such as 2% decay per second) to prevent historical demand from occupying a high position for a long time.
[0018] In one embodiment, the present invention provides a method for adjusting live content based on user interactive behavior analysis, wherein a multi-dimensional scoring model is established based on basic popularity, user group value, and timeliness coefficient, and the step of outputting the TOP3 demands and confidence scores through the multi-dimensional scoring model specifically includes:
[0019] A multi-dimensional scoring model is established based on basic popularity, user group value, and time efficiency coefficient. Basic popularity (60%) is calculated by taking the normalized value of the total popularity of the current cluster; user value (30%) is calculated by calculating the proportion of high-paying users with different needs; time efficiency coefficient (10%) is calculated by combining the activity countdown (e.g., if there are 10 minutes left, the coefficient is multiplied by 1.5).
[0020] User demand is calculated using the total score = 0.6 × basic popularity + 0.3 × user value + 0.1 × time efficiency coefficient. The top three demands must meet the requirement that the total score difference is less than 15% (to avoid monopoly). The confidence level is the percentage of independent users who have encountered this demand in the last minute.
[0021] Establish a demand status cache. If a demand is the TOP1 demand for three consecutive times, it will be added to the 30-minute blocking list and will no longer participate in the scoring during this period.
[0022] In one embodiment, the present invention provides a method for adjusting live content based on user interactive behavior analysis, wherein a multi-dimensional scoring model is established based on basic popularity, user group value, and timeliness coefficient, and the step of outputting TOP3 demands and confidence scores through the multi-dimensional scoring model further includes:
[0023] From the requirements ranked 4-10 by popularity, exclude semantically duplicate requirements with the top 3 (similarity > 0.6), and store the remaining requirements in the candidate pool in descending order of basic popularity;
[0024] After the anchor completes the response to a TOP1 demand, a demand will be selected from the candidate pool with a probability of 20% (such as the 4th, 5th, and 6th places are drawn in a cycle) as the next TOP1 demand. At the same time, the interaction volume of low-paying users' barrage will be monitored. If the interaction increases by less than 10% within 5 minutes after insertion, the probability will be increased to 30% (the activity level of low-paying users is not ideal, and low-paying users need to be retained further).
[0025] In one embodiment, the present invention provides a live content adjustment system based on user interaction behavior analysis, comprising:
[0026] The barrage processing module processes real-time barrages, using NLP (natural language processing) technology to filter out invalid information (such as emoticons and advertisements) to obtain a first-processed barrage. This module then standardizes the first-processed barrage by combining it with the domain knowledge base (correcting typos and adding or removing words to make it coherent) to obtain a second-processed barrage. The module then extracts key demand terms from the second-processed barrage to obtain a third-processed barrage (e.g., extracting "waterproof" from the question "Is XX waterproof?").
[0027] The cluster popularity calculation module uses an improved BIRCH clustering algorithm (an algorithm) to merge requirements from the third processing barrage based on semantic similarity (for example, "waterproof" + "fall-resistant" are grouped as "product durability verification") and calculate cluster popularity in real time: popularity = frequency of occurrence × user value weight (high-paying users have higher value weights, which can effectively filter out low-value requirements and avoid wasting resources processing invalid information later, such as 10 free accounts flooding the screen with a "lucky draw").
[0028] The demand output module is used to establish a multi-dimensional scoring model based on basic popularity (60%), user group value (30%, weighted by high-paying user demand), and timeliness coefficient (10%, such as demand for limited-time events). The multi-dimensional scoring model outputs the top three demands and confidence scores;
[0029] The feedback adjustment module is used to monitor user feedback on the host's response to adjusting the live broadcast content according to demand (such as the subsequent interaction volume), and dynamically adjust the parameters of the improved BIRCH clustering algorithm through reinforcement learning.
[0030] In one embodiment, the present invention provides a live content adjustment system based on user interactive behavior analysis, wherein the bullet screen processing module includes:
[0031] The first processing barrage acquisition unit is used to filter invalid information in real time using predefined rules (such as regular expressions) and NLP tools (such as TextBlob). Invalid information includes emoticons, URLs (URLs), and advertising keywords. At the same time, it removes duplicates from short text (less than 3 characters) and repeated barrages, retaining valid semantic segments to obtain the first processed barrage.
[0032] The second processing barrage acquisition unit combines domain knowledge bases (such as e-commerce product dictionaries) with a text error correction model (such as the BERT-based spell corrector) to correct typos. It also uses dependency analysis (a key technique for parsing sentence grammatical structure in natural language processing) to complete missing subjects and predicates (e.g., "Is it waterproof?" → "Is this phone waterproof?"), generating grammatically coherent second processing barrages.
[0033] The third processing barrage acquisition unit uses named entity recognition (NER) and dependency analysis to extract verb-noun phrases (such as "waterproof performance" and "is it drop-resistant") from the second processing barrage, normalizes them into core demand words (such as "waterproof" and "drop-resistant"), and obtains the third processing barrage.
[0034] In one embodiment, the present invention provides a live content adjustment system based on user interaction behavior analysis, wherein the cluster heat calculation module includes:
[0035] The demand merging unit uses the lightweight Sentence-BERT model to encode demand words into 768-dimensional vectors (high-dimensional numerical arrays used to represent text or semantic features), calculates semantic relevance using cosine similarity (for example, the similarity between "waterproof" and "moisture-proof" is greater than 0.8), and merges the demands from the third-processed bullet comments.
[0036] The initial heat unit is used to integrate user value weights (payment levels are mapped to weight coefficients, such as VIP users = 5.0 and free users = 0.2) into cluster radius calculations using the improved BIRCH clustering algorithm. The sum of the demand frequency and weight within the cluster is updated in real time as the initial heat. The threshold for merging similar clusters is set to 0.75.
[0037] The real-time heat calculation unit is used to scan clusters (collections formed by data grouping) every 5 seconds, sort them according to the formula heat = Σ(demand frequency × user weight), and set a heat decay mechanism (such as 2% decay per second) to prevent historical demand from occupying a high level for a long time.
[0038] In one embodiment, the present invention provides a live content adjustment system based on user interaction behavior analysis, wherein the demand output module includes:
[0039] The model building unit is used to establish a multi-dimensional scoring model based on basic popularity, user group value, and time efficiency coefficient. Basic popularity (60%) is calculated by taking the normalized value of the total popularity of the current cluster; user value (30%) is calculated by calculating the proportion of high-paying users with different needs; time efficiency coefficient (10%) is calculated by combining the activity countdown (if there are 10 minutes left, the coefficient is ×1.5).
[0040] The user demand calculation unit is used to calculate user demand based on the total score = 0.6 × basic popularity + 0.3 × user value + 0.1 × timeliness coefficient. The top three demands must meet the requirement that the total score difference is less than 15% (to avoid monopoly). The confidence level is the percentage of independent users who have encountered this demand in the last minute.
[0041] The demand shielding unit is used to establish a demand status cache. If a demand is the TOP1 demand for three consecutive times, the demand will be added to the 30-minute shielding list and will no longer participate in the scoring during this period.
[0042] In one embodiment, the present invention provides a live content adjustment system based on user interaction behavior analysis, wherein the demand output module further includes:
[0043] The candidate pool construction unit is used to exclude semantically duplicate requirements (similarity > 0.6) from the top 10 requirements ranked by popularity, and store the remaining requirements in the candidate pool in descending order of basic popularity;
[0044] The demand insertion unit is used to select demands from the candidate pool with a probability of 20% (such as the 4th, 5th, and 6th places are drawn in a cycle) as the next TOP1 demand after the anchor completes the response to a TOP1 demand. At the same time, the interaction volume of low-paying users' barrage is monitored. If the interaction increases by less than 10% within 5 minutes after the insertion, the probability is increased to 30% (the activity level of low-paying users does not increase ideally, and low-paying users are further retained).
[0045] Compared with the existing technology, the beneficial effect of the present invention is: the present invention ranks current user needs by processing the barrage content during the live broadcast, so that the anchor can intuitively obtain the common needs of the majority of viewers, and there is no need to initiate voting to aggregate user opinions, which is highly efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A flowchart of a method for adjusting live content based on user interactive behavior analysis provided by an embodiment of the present invention.
[0047] Figure 2 A schematic diagram of the process of processing barrage content provided by an embodiment of the present invention.
[0048] Figure 3 A schematic diagram of the process of calculating cluster heat provided by an embodiment of the present invention.
[0049] Figure 4 A schematic diagram of the process of outputting TOP3 requirements and demand shielding provided in an embodiment of the present invention.
[0050] Figure 5 A schematic diagram of the process of inserting a low-heat requirement provided by an embodiment of the present invention.
[0051] Figure 6 A schematic diagram of a live content adjustment system based on user interactive behavior analysis provided by an embodiment of the present invention.
[0052] Figure 7 A schematic diagram of a bullet comment processing module provided in an embodiment of the present invention.
[0053] Figure 8 A schematic diagram of a cluster heat calculation module provided in an embodiment of the present invention.
[0054] Figure 9This is a schematic diagram of the first part of the demand output module provided in an embodiment of the present invention.
[0055] Figure 10 This is a schematic diagram of the second part of the demand output module provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0057] It is understood that the terms "first," "second," etc., used herein may be used to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script without departing from the scope of this application.
[0058] In one embodiment, Figure 1 As shown, a live broadcast content adjustment method based on user interaction behavior analysis includes the following steps:
[0059] Step S1: Processing the real-time bullet comments, using NLP (natural language processing) technology to filter out invalid information (such as emoticons and advertisements) to obtain a first processed bullet comment. Based on the domain knowledge base, the first processed bullet comment is standardized (correcting typos and adding or removing words to make the expression smooth) to obtain a second processed bullet comment. Key demand words in the second processed bullet comment are extracted to obtain a third processed bullet comment (e.g., extracting "waterproof" from "Is XX waterproof").
[0060] In step S2, an improved BIRCH clustering algorithm (an algorithm) is used to merge requirements from the third processed bullet comments based on semantic similarity (e.g., "waterproof" + "fall-resistant" are grouped as "product durability verification"), and cluster popularity is calculated in real time: popularity = frequency of occurrence × user value weight (high-paying users have higher value weights, which can effectively filter out low-value requirements and avoid wasting resources processing invalid information later, such as 10 free accounts flooding the screen with "lucky draw").
[0061] Step S3: Build a multi-dimensional scoring model based on basic popularity (60%), user group value (30%, weighted by high-paying user demand), and timeliness coefficient (10%, such as demand for limited-time events). Output the top three demands and confidence scores through the multi-dimensional scoring model.
[0062] Step S4: monitor user feedback (such as subsequent interaction volume) on the host's response to adjusting the live broadcast content according to demand, and dynamically adjust the parameters of the improved BIRCH clustering algorithm through reinforcement learning.
[0063] Step S1 constructs analyzable demand terms through three layers of text purification: first filtering out noise (emojis / advertisements), then standardizing semantics (error correction and completion), and finally extracting core demand terms. This design ensures that raw comments are converted into structured data, addressing the fragmentation and high noise issues of text in live streaming scenarios. Step S2 employs value-weighted clustering: an improved BIRCH algorithm incorporates user payment weights, achieving a value of popularity = frequency × user value. This not only consolidates semantically similar demands (such as waterproof / moisture-proof) but also suppresses low-value spam (such as free user lottery requests). Step S3 establishes a multi-dimensional scoring model: 60% base popularity ensures coverage of mainstream demands, 30% weighting high-paying users ensures commercial value, and 10% timeliness coefficient responds to limited-time events. This balance avoids bias in a single metric. Step S4 introduces a reinforcement learning feedback loop: monitoring changes in user interaction volume dynamically adjusts algorithm parameters, achieving system adaptive optimization and ensuring that content adjustment strategies consistently align with actual results.
[0064] In one embodiment, Figure 2 As shown, a live content adjustment method based on user interactive behavior analysis is provided. Step S1 processes the real-time bullet screen, filters invalid information using NLP technology, obtains a first processed bullet screen, standardizes the first processed bullet screen in combination with a domain knowledge base, obtains a second processed bullet screen, extracts key demand words from the second processed bullet screen, and obtains a third processed bullet screen. The steps specifically include:
[0065] Step S11: Invalid information is filtered in real time using predefined rules (e.g., regular expressions) and NLP tools (e.g., TextBlob). Invalid information includes emoticons, URLs (URLs), and advertising keywords. Short text (less than 3 characters) and repeated comments are deduplicated, and valid semantic segments are retained to obtain a first processed comment.
[0066] Step S12: In combination with a domain knowledge base (e.g., an e-commerce product dictionary), a text error correction model (e.g., a BERT-based spell corrector) is used to correct typos and dependency analysis (a key technique for parsing sentence grammatical structure in natural language processing) is used to complete missing subjects / predicates (e.g., "Is it waterproof?" → "Is this phone waterproof?"), generating a grammatically fluent second processed comment.
[0067] In step S13, named entity recognition (NER) and dependency analysis are used to extract verb-noun phrases (such as "waterproof performance" and "is it drop-resistant") from the second processed barrage, normalize them into core demand words (such as "waterproof" and "drop-resistant"), and obtain the third processed barrage.
[0068] Step S11 uses a dual-track filtering approach of rules and natural language processing (NLP): regular expressions remove URL / advertising keywords, and TextBlob (an open-source natural language processing library) processes emoticons and deduplicates short text. This design addresses the high noise characteristics of live broadcast comments and quickly retains valid semantic fragments. Step S12 incorporates domain knowledge bases for deep correction: BERT-based spellcorrector addresses pinyin input errors, and dependency analysis completes missing components, converting unstructured spoken language into standardized, machine-parseable text. Step S13 extracts requirements: Named entity recognition captures key nouns, and dependency analysis associates verb phrases, normalizing diverse expressions such as "is it drop-resistant?" to the core term "fall-resistant," providing standardized input for clustering.
[0069] In one embodiment, Figure 3 As shown, a live content adjustment method based on user interactive behavior analysis is described. In step S2, an improved BIRCH clustering algorithm is used to process the barrage merging requirements from the third processing based on semantic similarity, and the cluster heat is calculated in real time: heat = frequency of occurrence × speaking user value weight step, specifically including:
[0070] Step S21: Use the lightweight Sentence-BERT model to encode the demand words into 768-dimensional vectors (high-dimensional numerical arrays used to represent text or semantic features). Calculate semantic relevance using cosine similarity (e.g., the similarity between "waterproof" and "moisture-proof" is greater than 0.8). Merge the demand words from the third processed bullet comments.
[0071] Step S22: The user value weight (payment level is mapped to a weight coefficient, such as VIP user = 5.0, free user = 0.2) is incorporated into the cluster radius calculation through the improved BIRCH clustering algorithm. The demand frequency in the cluster is updated in real time × the sum of the weights as the initial heat. The threshold for merging similar clusters is set to 0.75.
[0072] In step S23, clusters (sets formed by data grouping) are scanned every 5 seconds and sorted according to the formula heat = Σ(demand frequency × user weight). A heat decay mechanism is set (e.g., decaying by 2% per second) to prevent historical demand from occupying a high position for a long time.
[0073] Step S21 uses a lightweight semantic encoding: the Sentence-BERT model (an improved sentence embedding model based on BERT) generates a 768-dimensional vector. Synonyms (waterproof / moisture-proof) are merged when cosine similarity exceeds 0.8 to achieve semantic association. Step S22 improves value-sensitive clustering: user payment levels (VIP = 5.0 / Free = 0.2) are incorporated into cluster radius calculation, making the popularity = Σ(frequency × weight). This algorithmically suppresses interference from low-value requests. Step S23 designs a dynamic decay mechanism: clusters are scanned and sorted every 5 seconds, with popularity decaying by 2% per second to prevent historical requests from being retained (such as resolved "price questions") and ensure that top requests reflect current audience focus in real time.
[0074] In one embodiment, Figure 4 As shown, a live content adjustment method based on user interactive behavior analysis is described. Step S3 establishes a multi-dimensional scoring model based on basic popularity, user group value, and timeliness coefficient. The TOP3 demand and confidence scoring steps are output through the multi-dimensional scoring model. Specifically, the steps include:
[0075] Step S31: A multi-dimensional scoring model is established based on basic popularity, user group value, and time efficiency coefficient. Basic popularity (60%): takes the normalized value of the total popularity of the current cluster; User value (30%): calculates the proportion of high-paying users with different needs; Time efficiency coefficient (10%): combines the activity countdown (e.g., if there are 10 minutes left, multiply by 1.5);
[0076] Step S32: Calculate user demands based on the total score = 0.6 × base popularity + 0.3 × user value + 0.1 × timeliness coefficient. The top three demands must have a total score difference less than 15% (to avoid monopoly). The confidence level is the percentage of independent users who have encountered the demand in the last minute.
[0077] Step S33: Create a demand status cache. If a demand is the TOP1 demand for three consecutive times, the demand will be added to the 30-minute blocking list and will no longer participate in the scoring during this period.
[0078] Step S31 establishes a multi-dimensional scoring system: 60% base popularity ensures coverage, 30% high-paying user share strengthens commercial value, and 10% timeliness factor (e.g., limited-time events x 1.5) enhances strategic flexibility. Step S32 establishes an anti-monopoly mechanism: the total score difference between the top three requests must be less than 15% to prevent a single request from dominating the screen, and confidence is calculated based on the number of independent users to prevent spamming and cheating. Step S33 introduces a demand cooling system: three consecutive top-ranking requests are added to a 30-minute blacklist, forcing the streamer to respond to diverse requests (e.g., avoiding repeated "price questions"), thereby enhancing the richness of live content.
[0079] In one embodiment, Figure 5As shown, a live content adjustment method based on user interactive behavior analysis, said step S3, based on basic popularity, user group value, and timeliness coefficient, establishes a multi-dimensional scoring model, and outputs the TOP3 demands and confidence scores through the multi-dimensional scoring model, further comprising:
[0080] Step S34: Eliminate semantically duplicate requirements (similarity > 0.6) from the top 3 requirements ranked 4-10 by popularity, and store the remaining requirements in a candidate pool in descending order of base popularity.
[0081] In step S35, after the anchor completes the response to a TOP1 demand, a demand is selected from the candidate pool with a probability of 20% (such as the 4th, 5th, and 6th are drawn in a cycle) as the next TOP1 demand, and the interaction volume of low-paying users' barrage is monitored at the same time. If the interaction increases by less than 10% within 5 minutes after insertion, the probability is increased to 30% (the increase in the activity of low-paying users is not ideal, and low-paying users are further retained).
[0082] Step S34 establishes dynamic candidate pool management: Duplicates with a similarity greater than 0.6 with the top 3 requests are excluded from the top 4-10 requests, and the requests are sorted in descending order by base popularity, ensuring both differentiation and value. Step S35 employs an exploratory response strategy: After the streamer completes the top 1 request, they rotate through the candidate pool (e.g., the fourth request) with a 20% probability. They monitor the engagement of low-paying users and make adjustments to retain them and maintain the number of people in the livestream. This design balances responses to both top and long-tail requests, using a probabilistic mechanism to gauge the interests of low-paying users and improve retention rates for this group.
[0083] In one embodiment, Figure 6 As shown, a live content adjustment system based on user interaction behavior analysis includes:
[0084] Danmu processing module 1 is used to process real-time Danmu comments, using NLP (natural language processing) technology to filter out invalid information (such as emoticons and advertisements) to obtain a first processed Danmu comment. Based on the domain knowledge base, the first processed Danmu comment is standardized (correcting typos and adding or removing words to make the expression smooth) to obtain a second processed Danmu comment. Key demand words in the second processed Danmu comment are extracted to obtain a third processed Danmu comment (for example, extracting "waterproof" from the question "Is XX waterproof?").
[0085] Cluster heat calculation module 2 is used to use the improved BIRCH clustering algorithm (an algorithm) to merge requirements from the third processing barrage based on semantic similarity (for example, "waterproof" + "fall-resistant" is classified as "product durability verification"), and calculate cluster heat in real time: heat = frequency of occurrence × speaking user value weight (high-paying users have higher value weight, which can effectively filter out low-value requirements and avoid the subsequent waste of resources processing invalid information, such as 10 free accounts flooding the screen with "lucky draw");
[0086] Demand output module 3 is used to establish a multi-dimensional scoring model based on basic popularity (60%), user group value (30%, weighted by high-paying user demand), and timeliness coefficient (10%, such as demand for limited-time events). The multi-dimensional scoring model outputs the top three demands and confidence scores;
[0087] Feedback adjustment module 4 is used to monitor user feedback (such as subsequent interaction volume) on the host's response to adjusting the live broadcast content according to demand, and dynamically adjust the parameters of the improved BIRCH clustering algorithm through reinforcement learning.
[0088] Feedback Adjustment Module 4 specifically monitors user behavior (e.g., barrage growth rate, number of gifts, and duration of stay) within five minutes after the streamer responds to the top three / non-preferred requests. The feedback score is then calculated using a weighted formula (e.g., feedback score = 0.6 × barrage increase + 0.3 × gift value + 0.1 × retention rate). Parameter Adjustment Strategy: If the feedback score falls below a threshold (e.g., <60 points), parameter optimization is initiated: the modified BIRCH clustering radius is reduced (e.g., 0.35 to 0.30) to more strictly merge similar requests and avoid overgeneralization; the user value weight coefficient is increased (e.g., VIP weight 5.0 to 6.0) to strengthen the identification of high-paying user requests; and the timeliness coefficient is increased (10% to 15%) to prioritize requests for limited-time events. Reinforcement Learning Training: Using a DQN (Deep Q-Network) model, with parameter combinations as actions and feedback score increases as rewards, the algorithm dynamically generates an optimal parameter mapping table through offline simulation (using historical live broadcast data) and online fine-tuning to ensure its adaptability to different live broadcast scenarios.
[0089] In one embodiment, Figure 7 As shown, a live content adjustment system based on user interactive behavior analysis, the barrage processing module 1 includes:
[0090] The first processed bullet comment acquisition unit 11 is configured to filter invalid information in real time using predefined rules (e.g., regular expressions) and NLP tools (e.g., TextBlob). Invalid information includes emoticons, URLs (URLs), and advertising keywords. The first processed bullet comment is then deduplicated from short text (less than 3 characters) and repeated bullet comments, retaining valid semantic segments to obtain a first processed bullet comment.
[0091] The second processed comment acquisition unit 12 combines domain knowledge (e.g., an e-commerce product dictionary) with a text error correction model (e.g., a BERT-based spell corrector) to correct typos and completes missing subjects / predicates (e.g., "Is it waterproof?" → "Is this phone waterproof?") through dependency analysis (a key technique for parsing sentence grammatical structure in natural language processing), generating a grammatically coherent second processed comment.
[0092] The third processing barrage acquisition unit 13 uses named entity recognition (NER) and dependency analysis to extract verb-noun phrases (such as "waterproof performance" and "is it drop-resistant") from the second processing barrage, normalizes them into core demand words (such as "waterproof" and "drop-resistant"), and obtains the third processing barrage.
[0093] The second bullet comment acquisition unit 12 can also further add dialect conversion functions (e.g., converting Minnan dialect and Cantonese to standard language) and homophonic error correction functions (e.g., "anti-tax" → "waterproof"). Dynamically loading dialect dictionaries based on user region tags can address the issue of missing information from non-Mandarin speaking users.
[0094] In one embodiment, Figure 8 As shown, a live content adjustment system based on user interactive behavior analysis, the cluster heat calculation module 2 includes:
[0095] The demand merging unit 21 is used to encode the demand words into 768-dimensional vectors (high-dimensional numerical arrays used to represent text or semantic features) using the lightweight Sentence-BERT model, calculate the semantic relevance using cosine similarity (e.g., the similarity between "waterproof" and "moisture-proof" is greater than 0.8), and merge the demands from the third processed bullet comments;
[0096] Initial heat unit 22 is used to integrate user value weight (payment level is mapped to weight coefficient, such as VIP user = 5.0, free user = 0.2) into cluster radius calculation through the improved BIRCH clustering algorithm, and update the demand frequency in the cluster × the sum of weights in real time as the initial heat. The threshold for merging similar clusters is set to 0.75;
[0097] The real-time heat calculation unit 23 is used to scan the clusters (a collection of data groups) every 5 seconds, sort them according to the formula heat = Σ(demand frequency × user weight), and set a heat decay mechanism (such as decaying by 2% per second) to prevent historical demand from occupying a high position for a long time.
[0098] The real-time popularity calculation unit 22 can further detect abnormal speech patterns (e.g., 10 identical comments within 0.1 seconds) and combine device fingerprints to identify fraudulent groups. It automatically resets the weight of fraudulent clusters to zero, reducing the pollution of popularity caused by spamming and suppressing false demand.
[0099] In one embodiment, Figure 9 As shown, a live content adjustment system based on user interaction behavior analysis, the demand output module 3 includes:
[0100] The model building unit 31 is used to establish a multi-dimensional scoring model based on basic popularity, user group value, and time efficiency coefficient. The basic popularity (60%) is calculated by taking the normalized value of the total popularity of the current cluster; the user value (30%) is calculated by calculating the proportion of high-paying users with different needs; and the time efficiency coefficient (10%) is calculated by combining the activity countdown (e.g., if there are 10 minutes left, the coefficient is multiplied by 1.5).
[0101] User demand calculation unit 32 is used to calculate user demand based on the total score = 0.6 × basic popularity + 0.3 × user value + 0.1 × timeliness coefficient. The top three demands must meet the requirement that the total score difference is less than 15% (to avoid monopoly). The confidence level is the percentage of independent users who have encountered the demand in the last minute.
[0102] The demand shielding unit 33 is used to establish a demand status cache. If a demand is the TOP1 demand for three consecutive times, the demand will be added to the 30-minute shielding list and will no longer participate in the scoring during this period.
[0103] The shielding time in the demand shielding unit 33 can be further changed to a dynamic calculation, setting a benchmark time based on the demand type (price consultation = 10 minutes, function question = 30 minutes), and adding the recent frequency coefficient (each time exceeding the threshold + 5 minutes) to avoid high-value demands being mistakenly shielded.
[0104] In one embodiment, Figure 10 As shown, a live content adjustment system based on user interactive behavior analysis, the demand output module 3 also includes:
[0105] The candidate pool construction unit 34 is used to exclude semantically duplicate requirements (similarity > 0.6) with the top 3 requirements from the requirements ranked 4-10 by popularity, and store the remaining requirements into the candidate pool in descending order of basic popularity;
[0106] The demand insertion unit 35 is used to select a demand from the candidate pool with a probability of 20% (such as the 4th, 5th, and 6th positions are drawn in a cycle) as the next TOP1 demand after the anchor completes the response to a TOP1 demand, and monitor the interaction volume of low-paying users' barrage at the same time. If the interaction increases by less than 10% within 5 minutes after the insertion, the probability is increased to 30% (the activity level of low-paying users does not increase ideally, and low-paying users are further retained).
[0107] In the demand insertion unit 35, the probability adjustment can be further optimized. When the candidate demand is responded to, the interaction increase of the user group is tracked in real time and the probability is adjusted. For example, if the low-paying user barrage is +15%, the probability is +0.5%.
[0108] It should be understood that, although the various steps in the flow chart of each embodiment of the present invention are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0109] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0110] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
[0111] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
[0112] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A method for adjusting live content based on user interaction behavior analysis, characterized in that: The live broadcast content adjustment method based on user interaction behavior analysis includes the following steps: For real-time bullet comments, NLP technology is used to filter out invalid information to obtain the first processed bullet comments. The first processed bullet comments are standardized based on the domain knowledge base to obtain the second processed bullet comments. The key demand words in the second processed bullet comments are extracted to obtain the third processed bullet comments. Adopting the improved BIRCH clustering algorithm, based on semantic similarity, the third processing barrage merging demand is used to calculate the cluster popularity in real time: popularity = frequency of occurrence × weight of the user value of the speaker; A multi-dimensional scoring model is established based on basic popularity, user group value, and timeliness coefficient, and the top 3 demands and confidence scores are output through the multi-dimensional scoring model; Monitor user feedback on the host's response to adjusting the live broadcast content according to demand, and dynamically adjust the parameters of the improved BIRCH clustering algorithm through reinforcement learning.
2. The method for adjusting live content based on user interactive behavior analysis according to claim 1, characterized in that: The processing of real-time bullet comments, using NLP technology to filter invalid information to obtain a first processed bullet comment, combining the first processed bullet comment with a domain knowledge base to standardize the expression to obtain a second processed bullet comment, extracting key demand words in the second processed bullet comment, and obtaining a third processed bullet comment, specifically includes: Predefined rules and NLP tools are used to filter invalid information in real time, including emoticons, URLs, and advertising keywords. Short text and repeated comments are also removed, retaining valid semantic fragments to obtain the first-level processed comments. Combined with the domain knowledge base, the text error correction model is used to correct typos and missing subjects / predicates are completed through dependency analysis to generate grammatically fluent second-processed comments. Using named entity recognition and dependency analysis, verb-noun phrases are extracted from the second-processed barrage, normalized into core demand words, and the third-processed barrage is obtained.
3. The method for adjusting live content based on user interactive behavior analysis according to claim 1, characterized in that: The improved BIRCH clustering algorithm is used to process the barrage merging requirements from the third step based on semantic similarity, and the cluster heat is calculated in real time: heat = frequency of occurrence × speaking user value weight step, specifically including: The lightweight Sentence-BERT model is used to encode the demand words into 768-dimensional vectors, and the semantic relevance is calculated by cosine similarity to merge the demands from the third-processed comments. The improved BIRCH clustering algorithm is used to incorporate user value weight into cluster radius calculation. The sum of demand frequency and weight within the cluster is updated in real time as the initial heat, and the threshold for merging similar clusters is set to 0.
75. The clusters are scanned every 5 seconds and sorted according to the formula heat = Σ(demand frequency × user weight). A heat decay mechanism is set to prevent historical demand from occupying a high position for a long time.
4. The method for adjusting live content based on user interactive behavior analysis according to claim 1, characterized in that: The step of establishing a multi-dimensional scoring model based on basic popularity, user group value, and timeliness coefficient, and outputting the TOP3 demands and confidence scores through the multi-dimensional scoring model specifically includes: A multi-dimensional scoring model is established based on basic popularity, user group value, and time efficiency coefficient. Basic popularity: takes the normalized value of the total popularity of the current cluster; user value: calculates the proportion of high-paying users with different needs; time efficiency coefficient: combines the activity countdown; User demands are calculated using the total score = 0.6 × basic popularity + 0.3 × user value + 0.1 × timeliness coefficient. The top three demands must meet a total score difference of less than 15%. The confidence level is the percentage of independent users who have encountered this demand in the last minute. Establish a demand status cache. If a demand is the TOP1 demand for three consecutive times, it will be added to the 30-minute blocking list and will no longer participate in the scoring during this period.
5. The method for adjusting live content based on user interactive behavior analysis according to claim 4, characterized in that: The step of establishing a multi-dimensional scoring model based on basic popularity, user group value, and timeliness coefficient, and outputting the TOP3 demands and confidence scores through the multi-dimensional scoring model, further includes: Eliminate semantically duplicate requirements from the top 10 most popular requirements and place them in the candidate pool in descending order of popularity. When the anchor completes the response to a TOP1 demand, a demand will be selected from the candidate pool with a probability of 20% as the next TOP1 demand. At the same time, the interaction volume of low-paying users' barrage will be monitored. If the interaction increase is less than 10% within 5 minutes after insertion, the probability will be increased to 30%.
6. A live content adjustment system based on user interactive behavior analysis, characterized in that: The live content adjustment system based on user interactive behavior analysis includes: The barrage processing module is used to process real-time barrages, filter invalid information using NLP technology, obtain a first processed barrage, standardize the first processed barrage based on the domain knowledge base, obtain a second processed barrage, extract key demand words from the second processed barrage, and obtain a third processed barrage; The cluster heat calculation module is used to use the improved BIRCH clustering algorithm to calculate the cluster heat in real time based on semantic similarity from the third processing barrage merging requirements: heat = frequency of occurrence × speaking user value weight; The demand output module is used to establish a multi-dimensional scoring model based on basic popularity, user group value, and timeliness coefficient, and output the top 3 demands and confidence scores through the multi-dimensional scoring model; The feedback adjustment module is used to monitor users' feedback on the host's response to adjusting the live broadcast content according to demand, and dynamically adjust the parameters of the improved BIRCH clustering algorithm through reinforcement learning.
7. The live content adjustment system based on user interactive behavior analysis according to claim 6 is characterized in that: The barrage processing module includes: The first processing barrage acquisition unit is used to filter invalid information in real time using predefined rules and NLP tools. Invalid information includes emoticons, URLs, and advertising keywords. At the same time, it removes duplicate short text and repeated barrages, retains valid semantic segments, and obtains the first processed barrage; The second processing barrage acquisition unit, in combination with the domain knowledge base, uses a text error correction model to correct typos and completes missing subjects / predicates through dependency analysis to generate a grammatically fluent second processing barrage; The third processing barrage acquisition unit uses named entity recognition and dependency analysis to extract verb-noun phrases from the second processing barrage, normalizes them into core demand words, and obtains the third processing barrage.
8. The live content adjustment system based on user interactive behavior analysis according to claim 6 is characterized in that: The cluster heat calculation module includes: The demand merging unit is used to encode demand words into 768-dimensional vectors using the lightweight Sentence-BERT model, calculate semantic relevance through cosine similarity, and merge demands from the third-processed bullet comments; The initial heat unit is used to incorporate user value weights into cluster radius calculations using the improved BIRCH clustering algorithm. The sum of the demand frequency and weight within the cluster is updated in real time as the initial heat, and the threshold for merging similar clusters is set to 0.
75. The real-time heat calculation unit is used to scan clusters every 5 seconds, sort them according to the formula heat = Σ(demand frequency × user weight), and set a heat decay mechanism to prevent historical demand from occupying a high position for a long time.
9. The live content adjustment system based on user interactive behavior analysis according to claim 6, characterized in that: The demand output module includes: The model building unit is used to establish a multi-dimensional scoring model based on basic popularity, user group value, and time efficiency coefficient. Basic popularity: takes the normalized value of the total popularity of the current cluster; user value: calculates the proportion of high-paying users with different needs; time efficiency coefficient: combines the activity countdown; The user demand calculation unit is used to calculate user demand based on the total score = 0.6 × basic popularity + 0.3 × user value + 0.1 × timeliness coefficient. The top three demands must meet the total score difference of less than 15%. The confidence level is the percentage of independent users who have encountered this demand in the last minute. The demand shielding unit is used to establish a demand status cache. If a demand is the TOP1 demand for three consecutive times, the demand will be added to the 30-minute shielding list and will no longer participate in the scoring during this period.
10. The live content adjustment system based on user interactive behavior analysis according to claim 9, characterized in that: The demand output module also includes: The candidate pool construction unit is used to exclude semantically duplicate requirements from the top 3 requirements ranked 4-10 by popularity, and store the remaining requirements in the candidate pool in descending order of basic popularity; The demand insertion unit is used to select a demand from the candidate pool as the next TOP1 demand with a 20% probability after the anchor completes the response to a TOP1 demand, and at the same time monitor the interaction volume of low-paying users' barrage. If the interaction increase is less than 10% within 5 minutes after insertion, the probability is increased to 30%.
Citation Information
Cited By
Scenarized man-machine collaborative optimization method and system based on artificial intelligence
CN121418586A
Interaction system and method based on Internet live broadcast
CN121842416A