Live broadcast scene perception recommendation system

By collecting multi-source data and dynamic user portraits in real time, and generating personalized recommendation strategies, it solves the problem that traditional live broadcast recommendation systems cannot accurately capture user needs, and improves user experience and platform stickiness.

CN120264033AInactive Publication Date: 2025-07-04SHANGHAI YIXING NETWORK TECHNOLOGY CO LTD +1

Patent Information

Application Number
CN202510752063.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional live broadcast recommendation systems cannot accurately capture users' potential needs in live broadcast real-time situations, resulting in the recommended content being inconsistent with user interests and affecting user viewing experience and participation.

Method used

Design a live scene perception recommendation system, and use real-time collection of multi-source data, including live broadcast screen, anchor voice, audience interaction, etc., and combine user behavior trajectory and historical data to build dynamic user portraits, generate personalized recommendation strategies, and push content through sidebar pop-up windows, barrage reminders, etc.

Benefits of technology

It has achieved personalized and precise recommendations, improved user viewing experience and participation, and improved the user stickiness and commercial conversion of the live broadcast platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264033A_ABST
    Figure CN120264033A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of live broadcast, and discloses a live broadcast scene perception recommendation system, which is characterized in that a data acquisition module is used for acquiring multi-source data in a live broadcast process in real time; the scene analysis module is connected with the data acquisition module and is used for receiving and integrating the acquired data and analyzing the live broadcast scene; the user portrait construction module is used for tracking a behavior track of a user on a live broadcast platform in real time, and constructing a dynamic user portrait in combination with historical watching data, a collection record, a consumption record and an interactive behavior in current live broadcast; a recommendation strategy generation module generates a personalized recommendation strategy according to the result of the scene analysis module and the user portrait constructed by the user portrait construction module; the recommendation content pushing module pushes the recommendation content generated by the recommendation strategy generation module to the user; according to the method, personalized and precise recommendation is realized, and the watching experience and participation degree of the user are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of live broadcast technology, and particularly to a live broadcast scene perception recommendation system. Background Art

[0002] With the rapid development of the Internet, the live broadcast industry has become increasingly popular, and various live broadcast platforms have emerged like mushrooms after a spring rain. When users are faced with a vast amount of live broadcast content, it is often difficult for them to quickly find the live broadcast rooms or hosts that they are really interested in. Most traditional recommendation systems recommend based on simple behavioral data such as users' historical viewing records, likes, and comments, lacking in-depth analysis of the real-time live broadcast scene and unable to accurately capture the potential needs of users in the current live broadcast situation. For example, when a user is watching a live sports event and the event reaches a crucial moment, if relevant content such as highlights of the event, instant expert comments, and purchase links for the same sports equipment can be recommended at this time, it will greatly enhance the user's viewing experience and meet their diverse needs. Summary of the Invention

[0003] The purpose of the present invention is to solve the above problems and design a live broadcast scene perception recommendation system.

[0004] The present invention provides a live broadcast scene perception recommendation system, including a data acquisition module for real-time acquisition of multi-source data during the live broadcast process, where the multi-source data at least includes live broadcast video data, host voice data, audience interaction data, and live broadcast platform metadata; A scene analysis module, connected to the data acquisition module, receiving and integrating the acquired data, and parsing the live broadcast scene; A user profile construction module for real-time tracking of the user's behavior trajectory on the live broadcast platform, and constructing a dynamic user profile in combination with historical viewing data, collection records, consumption records, and interactive behaviors in the current live broadcast; A recommendation strategy generation module for generating personalized recommendation strategies according to the results of the scene analysis module and the user profile constructed by the user profile construction module; A recommended content push module for pushing the recommended content generated by the recommendation strategy generation module to the user, and the push methods include pop-up display in the sidebar of the live broadcast interface, bullet screen reminder, and private message push.

[0005] Optionally, in the first implementation manner of the present invention, the scene analysis module includes: An extraction sub-module for extracting video, audio, and text features, and respectively inputting them into the latent space encoder of Stable Diffusion to align multi-modal data at the semantic level and generate aligned feature vectors; An input sub-module for inputting the aligned multi-modal feature vectors and the live platform metadata into a scene classification model, and determining the current live scene type through the scene classification model; An analysis sub-module for performing word segmentation, part-of-speech tagging, and syntactic analysis on the host's speech data and the audience interaction data by using natural language processing technology, and mining the products and topics of user interest through semantic analysis; A judgment sub-module for inputting the processed text into a sentiment analysis algorithm model, judging the sentiment tendency of the live atmosphere, and extracting the scene intention through the LLaMA-3-8B large language model; A construction sub-module for integrating multi-modal information, constructing a spatio-temporal relationship graph and a scene semantic graph, and determining the association between the elements in the live scene.

[0006] Optionally, in the second implementation manner of the present invention, the extraction sub-module includes: A visual feature unit for using WebAssembly to accelerate the decoding of the H.265 encoded stream for the received video data, generating an RGB frame sequence and an optical flow feature, and inputting the RGB frame sequence and the optical flow feature into the SwinTransformerV2 model to extract the visual features in the picture; An audio feature unit for separating the human voice and the background sound from the received audio data through the Audio Super-Resolution network, and inputting the separated audio data into the Wav2Vec 3.0 model to learn the semantic, emotional, and acoustic features in the audio signal; A text feature unit for extracting keywords from the received text data through an FPGA-hardware-accelerated BERT tokenizer.

[0007] Optionally, in the third implementation manner of the present invention, the input sub-module includes: A processing unit for inputting the aligned multi-modal feature vectors and the live platform metadata into a scene classification model, where the scene classification model contains multiple Transformer modules for respectively processing visual and language information; A fusion unit for interacting and fusing visual features, language features, and live platform metadata during the calculation process of multiple Transformer modules; A generation unit for generating a feature representation after the calculation and feature interaction and fusion of multiple Transformer modules, calculating the probability scores of each possible live scene type according to the feature representation, and selecting the scene type with the highest probability score as the judgment result of the current live scene type.

[0008] Optionally, in the fourth implementation manner of the present invention, the user profile construction module includes: The first determination sub-module is used to obtain the user's historical viewing data, collection records, consumption records, and interaction behaviors, and determine the main fields that the user is most interested in based on the live broadcast types and collection records in the historical viewing data; The calculation sub-module is used to judge the user's interest points from the live broadcast interaction behaviors, calculate the total consumption amount in the user's historical consumption records, and evaluate the user's consumption ability in combination with the consumption frequency; The second determination sub-module is used to analyze the viewing time periods in the historical viewing data, determine the time intervals when the user often watches live broadcasts, integrate the user's interest fields, consumption ability levels, and viewing habits, and construct a user portrait.

[0009] Optionally, in the fifth implementation manner of the present invention, the recommendation strategy generation module includes: The establishment sub-module is used to obtain the live broadcast scene parsing result from the scene analysis module, extract the complete dynamic user portrait from the user portrait construction module, and establish an associated mapping between the user's interest fields and the live broadcast scene types; The screening sub-module is used to screen out preliminary candidate recommended contents from the live broadcast platform content library according to the user's interest fields and live broadcast scene types, and filter out the recommended contents that exceed the user's consumption ability; The sorting sub-module is used to sort the candidate recommended contents by using a collaborative filtering algorithm, and determine the push timing of the recommended contents according to the rhythm of the live broadcast scene and the user's viewing habits.

[0010] Optionally, in the sixth implementation manner of the present invention, the sorting sub-module includes: The collection unit is used to collect the user's historical viewing data, interaction behavior data, and candidate recommended content information, and construct a user-content interaction matrix; The traversal unit is used to traverse the user-content interaction matrix, regard the user's interaction vector with the content as a spatial vector by using the cosine similarity, calculate the cosine value of the vector included angle, the closer the value is to 1, the more similar the user interests are, generate a user list with higher similarity for each user, and sort it in descending order according to the similarity score; The calculation unit is used to obtain the users among the similar users who have had interaction behaviors with the candidate content, and calculate the predicted preference score of the target user for the candidate content by weighted calculation according to the interaction intensity of the similar users with the content and the similarity score with the target user; The arrangement unit is used to arrange the candidate recommended contents in descending order from high to low according to the calculated predicted preference scores, and the content with a higher score is arranged in a more forward position; A matching unit, which is used to obtain real-time information of the live broadcast scene from the scene analysis module, determine the viewing habits of users extracted from the user portrait construction module according to the rhythm of the live broadcast scene, and match the live broadcast real-time information with the viewing habit data to determine the push timing of the recommended content.

[0011] Optionally, in the seventh implementation manner of the present invention, a method for implementing a live broadcast scene-aware recommendation system is provided, and the method includes the following steps: Collect multi-source data during the live broadcast in real time, and the multi-source data at least includes live broadcast screen data, host voice data, audience interaction data, and live broadcast platform metadata; Receive and integrate the collected data, and analyze the live broadcast scene; Track the user's behavior trajectory on the live broadcast platform in real time, and combine historical viewing data, collection records, consumption records, and interaction behaviors in the current live broadcast to construct a dynamic user portrait; Generate a personalized recommendation strategy according to the live broadcast scene analysis result and the constructed user portrait; Push the generated recommended content to the user, and the push methods include pop-up display in the sidebar of the live broadcast interface, barrage reminder, and private message push.

[0012] Optionally, in the eighth implementation manner of the present invention, a method for implementing a live broadcast scene-aware recommendation system is provided, and the method includes the following steps: Extract video, audio, and text features, and input them into the latent space encoder of Stable Diffusion respectively to align the multi-modal data at the semantic level and generate the aligned feature vectors; Input the aligned multi-modal feature vectors and the live broadcast platform metadata into the scene classification model to judge the current live broadcast scene type through the scene classification model; Use natural language processing technology to perform word segmentation, part-of-speech tagging, and syntactic analysis on the host voice data and audience interaction data, and mine the products and topics that users are interested in through semantic analysis; Input the processed text into the sentiment analysis algorithm model to judge the sentiment tendency of the live broadcast atmosphere, and extract the scene intention through the LLaMA-3-8B large language model; Integrate multi-modal information, construct a spatio-temporal relationship graph and a scene semantic graph, and determine the association between elements in the live broadcast scene.

[0013] Optionally, in the ninth implementation manner of the present invention, a method for implementing a live broadcast scene-aware recommendation system is provided, and the method includes the following steps: Obtain the user's historical viewing data, collection records, consumption records, and interaction behaviors, and determine the main fields that the user is most interested in based on the live broadcast type and collection records in the historical viewing data; Based on the live interaction behavior, determine the user's points of interest, calculate the total consumption amount in the user's historical consumption records, and evaluate the user's consumption ability in combination with the consumption frequency; Analyze the viewing time period in the historical viewing data, determine the time interval when the user often watches the live broadcast, integrate the user's interest fields, consumption ability level, and viewing habits, and construct a user profile.

[0014] In the technical solution provided by the present invention, multi-source data during the live broadcast process is collected in real time, the collected data is received and integrated, and the live broadcast scenario is analyzed; the user's behavior trajectory on the live broadcast platform is tracked in real time, and combined with the historical viewing data, collection records, consumption records, and interactive behaviors in the current live broadcast, a dynamic user profile is constructed; according to the live broadcast scenario analysis results and the constructed user profile, a personalized recommendation strategy is generated; the generated recommended content is pushed to the user; through multi-source data collection and in-depth scenario analysis, the present invention can accurately grasp the real-time situation of the live broadcast, is no longer limited to traditional static data recommendations, and greatly improves the fit between the recommended content and the live broadcast scenario; combined with the dynamic user profile, fully considering the individual differences of users and their immediate needs in the current live broadcast, realizing personalized and accurate recommendations, effectively improving the user's viewing experience and participation; the intelligent recommendation strategy generation and timely push mechanism not only ensure the timeliness and relevance of the recommended content, but also avoid over-disturbing the user, optimize the interaction process between the user and the live broadcast platform, and help the live broadcast platform improve user stickiness and promote commercial conversion. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.

[0016] Figure 1 Schematic diagram of the first embodiment of the live broadcast scenario perception recommendation system provided by the embodiment of the present invention; Figure 2 Schematic diagram of the second embodiment of the live broadcast scenario perception recommendation system provided by the embodiment of the present invention; Figure 3 Schematic diagram of the third embodiment of the live broadcast scenario perception recommendation system provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] In the description and claims of the present invention and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that shown or described herein. In addition, the term "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0018] The present invention will be specifically described below with reference to the accompanying drawings. As Figure 1 shown, the following is a schematic diagram of the first embodiment of the live broadcast scenario perception recommendation system provided by the embodiment of the present invention. A live broadcast scenario perception recommendation system includes the following modules: The data acquisition module 101 is used to collect multi-source data during the live broadcast in real time. The multi-source data at least includes live broadcast screen data, host voice data, audience interaction data, and live broadcast platform metadata; The scenario analysis module 102 is connected to the data acquisition module, receives and integrates the collected data, and analyzes the live broadcast scenario; The user portrait construction module 103 is used to track the user's behavior trajectory on the live broadcast platform in real time, and construct a dynamic user portrait by combining historical viewing data, collection records, consumption records, and interaction behaviors in the current live broadcast; Specifically, this embodiment further includes a first determination sub-module, which is used to obtain the user's historical viewing data, collection records, consumption records, and interaction behaviors, and determine the main fields that the user is most interested in based on the live broadcast type and collection records in the historical viewing data; The calculation sub-module is used to judge the user's interest points from the live broadcast interaction behaviors, calculate the total consumption amount in the user's historical consumption records, and evaluate the user's consumption ability in combination with the consumption frequency; The second determination sub-module is used to analyze the viewing time period in the historical viewing data, determine the time interval when the user often watches the live broadcast, and integrate the user's interest fields, consumption ability level, and viewing habits to construct a user portrait.

[0019] The recommendation strategy generation module 104 generates a personalized recommendation strategy according to the results of the scenario analysis module and the user portrait constructed by the user portrait construction module; In this embodiment, the recommendation strategy generation module first obtains the analysis results of the live broadcast scene from the scene analysis module, which covers the live broadcast scene type (such as e-commerce live broadcast, game live broadcast, knowledge lecture live broadcast, etc.), scene emotional tendency (enthusiastic, nervous, relaxed, etc.), potential recommendation clues (product information, topic keywords, key events, etc.). At the same time, the dynamic user portrait data is extracted from the user portrait construction module, including user interest areas (such as beauty, sports, technology, etc.), consumption ability level, viewing habits (viewing time period, duration preference), real-time interest preference (generated based on the current live broadcast interactive behavior); according to the rhythm of the live broadcast scene and the user's viewing habits, the timing of pushing the recommended content is determined. For users who watch for a long time and the current live broadcast is in the halftime break, relevant content will be pushed first; avoid pushing disruptive content at key live broadcast nodes (such as game decisive rounds, product rush climaxes), and formulate corresponding display forms and recommendation copy for different types of recommended content (live broadcast, video, goods). For product recommendations, highlight the attributes and discount information that users are interested in. For live broadcast recommendations, emphasize the host's characteristics and live broadcast highlights. Establish an evaluation index system for recommendation strategies, including data indicators such as click-through rate, conversion rate, and user feedback (likes, comments, favorites), monitor user behavior data after the execution of the recommendation strategy in real time, compare the differences between the evaluation indicators and the preset targets, and use reinforcement learning or heuristic algorithms to dynamically adjust the recommendation strategy based on the monitoring results. If a certain type of recommended content has a low click-through rate, reduce its recommendation weight. If users have good feedback on a specific form of recommendation, increase the recommendation ratio of that form, and continuously optimize the effectiveness of the recommendation strategy. Ultimately, output a personalized recommendation strategy that fits user needs and live broadcast scenarios for execution by the recommended content push module.

[0020] See also Figure 2 , a schematic diagram of a second embodiment of the live broadcast scene perception recommendation system provided by an embodiment of the present invention, wherein the scene analysis module includes: The extraction submodule 201 is used to extract video, audio, and text features, and input them into the latent space encoder of Stable Diffusion respectively, so as to align the multimodal data at the semantic level and generate an aligned feature vector; An input submodule 202 is used to input the aligned multimodal feature vector and the live broadcast platform metadata into a scene classification model, and determine the current live broadcast scene type through the scene classification model; The analysis submodule 203 is used to use natural language processing technology to perform word segmentation, part-of-speech tagging and syntactic analysis on the host voice data and the audience interaction data, and to mine products and topics that users are interested in through semantic analysis; The judgment submodule 204 is used to input the processed text into the sentiment analysis algorithm model, judge the sentiment tendency of the live broadcast atmosphere, and extract the scene intention through the LLaMA-3-8B large language model; The construction sub-module 205 is used to synthesize multi-modal information, construct a spatio-temporal relationship graph and a scene semantic graph, and determine the associations between elements in the live broadcast scene.

[0021] Specifically, this embodiment further includes a visual feature unit, which is used to accelerate the decoding of the H.265 encoded stream using WebAssembly for the received video data, generate an RGB frame sequence and optical flow features, and input the RGB frame sequence and optical flow features into the Swin Transformer V2 model to extract visual features in the picture. The audio feature unit is used to separate the human voice and background sound from the received audio data through the Audio Super-Resolution network, and input the separated audio data into the Wav2Vec 3.0 model to learn semantic, emotional, and acoustic features in the audio signal. The text feature unit is used to extract keywords from the received text data through an FPGA-hardware-accelerated BERT tokenizer.

[0022] Specifically, this embodiment further includes a processing unit, which is used to input the aligned multi-modal feature vectors and live broadcast platform metadata into a scene classification model. The scene classification model contains multiple Transformer modules, which process visual and language information respectively. The fusion unit is used to interact and fuse visual features, language features, and live broadcast platform metadata during the calculation of multiple Transformer modules. The generation unit is used to generate a feature representation after the calculation and feature interaction fusion of multiple Transformer modules, calculate the probability score of each possible live broadcast scene type according to the feature representation, and select the scene type with the highest probability score as the judgment result of the current live broadcast scene type.

[0023] In this embodiment, the live video data transmitted by the data acquisition module is received, and the video frame is feature extracted through SwinTransformerV2 to capture the features of visual elements such as objects, scenes, and characters in the picture, and the video feature vector is output; the live audio data is obtained, and the audio signal is processed by Wav2Vec 3.0 to extract semantic, emotional and acoustic features, and generate an audio feature vector; for the host voice text and barrage text, the natural language processing model is used to perform operations such as word segmentation and semantic understanding to refine the text semantic features and obtain the text feature vector; the video, audio, and text feature vectors are respectively input into the latent space encoder of Stable Diffusion, and the encoder maps different modal features to a unified latent space to eliminate the semantic gap and output the aligned multimodal feature vector; the aligned multimodal feature vector and the live platform metadata (live classification, host label, live room popularity, etc.) are collected for format unification and standardization; the integrated data is input into the scene classification model (such as based on ViLBERT The model calculates the probability score of each live scene type (e-commerce live, game live, etc.) based on the input data, and selects the type with the highest probability as the current live scene type judgment result; receives the host's voice data and audience interaction data, uses natural language processing technology to perform word segmentation, and divides the continuous text into single words or phrases; performs part-of-speech tagging on the segmented text to determine the part of speech of each word (noun, verb, etc.), and performs syntactic analysis to parse the grammatical structure of the sentence; through the semantic analysis algorithm, extracts the product names, topic keywords and other information mentioned by the user from the processed text, and mines the products and topics that the user is interested in; the text processed by the analysis submodule is input into the sentiment analysis algorithm model, and the model calculates the sentiment tendency score of the text based on the preset sentiment classification rules (positive, negative, neutral), and determines the sentiment tendency of the live atmosphere; the text is simultaneously input into LLaMA-3-8B Large language model, based on text semantics and context, the model extracts the scene intentions contained in it, such as key information such as promotion nodes and topic transitions; summarizes multi-source information such as video features, audio semantics, text keywords, scene classification results, emotional tendencies, etc.; based on the position changes and time series of objects and characters in the video screen, combined with the time clues in the audio and text information, constructs a graph structure that describes the spatiotemporal relationship of elements in the live broadcast scene; using the results of semantic analysis, associates the products and topics that users are interested in with other elements in the live broadcast scene, forming a scene semantic map containing entities and relationships, and clarifying the intrinsic connection between the elements in the live broadcast scene.

[0024] In this embodiment, the visual feature unit is mainly responsible for processing and feature extraction of live video data. Specifically, first, the WebAssembly technology is used to accelerate the decoding of the H.265 encoded video stream. H.265 is an efficient video coding format that can reduce the data volume while ensuring the picture quality, but the decoding process is relatively complex. WebAssembly provides an execution efficiency close to native performance in environments such as browsers, accelerating the decoding speed, thereby generating an RGB frame sequence and optical flow features. The RGB frame sequence records the color information of each frame of the video picture, and the optical flow features describe the motion information of objects or pixels in the video between consecutive frames. Then, the generated RGB frame sequence and optical flow features are input into the Swin Transformer V2 model, which is an advanced visual processing model. Through its unique window attention mechanism, it can deeply analyze the video picture, capture visual element features such as objects, scenes, and people in it, and finally output the visual feature vector of the video, providing visual-level information support for subsequent multi-modal data processing and scene analysis. The audio feature unit focuses on the processing of live audio data. First, the received audio data is processed using the Audio Super-Resolution network, which can enhance the sound quality of low-quality audio, improve audio clarity, and separate the human voice and background sound in the audio, making the audio data cleaner and more convenient for subsequent analysis. The separated audio data is input into the Wav2Vec 3.0 model, which is a powerful audio processing model. It can learn semantic information (understand the content expressed by the audio), emotional information (judge the emotion conveyed by the audio, such as cheerful, sad, etc.) and acoustic features (such as physical characteristics such as pitch, timbre, volume, etc.) from the audio signal, thereby generating an audio feature vector containing rich audio semantic and emotional information, providing a feature basis in the audio dimension for multi-modal data fusion and scene analysis. The text feature unit mainly targets the text data in the live broadcast, including the text converted from the host's speech and the barrage text sent by the audience, etc. It uses the FPGA hardware-accelerated BERT tokenizer to process the text data. FPGA (Field Programmable Gate Array) has the characteristics of high-speed parallel computing, which can greatly improve the processing speed and achieve keyword extraction with microsecond-level latency. The BERT tokenizer is based on the pre-trained language model BERT and can understand the text semantics, accurately extract the key words and phrases with key meanings from the text. These keywords can reflect the core content of the text and the focus of user attention. The extracted keyword information will be used as text features for subsequent multi-modal data processing and scene analysis. The processing unit undertakes the important task of inputting multi-modal data and live platform metadata into the scene classification model.Before data input, it is necessary to ensure that the aligned multi-modal feature vectors (including multi-dimensional features such as vision, audio, and text) are in the same format and complete as the metadata of the live streaming platform (such as live streaming classification, anchor tags, and popularity of the live streaming room); the scene classification model contains multiple Transformer modules, which are the core components in the fields of natural language processing and multi-modal processing; when processing data, the Transformer modules will encode and process the input visual information and language information respectively, and through technologies such as self-attention mechanism, mine the semantic and structural information in the data, laying a foundation for subsequent feature fusion and scene judgment, which is a key step to achieve accurate scene classification; the fusion unit plays a core role in the operation of the scene classification model; when the Transformer modules calculate and process visual features, language features, and metadata of the live streaming platform, the fusion unit is responsible for promoting the interaction and fusion between these data from different sources; it enables visual features, language features, and metadata to pay attention to and be associated with each other through means such as the attention mechanism; for example, when processing the live streaming scene of selling goods, the fusion unit will guide the model to organically combine the visual features of the goods in the video screen, the language features of the anchor's description of the goods, and the metadata such as the product classification provided by the live streaming platform, mine the potential connections between the data, enable the model to comprehensively understand the live streaming scene from multiple dimensions, so as to generate more representative and accurate feature representations, and provide rich and comprehensive information for scene type judgment; the generation unit is the last link in the scene classification process; after multiple Transformer modules repeatedly calculate and interactively fuse the data, the generation unit will generate a comprehensive feature representation based on the processing results, which integrates information from multiple aspects such as vision, audio, text, and platform metadata, and comprehensively describes the content and background of the current live stream; then, the generation unit will calculate the probability scores of the current live stream belonging to each possible scene type (such as e-commerce live stream, game live stream, music live stream, etc.) according to this feature representation, and by comparing these probability scores, select the scene type with the highest probability as the scene type judgment result of the current live stream, so as to achieve accurate classification of the live streaming scene, providing an important basis for subsequent functions such as generating recommendation strategies based on scenes and analyzing user behavior.;

[0025] Please refer to Figure 3 , the schematic diagram of the third embodiment of the live streaming scene perception recommendation system provided by the embodiment of the present invention, the recommendation strategy generation module includes: The establishment sub-module 301 is used to obtain the live streaming scene parsing result from the scene analysis module, extract the complete dynamic user portrait from the user portrait construction module, and establish the association mapping between the user interest field and the live streaming scene type; The screening sub-module 302 is used to screen out preliminary candidate recommended content from the content library of the live streaming platform according to the user's interest fields and the types of live streaming scenarios, and filter out recommended content that exceeds the user's consumption ability; The sorting sub-module 303 is used to sort the candidate recommended content by using the collaborative filtering algorithm, and determine the pushing time of the recommended content according to the rhythm of the live streaming scenario and the user's viewing habits.

[0026] Specifically, this embodiment further includes a collection unit for collecting the user's historical viewing data, interaction behavior data, and candidate recommended content information to construct a user-content interaction matrix; The traversal unit is used to traverse the user-content interaction matrix, regard the interaction vector of the user with respect to the content as a spatial vector by using the cosine similarity, calculate the cosine value of the vector angle, the closer the value is to 1, the more similar the user interests are, generate a user list with higher similarity to each user, and sort it in descending order according to the similarity score; The calculation unit is used to obtain the users among the similar users who have had interaction behaviors with the candidate content, and calculate the predicted preference score of the target user for the candidate content by weighted calculation according to the interaction intensity of the similar users with respect to the content and the similarity score with the target user; The arrangement unit is used to arrange the candidate recommended content in descending order from high to low according to the calculated predicted preference score, and the content with a higher score is arranged in a more forward position; The matching unit is used to obtain the real-time information of the live streaming scenario from the scenario analysis module, determine the user's viewing habits extracted from the user portrait construction module according to the rhythm of the live streaming scenario, and match the live streaming real-time information with the viewing habit data to determine the pushing time of the recommended content.

[0027] In this embodiment, the core function of the establishment sub-module is to build a connection bridge between user interests and live broadcast scenarios, providing the basic logic for subsequent recommendations. Specifically, this module first obtains the parsed live broadcast scenario results from the scenario analysis module. These results cover multi-dimensional information such as the type of live broadcast (e.g., e-commerce live broadcast, game live broadcast), emotional tendency (enthusiastic, calm), potential recommendation clues (product information, topic keywords), etc. At the same time, it extracts the complete dynamic user portrait from the user portrait construction module, including personalized data such as the user's interest field, consumption ability level, viewing habits, real-time interest preferences, etc. After obtaining the data, the establishment sub-module analyzes historical data to statistically calculate the preference degrees of different user interests in various scenarios, thereby assigning initial association weights. For example, through analysis, it is found that users with a beauty interest have a higher interaction frequency in the e-commerce beauty live broadcast scenario, so a higher association weight is assigned to this interest field and the e-commerce beauty live broadcast scenario. In addition, the module will further adjust the association weight by combining the emotional tendency of the live broadcast scenario and the user's emotional preference (obtained from the user portrait). If the user prefers an enthusiastic atmosphere, when the live broadcast scenario is an enthusiastic promotional live broadcast, the corresponding association weight will be further increased. Finally, through this method, a dynamic association mapping between the user's interest field and the live broadcast scenario type is established, providing a basis for accurate recommendations. The main function of the screening sub-module is to screen out eligible candidate recommended content from the massive live broadcast platform content library according to the user's interests and scenario types, and eliminate options that exceed the user's consumption ability, narrowing the recommendation scope and improving the recommendation efficiency and accuracy. This module first preliminarily screens out content that may match the user's interests from the live broadcast platform content library according to the association mapping between the user's interest field and the live broadcast scenario type generated by the establishment sub-module. For example, for users with a game interest and in a game live broadcast scenario, it screens out the same type of game live broadcasts, game strategy videos, game peripheral products, etc. as candidate recommended content. Then, the screening sub-module 302 will filter the preliminarily screened candidate recommended content in combination with the consumption ability level information in the user portrait. For users with a high consumption ability, the range of recommended content can include high-end products or services with higher prices. For users with a low consumption ability, it will eliminate recommended content that exceeds their consumption level and only retain candidate items that match their consumption ability. Through these two steps of screening, a list of candidate recommended content that is more in line with the user's actual needs and consumption ability is obtained, preparing for subsequent accurate sorting and pushing. The core task of the sorting sub-module is to prioritize the candidate recommended content obtained by the screening sub-module and determine the best pushing time in combination with the live broadcast scenario rhythm and the user's viewing habits to achieve accurate and efficient personalized recommendations. In terms of sorting, the sorting sub-module uses a collaborative filtering algorithm to calculate the similarity or correlation between the candidate recommended content and the user's existing behaviors by analyzing the user's historical behavior data, thereby sorting the candidate content from high to low.The specific process includes steps such as constructing a user-content interaction matrix, calculating the similarity between users, and predicting the preference scores of users for candidate content. Finally, the order of recommended content is determined according to the scores. When determining the push timing, the sorting sub-module will obtain the real-time rhythm information of the live broadcast scene from the scene analysis module, such as whether the live broadcast is in the opening, product introduction, interaction session or end stage, and the frequency of scene changes. At the same time, it combines the user viewing habit data extracted from the user portrait construction module, including information such as the average viewing duration, high-frequency viewing time period, and viewing continuity. By matching the live broadcast scene rhythm with the user viewing habits, the coincidence point between the two is found, so as to determine the best push timing of the recommended content. For example, when the live broadcast is in the e-commerce product introduction session and the user's high-frequency viewing time period coincides with this, it is preferred to push the recommended content at this time. For different types of recommended content (live broadcast, video, product), the sorting sub-module will also formulate differential push timing rules to avoid disturbing users by pushing a large number of recommended content in a short period of time, ensure that the recommended content is displayed during the period when users may be active, and improve the user experience and recommendation effect.

[0028] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A live broadcast scene perception recommendation system, characterized in that The live scene perception recommendation system described above includes: A data collection module, which is used to collect multi-source data during the live broadcast in real time. The multi-source data at least includes live video data, host voice data, audience interaction data, and live platform metadata; A scene analysis module, which is connected to the data collection module, receives and integrates the collected data, and analyzes the live scene; A user profile construction module, which is used to track the user's behavior trajectory on the live platform in real time, and construct a dynamic user profile by combining historical viewing data, collection records, consumption records, and interactive behaviors in the current live broadcast; A recommendation strategy generation module, which generates a personalized recommendation strategy according to the results of the scene analysis module and the user profile constructed by the user profile construction module; A recommended content push module, which pushes the recommended content generated by the recommendation strategy generation module to the user. The push methods include pop-up display in the sidebar of the live broadcast interface, bullet screen reminder, and private message push.

2. The live broadcast scene perception recommendation system according to claim 1, characterized in that, The scene analysis module includes: An extraction sub-module, which is used to extract video, audio, and text features, and respectively input them into the latent space encoder of Stable Diffusion to align multi-modal data at the semantic level and generate aligned feature vectors; An input sub-module, which is used to input the aligned multi-modal feature vectors and live platform metadata into a scene classification model, and judge the current live scene type through the scene classification model; An analysis sub-module, which is used to use natural language processing technology to perform word segmentation, part-of-speech tagging, and syntactic analysis on the host voice data and audience interaction data, and mine the products and topics that users are interested in through semantic analysis; A judgment sub-module, which is used to input the processed text into an emotion analysis algorithm model to judge the emotional tendency of the live broadcast atmosphere, and extract the scene intention through the LLaMA-3-8B large language model; A construction sub-module, which is used to synthesize multi-modal information, construct a spatio-temporal relationship graph and a scene semantic graph, and determine the association between elements in the live scene.

3. The live broadcast scene perception recommendation system according to claim 2, wherein, The extraction sub-module includes: A visual feature unit, which is used to accelerate the decoding of the H.265 encoded stream using WebAssembly for the received video data, generate an RGB frame sequence and optical flow features, and input the RGB frame sequence and optical flow features into the SwinTransformerV2 model to extract visual features in the picture; An audio feature unit, which is used to separate the human voice and background sound from the received audio data through the Audio Super-Resolution network, and input the separated audio data into the Wav2Vec 3.0 model to learn semantic, emotional, and acoustic features in the audio signal; A text feature unit, which is used to extract keywords from the received text data through an FPGA hardware-accelerated BERT tokenizer.

4. The live broadcast scene perception recommendation system according to claim 2, wherein The input sub-module includes: A processing unit, which is used to input the aligned multi-modal feature vectors and live platform metadata into a scene classification model, where the scene classification model contains multiple Transformer modules to process visual and language information respectively; A fusion unit, which is used to interact and fuse visual features, language features, and live platform metadata during the calculation processes of multiple Transformer modules; A generation unit, which is used to generate a feature representation after the calculation and feature interaction and fusion of multiple Transformer modules, calculate the probability scores of each possible live scene type according to the feature representation, and select the scene type with the highest probability score as the judgment result of the current live scene type.

5. The live broadcast scene perception recommendation system according to claim 1, characterized in that, The user portrait construction module includes: A first determination sub-module, which is used to obtain the user's historical viewing data, collection records, consumption records, and interaction behaviors, and determine the main fields that the user is most interested in based on the live types and collection records in the historical viewing data; A calculation sub-module, which is used to judge the user's interest points from the live interaction behaviors, and calculate the total consumption amount in the user's historical consumption records and evaluate the user's consumption ability in combination with the consumption frequency; A second determination sub-module, which is used to analyze the viewing time periods in the historical viewing data, determine the time intervals when the user often watches live broadcasts, and integrate the user's interest fields, consumption ability levels, and viewing habits to construct a user portrait.

6. The live broadcast scene perception recommendation system according to claim 1, wherein The recommendation strategy generation module includes: A establishment sub-module, which is used to obtain the live scene parsing result from the scene analysis module, extract the complete dynamic user portrait from the user portrait construction module, and establish an association mapping between the user's interest fields and the live scene types; A screening sub-module, which is used to screen out preliminary candidate recommended contents from the live platform content library according to the user's interest fields and live scene types, and filter out the recommended contents that exceed the user's consumption ability; A sorting sub-module, which is used to sort the candidate recommended contents by using a collaborative filtering algorithm, and determine the push timing of the recommended contents according to the rhythm of the live scene and the user's viewing habits.

7. The live broadcast scene perception recommendation system according to claim 6, characterized in that, The sorting sub-module includes: A collection unit, which is used to collect the user's historical viewing data, interaction behavior data, and candidate recommended content information, and construct a user-content interaction matrix; A traversal unit, which is used to traverse the user-content interaction matrix, regard the user's interaction vector with the content as a spatial vector by using the cosine similarity, calculate the cosine value of the vector angle, the closer the value is to 1, the more similar the user interests are, generate a list of users with higher similarity for each user, and sort them from high to low according to the similarity scores; A calculation unit, which is used to obtain the users who have interacted with the candidate content among the similar users, and calculate the predicted preference score of the target user for the candidate content by weighted calculation according to the interaction intensity of the similar users with the content and the similarity score with the target user; An arrangement unit, which is used to arrange the candidate recommended contents in descending order from high to low according to the calculated predicted preference scores, and the content with a higher score is ranked in a more forward position; A matching unit, which is used to obtain the real-time information of the live scene from the scene analysis module, determine the user's viewing habits from the user portrait construction module according to the rhythm of the live scene, and match the live real-time information with the viewing habit data to determine the push timing of the recommended content.

8. A method for implementing a live broadcast scene perception recommendation system as described in claim 1, characterized in that, The method includes the following steps: Collect multi-source data during the live broadcast in real time. The multi-source data includes at least live video data, host voice data, audience interaction data, and live platform metadata; Receive and integrate the collected data, and analyze the live broadcast scenario; Track the user's behavior trajectory on the live platform in real time, and combine historical viewing data, collection records, consumption records, and interaction behaviors in the current live broadcast to construct a dynamic user profile; Generate personalized recommendation strategies based on the live broadcast scenario analysis results and the constructed user profile; Push the generated recommended content to the user. The push methods include displaying a pop-up window in the sidebar of the live broadcast interface, sending a barrage reminder, and sending a private message.

9. A method for implementing a live scene perception recommendation system as described in claim 1, characterized in that, The method includes the following steps: Extract video, audio, and text features, and input them into the latent space encoder of Stable Diffusion respectively to align multi-modal data at the semantic level and generate aligned feature vectors; Input the aligned multi-modal feature vectors and live platform metadata into the scene classification model to determine the current live broadcast scene type through the scene classification model; Use natural language processing technology to perform word segmentation, part-of-speech tagging, and syntactic analysis on the host voice data and audience interaction data, and mine the products and topics that users are interested in through semantic analysis; Input the processed text into the sentiment analysis algorithm model to judge the emotional tendency of the live broadcast atmosphere, and extract the scene intention through the LLaMA-3-8B large language model; Integrate multi-modal information, construct a spatio-temporal relationship graph and a scene semantic graph, and determine the associations between elements in the live broadcast scene.

10. A method for implementing a live broadcast scenario perception recommendation system as described in claim 1, characterized in that, The method includes the following steps: Obtain the user's historical viewing data, collection records, consumption records, and interaction behaviors, and determine the main fields that the user is most interested in based on the live broadcast type and collection records in the historical viewing data; Judge the user's interest points from the live broadcast interaction behaviors, calculate the total consumption amount in the user's historical consumption records, and evaluate the user's consumption ability in combination with the consumption frequency; Analyze the viewing time period in the historical viewing data, determine the time interval when the user often watches live broadcasts, and integrate the user's interest fields, consumption ability levels, and viewing habits to construct a user profile.

Citation Information

Patent Citations

  • Live broadcast accurate drainage method and system based on Internet big data analysis

    CN117768665A

  • Live broadcast room content identification and intelligent distribution method and system based on multi-modal fusion

    CN119377895A

  • Live broadcast stream pushing method, device and system, electronic equipment and storage medium

    CN119603487A

  • IPTV service operation management system and program content personalized intelligent recommendation method

    CN119967206A

Cited By

  • Direct broadcasting room automatic operation method, system and equipment using AIGC technology

    CN120455728A

  • Live room automation operation method, system and device using AIGC technology

    CN120455728B

  • Commodity recommendation matching method and system for intelligent live broadcast scene

    CN121481676A

  • Picture book live broadcast personalized book recommendation method and system based on artificial intelligence

    CN121486643A

  • Artificial intelligence-based personalized book recommendation method and system for picture book live broadcast

    CN121486643B