Live broadcast system based on AI interaction
By integrating data collection, emotional judgment, parameter adjustment and style transfer modules in the AI anchor system, dynamic matching between AI anchors and users' emotional needs is achieved, solving the problem of insufficient matching of existing AI anchors in dealing with user emotions, and improving the interactivity and user experience of live broadcasts.
Patent Information
- Application Number
- CN202510473046.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AI anchors have insufficient matching in dealing with users' emotional needs and are unable to accurately capture real-time emotional changes, resulting in a lack of emotional resonance and sense of reality in interactions with users.
A live broadcast system based on AI interaction is designed. The user's real-time barrage data and interactive behavior data are obtained through the data acquisition module. The emotional judgment module judges the user's emotional tendencies. The parameter adjustment module dynamically adjusts the emotional expression parameters of the AI anchor, and the style transfer module generates dynamic expression image data to achieve dynamic matching between the image generation of AI anchors and the emotional needs of users.
It realizes dynamic matching between the image generation of AI anchors and the emotional needs of users, improves the image generation ability of AI anchors and the emotional interaction with users, and enhances the interactiveness and user experience of live broadcasts.
Smart Images

Figure CN120091164A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a live broadcast system based on AI interaction. Background Art
[0002] Against the backdrop of the rapid development of e-commerce live streaming, AI anchors, as an innovative form of live streaming, have gradually become an important tool in the e-commerce industry with their advantages of high efficiency, stability, and freedom from time and space constraints. By using advanced technologies such as style transfer and augmented reality, AI anchors can achieve a high degree of customization in appearance, voice, and movement, thereby meeting the brand's demand for diversified live streaming images. However, although AI anchors have shown great potential in improving live streaming efficiency, in actual business applications, the problem of matching the image generation of AI anchors with the emotional needs of users has become increasingly prominent.
[0003] Specifically, the current image generation of AI anchors mainly relies on style transfer models and preset style templates. Although these technologies can flexibly adjust the appearance and voice characteristics of AI anchors to keep them consistent with the brand image or live broadcast theme, they are unable to cope with user emotional needs. For example, during a live broadcast, users may express their love, doubts, or dissatisfaction with a product through barrages. However, existing AI anchor image generation systems are often unable to accurately capture these real-time emotional changes and dynamically adjust their expressions, tones, and movements accordingly. This results in the AI anchor lacking the necessary emotional resonance and authenticity when interacting with users, which affects the user's viewing experience and willingness to buy.
[0004] In addition, there is also a disconnect between emotion recognition and emotional response when AI anchors respond to barrages. Although existing sentiment analysis algorithms can identify users' emotional tendencies, since emotional response libraries are usually based on template designs, AI anchors' responses often appear mechanical and monotonous. This lack of flexibility and personalization in response further exacerbates the emotional gap between users and AI anchors, reducing user engagement and loyalty. Summary of the invention
[0005] The purpose of the present invention is to provide a live broadcast system based on AI interaction, which realizes the dynamic matching between the image generation of AI anchor and the emotional needs of users, not only improves the image generation ability of AI anchor, but also significantly enhances the emotional interaction effect between it and users, so as to solve at least one of the above-mentioned prior art problems.
[0006] The present invention discloses a live broadcast system based on AI interaction, and the system specifically comprises: The data collection module is used to obtain the real-time bullet screen data and interactive behavior data of users during the live broadcast; An emotion judgment module, which is used to determine the judgment result of the user's emotional tendency state during the live broadcast through the real-time barrage data and the interactive behavior data; A parameter adjustment module, which is used to dynamically adjust the emotional expression parameters of the AI anchor based on the judgment result of the emotional tendency state, according to the live broadcast content theme and the key information of commodity recommendation; A style transfer module, which is used to input the emotional expression parameters into a style transfer model, generate dynamic expression image data, and map the dynamic expression image data to the image model of the AI anchor.
[0007] Compared with the prior art, the present invention has at least one of the following technical effects: 1. The present invention realizes the dynamic matching between the generation of the AI anchor image and the emotional needs of users, not only improves the image generation ability of the AI anchor, but also significantly enhances the emotional interaction effect between it and users.
[0008] 2. By analyzing the barrage data and interactive behavior data in detail, the present invention can accurately capture the emotional changes of users during the live broadcast and identify key emotional nodes, which helps the AI anchor to understand the user's emotions more accurately, so as to make more considerate responses and adjustments, and improve the user experience.
[0009] 3. The present invention adopts an emotion classification template, user feature association matching and an emotion recognition algorithm model, which can analyze the user's emotions more meticulously, obtain an accurate emotional tendency value, provide reliable data support for the emotional expression of the AI anchor, and enable it to more realistically simulate human emotional reactions.
[0010] 4. By integrating the emotion vector, the live broadcast content theme vector and the commodity recommendation weight vector, the present invention can dynamically adjust the emotional expression parameters of the AI anchor, not only considering the emotional needs of users, but also combining the live broadcast content and the commodity recommendation strategy, making the response of the AI anchor closer to the user's expectations, and improving the conversion rate and user satisfaction of the live broadcast.
[0011] 5. The present invention uses a generative adversarial network model to generate dynamic expression image data and maps it to the image model of the AI anchor, which not only improves the richness and authenticity of the AI anchor's expressions, but also enables it to make instant adjustments according to the emotional changes of users, enhancing the interactivity and immersion of the live broadcast.
[0012] 6. Through real-time script generation, multi-modal material matching and crisis handling decision-making, the present invention realizes the intelligent generation and dynamic adjustment of the live broadcast content, not only improving the richness and attractiveness of the live broadcast content, but also being able to process sensitive content in a timely manner to ensure the compliance and security of the live broadcast.
[0013] 7. Through dynamic reward calculation, interactive policy network, and meta-learning adaptation, the present invention realizes intelligent interaction between the AI host and users, can dynamically adjust the live broadcast rhythm and the response method of the AI host according to real-time user behavior data, and improves user participation and interactivity.
[0014] 8. The design of the multi-objective reward function of the present invention enables the live broadcast interaction decision-making module to comprehensively evaluate according to multiple objectives (such as user stay duration, conversion rate, and interaction frequency), and generate a standardized reward value, which helps the AI host to balance between multiple objectives and maximize the overall benefit.
[0015] 9. Through dynamic user portrait modeling, multi-modal knowledge graph, and cognitive bias correction, the present invention realizes in-depth understanding and accurate grasp of user needs, which not only helps the AI host to provide more personalized services, but also can timely discover and correct user understanding biases, and improve the accuracy and effectiveness of the live broadcast. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 FIG. is a schematic structural diagram of a live broadcast system based on AI interaction provided by the first embodiment of the present invention; Figure 2 FIG. is a schematic structural diagram of a live broadcast system based on AI interaction provided by the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0019] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0020] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0021] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.
[0022] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for differential description and cannot be understood as indicating or implying relative importance.
[0023] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0024] Figure 1 The structural schematic diagram of the live broadcast system based on AI interaction disclosed in one embodiment of the present invention is shown as follows: The data acquisition module 101 is used to obtain the real-time barrage data and interaction behavior data of users during the live broadcast; The emotion judgment module 102 is used to determine the judgment result of the emotion tendency state of the user during the live broadcast through the real-time barrage data and the interaction behavior data; The parameter adjustment module 103 is used to dynamically adjust the emotion expression parameters of the AI anchor based on the judgment result of the emotion tendency state and according to the live broadcast content theme and the key information of product recommendation; The style transfer module 104 is used to input the emotion expression parameters into the style transfer model, generate dynamic expression image data, and map the dynamic expression image data to the image model of the AI anchor.
[0025] In this embodiment, the data collection module is used to integrate a bullet screen collection interface and an interactive behavior tracking script on the live streaming platform. The bullet screen data includes the text content, timestamp, etc. sent by users; the interactive behavior data includes likes, shares, comments, purchase clicks, etc.
[0026] The emotion judgment module is used to train an emotion analysis model using natural language processing (NLP) technology and machine learning algorithms. This model can analyze the emotion tendency (positive, negative, neutral) and intensity in the bullet screen text. Input the real-time bullet screen data and interactive behavior data into the emotion analysis model to obtain the judgment result of the user's emotion tendency state during the live broadcast. For example, if the positive emotion words in the bullet screen increase and the number of likes rises, it is judged that the user's emotion tendency is positive.
[0027] The parameter adjustment module is used to set a series of emotion expression parameter adjustment rules according to the live broadcast content theme and key information of product recommendations. Based on the emotion tendency state judgment result and the preset rules, dynamically adjust the emotion expression parameters of the AI anchor, including expression intensity, speech rate, intonation, etc. For example, when the live broadcast content is a new product release and the user's emotion tendency is positive, increase the excitement and anticipation expression parameters of the AI anchor.
[0028] The style transfer module is used to select a suitable style transfer model, such as a generative adversarial network (GAN) or a cycle generative network (CycleGAN), to convert the emotion expression parameters into dynamic expression image data. Input the adjusted emotion expression parameters into the style transfer model to generate dynamic expression image data that matches the image model of the AI anchor. Then, map these image data to the image model of the AI anchor to achieve dynamic changes in the AI anchor's expressions.
[0029] In this embodiment, by analyzing the user's emotion in real time and adjusting the performance of the AI anchor, it is possible to better attract the user's attention, improve the user's participation and interactivity. The AI anchor can show more scene-appropriate emotion expressions according to different live broadcast content themes and key information of product recommendations, thereby enhancing the overall effect of the live broadcast.
[0030] In some embodiments, determining the judgment result of the user's emotion tendency state during the live broadcast through the real-time bullet screen data and the interactive behavior data specifically includes: Adopt a text analysis algorithm to perform emotion analysis on the bullet screen data at each time point to obtain the emotion tendency value at each time point; According to the type and frequency of the interactive behavior data, combined with the emotion tendency value, calculate the comprehensive emotion index of the user at a specific time point; Mark the specific time points with the comprehensive sentiment index exceeding the preset sentiment index threshold as key sentiment nodes, and associate each key sentiment node with the live content to determine the key factors affecting the emotional changes of users.
[0031] In this embodiment, a text analysis algorithm is used to perform sentiment analysis on the barrage data to obtain the sentiment tendency value at each time point. According to the type and frequency of the interaction behavior data, combined with the sentiment tendency value, the comprehensive sentiment index of the user at a specific time point is calculated. For the barrage data at each time point, the number of barrages and the number of behaviors are extracted as input parameters for calculating the comprehensive sentiment index. The level of the sentiment tendency value is judged through a preset threshold. If the sentiment tendency value is higher than the threshold, the weight of the comprehensive sentiment index is increased. According to the behavior type value and the behavior frequency value, the calculation formula of the comprehensive sentiment index is adjusted to obtain a more accurate comprehensive sentiment index. If the comprehensive sentiment index exceeds the preset sentiment index threshold, then this time point is marked as a key sentiment node. For each key sentiment node, the live content of the corresponding time period is extracted. A text analysis method is used to identify the keywords related to emotional fluctuations from the live content. According to the appearance frequency of the keywords, the main factors affecting the emotional changes of users are determined. An emotional change trend model is constructed, and the key nodes and influencing factors are analyzed for correlation. The main factors affecting the emotional changes of users and the analysis results of their relevance to the live content are obtained.
[0032] Exemplarily, for a barrage like "The anchor is awesome", it may be given a relatively high positive sentiment value; while "This is too bad" may get a relatively low negative sentiment value. Such analysis can help understand the emotional changes of the audience during the live broadcast. The interactive behavior data includes user activities such as liking, rewarding, and following. Different types of interactive behaviors may be assigned different weights. For example, rewarding may be considered a stronger positive emotional expression than liking. By combining these interactive data with the sentiment tendency values, a more comprehensive comprehensive sentiment index can be calculated. This index can more accurately reflect the true emotional state of users. Setting a sentiment index threshold is to identify particularly prominent emotional fluctuations. Assuming the threshold is set to 0.8 (range -1 to 1), then when the comprehensive sentiment index exceeds 0.8, that time point is marked as a key emotional node. These nodes usually correspond to exciting or controversial moments in the live broadcast. Associating the key emotional nodes with the live broadcast content is an important step. For example, if at a certain time point, the comprehensive sentiment index suddenly soars to 0.9, then what happened at that moment will be carefully checked. It may be that the anchor demonstrated amazing skills or announced exciting news. Through this association, the key factors affecting the emotional changes of users can be determined. The significance of this analysis method is that it can help the live broadcast platform and the anchor better understand the audience's reactions. By identifying the content that triggers strong emotional reactions, the anchor can adjust their performance strategies, and the platform can also optimize content recommendations. For example, if the data shows that the audience reacts particularly enthusiastically to the anchor's talent shows, then the anchor may consider increasing the proportion of such content. In addition, this method can also be used to identify potential problems. If the negative sentiment index suddenly soars at a certain time point, it may mean that there are inappropriate remarks or behaviors, and the platform can intervene in a timely manner. This not only helps maintain a good live broadcast environment but also improves user satisfaction and loyalty.
[0033] Further, the text analysis algorithm is used to perform sentiment analysis on the barrage data at each time point to obtain the sentiment tendency value at each time point, specifically including: Based on a preset sentiment classification template, sentiment words are extracted from the barrage data at each time point; Obtain the user characteristics corresponding to each barrage, and associate and match the sentiment words with the corresponding user characteristics to form a preliminary sentiment tendency; Through an emotion recognition algorithm model, the preliminary sentiment tendency is classified by emotion to obtain the sentiment tendency value at each time point.
[0034] In this embodiment, a preset sentiment classification template is adopted to extract sentiment words from the barrage data at each time point and determine the categories of the sentiment words. The user characteristics corresponding to each barrage are obtained, including the user's age, gender, and regional information. The extracted sentiment words are associated and matched with the user characteristics to form a preliminary sentiment tendency. The preliminary sentiment value of the time point is obtained according to the preliminary sentiment tendency, and the preliminary sentiment value is classified through an emotion recognition algorithm model to obtain a sentiment tendency value. A preset sentiment classification model is used to process the sentiment tendency value to determine whether the sentiment tendency value exceeds a preset threshold. If it exceeds, the sentiment tendency value is determined as a strong sentiment value. The corresponding time point is obtained according to the strong sentiment value, and the time points are sorted through a time series analysis method to obtain a sentiment fluctuation time series. A clustering algorithm is used to process the sentiment fluctuation time series to determine whether the sentiment fluctuation time series has clustering characteristics. If it does, a sentiment fluctuation pattern is determined. The corresponding algorithm model parameters are obtained according to the sentiment fluctuation pattern, and the algorithm model parameters are adjusted through a parameter optimization algorithm to obtain an optimized emotion recognition model. The optimized emotion recognition model is used to classify the preliminary sentiment tendency to determine whether the preliminary sentiment tendency matches the optimized model. If it matches, the final sentiment value is determined. The corresponding time point is obtained according to the final sentiment value, and a sentiment change curve is generated through the mapping relationship between the time point and the sentiment value to obtain a sentiment analysis result.
[0035] Exemplarily, an emotion classification template usually contains multiple emotion categories, such as happiness, sadness, anger, etc. Each category has corresponding emotion words. For example, "happy" and "joyful" belong to the happiness category, and "disappointed" and "frustrated" belong to the sadness category. In actual applications, the template may be more complex, including the influence of degree adverbs and negative words. Extracting emotion words from the bullet comments is a key step. Suppose there is a bullet comment "The host is so great, and the program is extremely wonderful". The system will identify two positive emotion words, "great" and "wonderful". This process usually uses word segmentation technology and part-of-speech tagging to accurately locate emotion words. The acquisition and association of user characteristics are crucial for emotion analysis. User characteristics may include user level, activity, historical interaction records, etc. For example, a positive comment sent by a highly active and high-level user for a long time may be given a higher weight. This is because such users usually have a better understanding of the platform, and their opinions may be more representative. Matching emotion words with user characteristics to form a preliminary emotion tendency is a comprehensive process. For example, for a comment like "The program is really good-looking", it may be given a medium positive tendency if it comes from a new user, while it may be regarded as highly positive if it comes from a senior user. This differential processing helps to more accurately capture the true emotions of users. The emotion recognition algorithm model is the core of the entire analysis process. Such models usually rely on machine learning technologies, such as support vector machines or deep learning networks. The model will consider multiple factors such as words, context, and user characteristics. For example, a comment like "This is too exaggerated" may express surprise or dissatisfaction in different contexts, and the model needs to combine the context to accurately judge. The finally obtained emotion tendency value is usually a numerical value, such as ranging from -1 to 1, where a negative value indicates a negative emotion and a positive value indicates a positive emotion.
[0036] In some embodiments, based on the judgment result of the emotion tendency state, according to the live content theme and the key information of product recommendation, dynamically adjust the emotion expression parameters of the AI host, specifically including: According to the judgment result of the emotion tendency state, generate an emotion vector representation, and the emotion vector representation satisfies , where v represents the emotion vector representation, represents a number of emotion feature vectors, and MLP represents a multi-layer perceptron; Determine the live content theme vector through a theme classification model, and the theme classification model satisfies , where t represents the live content theme vector, represents the live content theme vector of the content theme category y and the live text content x, y represents the content theme category, x represents the live text content, represents model, and W and b represent learnable parameters; Determine the product recommendation weight vector through a product recommendation weight calculation model, and the product recommendation weight calculation model satisfies , where represents the recommendation weight of product i, represents the click-through rate, represents the conversion rate, represents the relevance score to the theme of the live content, , and respectively represent the weights of the click-through rate, conversion rate, and relevance score to the theme of the live content; According to the emotional vector representation, the live content theme vector, and the product recommendation weight vector, generate the emotional expression parameters of the AI anchor, and the emotional expression parameters satisfy , where represents the emotional expression parameter, v represents the emotional vector representation, t represents the live content theme vector, and w represents the product recommendation weight vector.
[0037] In this embodiment, the sentiment judgment module (such as a sentiment analysis model based on deep learning) processes the real-time barrage data and interaction behavior data to obtain the sentiment tendency state judgment result. Then, use a multi-layer perceptron (MLP) to map the sentiment tendency state to a high-dimensional space to generate the emotional vector representation v. The emotional vector v consists of several emotional feature vectors, and each feature vector represents the intensity of an emotional feature (such as happiness, sadness, anger, etc.). The input layer of the MLP receives the original representation of the sentiment tendency state (such as the category and intensity of the sentiment tendency), and the output layer generates the emotional vector v. The middle layer of the MLP includes multiple hidden layers for extracting and combining emotional features.
[0038] Use a theme classification model (such as a text classification model based on a convolutional neural network or a recurrent neural network) to process the live text content x to obtain the live content theme vector t. The output layer of the theme classification model corresponds to different content theme categories y, and each category has a corresponding vector representation. The model learns the mapping relationship from the live text content x to the content theme category y through training. During the live broadcast, the model receives the text content x in real time and outputs the corresponding theme vector t.
[0039] According to the click-through rate, conversion rate, and relevance score of the product to the theme of the live content, use the product recommendation weight calculation model to calculate the recommendation weight of each product. The model comprehensively considers multiple factors and generates a weight vector for each product. The weight calculation model may include one or more linear layers or non-linear layers for combining and processing input features such as click-through rate, conversion rate, and relevance score. The finally output weight vector is used to represent the priority of the product in the recommendation list.
[0040] Taking the emotion vector representation v, the live content theme vector t, and the product recommendation weight vector w as inputs, a fusion model (such as another MLP or attention mechanism model) is used to generate the emotion expression parameters of the AI anchor. These parameters may include expression intensity, speech rate, intonation, etc., which are used to guide the performance of the AI anchor during the live broadcast. The fusion model learns through training how to integrate information from different sources (emotion, theme, recommendation weight) to generate effective emotion expression parameters. These parameters are updated in real time during the live broadcast to adapt to different situations and requirements.
[0041] In this embodiment, by analyzing the user's emotion in real time and adjusting the performance of the AI anchor, the user can feel a more real and personalized interactive experience during the live broadcast, which helps to enhance the user's participation and loyalty. The AI anchor can show more scene-appropriate emotion expressions according to different live content themes and key product recommendation information, thus improving the overall effect of the live broadcast. This helps to attract more viewers and extend their viewing time.
[0042] In some embodiments, inputting the emotion expression parameters into a style transfer model to generate dynamic expression image data and mapping the dynamic expression image data to the image model of the AI anchor specifically includes: Using a generative adversarial network model to train and generate a style transfer model. Through the style transfer model, the emotion expression parameters are converted into dynamic expression image data. The style transfer model includes a generator, a discriminator, and a loss function. The generator is used to convert the emotion expression parameters into dynamic expression image data. The discriminator is used to detect the dynamic expression image data and output an image authenticity score. The loss function satisfies , where represents the total loss, represents the adversarial loss, represents the perceptual loss, represents the style loss, and respectively represent the weights of the perceptual loss and the style loss; Using a key point detection algorithm to determine the key point coordinates of the dynamic expression image data, and mapping the key point coordinates to the image model of the AI anchor through a deformation algorithm.
[0043] In this embodiment, the generator is a deep neural network, whose input is emotional expression parameters (such as expression intensity, speech rate, intonation, etc.) and output is dynamic expression image data. The goal of the generator is to generate dynamic expression images similar to real expression images. The discriminator is also a deep neural network, whose input is dynamic expression image data (which may be generated by the generator or real), and output is an image authenticity score. The goal of the discriminator is to distinguish between generated images and real images. The total loss consists of adversarial loss, perceptual loss, and style loss. The adversarial loss is used to encourage the generator to generate images that can deceive the discriminator; the perceptual loss is used to maintain the consistency of the generated images and real images in high-level features; the style loss is used to capture and maintain the style features of the generated images. The weights and are used to balance these three parts of the loss. By alternately optimizing the parameters of the generator and the discriminator, the total loss is gradually reduced until the training stop condition is reached (such as reaching the preset number of iterations or loss convergence).
[0044] After training is completed, the trained generator is used to convert emotional expression parameters into dynamic expression image data. These image data contain expression features corresponding to the emotional expression parameters. Use a keypoint detection algorithm (such as OpenPose or Dlib, etc.) to process the generated dynamic expression image data to determine the keypoint coordinates in the image (such as the positions of feature points like eyes, mouth, nose, etc.). According to the detected keypoint coordinates, use a deformation algorithm (such as affine transformation, mesh deformation, etc.) to map the expression features in the dynamic expression image data to the image model of the AI anchor. This step ensures that the generated dynamic expressions are consistent with the image of the AI anchor and can change dynamically with the change of emotional expression parameters.
[0045] Using a keypoint detection algorithm to determine the keypoint coordinates of the dynamic expression image data and mapping the keypoint coordinates to the image model of the AI anchor through a deformation algorithm specifically includes: using a keypoint detection algorithm to process the image data to obtain keypoint coordinate values. For the keypoint coordinate values, use a deformation algorithm to process them to obtain the deformed coordinate values. According to the deformed coordinate values, obtain the image model data of the AI anchor. Map the deformed coordinate values to the image model data through a mapping algorithm to obtain the mapped image model. According to the mapped image model, obtain the final image model data. If there is an abnormality in the image model data, use a preset threshold to make a judgment to obtain the corrected image model data. According to the corrected image model data, obtain the final image model.
[0046] In this embodiment, through the training of the generative adversarial network model, the generator can generate dynamic expression image data highly similar to real expression images, improving the expression authenticity of the AI anchor in the live broadcast and enhancing the immersion and substitution sense of the audience. The style transfer model can capture and maintain the style features of the generated images, enabling the AI anchor to show more diverse expression changes, which helps to improve the interactivity and interest of the live broadcast. By combining the key point detection algorithm and the deformation algorithm, an accurate mapping from the dynamic expression image data to the AI anchor image model is achieved, ensuring that the generated dynamic expressions are consistent with the image of the AI anchor and can dynamically change with the change of the emotional expression parameters, thereby improving the expressiveness and flexibility of the AI anchor.
[0047] In some embodiments, the system further includes: The intelligent live content generation module 105, including a real-time script generation unit 1051, a multi-modal material matching unit 1052, and a crisis handling decision-making unit 1053; The real-time script generation unit 1051 is used to analyze multi-source hot topics according to the heat evaluation model and dynamically generate a live broadcast script. The heat evaluation model satisfies , where represents the comprehensive heat value of the topic at the current time point , represents the current time point, represents the real-time weight of the jth topic, represents the time decay coefficient, represents the latest update timestamp of topic j, and n represents the total number of topics; The multi-modal material matching unit 1052 is used to semantically associate and match the product description with the 3D model and the demonstration animation through the cross-modal similarity calculation model. The cross-modal similarity calculation model satisfies , where represents the matching score between the query content Q and the material M, represents the text embedding vector, represents the visual embedding vector, and represent the modal weight coefficients; The crisis handling decision-making unit 1053 is used to detect sensitive content in the live broadcast script according to the sensitivity scoring model. The sensitivity scoring model satisfies , where represents the comprehensive risk score of the statement s, represents the rule matching score of the statement s based on the sensitive word library, represents the neural network semantic risk prediction value of the statement s, represents the mixed weight coefficient.
[0048] In this embodiment, a popularity evaluation model is constructed based on historical data and real-time data. This model is used to analyze multi-source popularity topics, including topics from channels such as social media, news websites, and user comments. The core of the model is to calculate the comprehensive popularity value of the topic at the current time point, which is jointly determined by the real-time weight of each topic, the time decay coefficient, and the latest update timestamp. During the live broadcast, the real-time script generation unit continuously obtains new popularity topics and dynamically generates a live broadcast script related to the topic according to the results of the popularity evaluation model. These scripts can include opening remarks, topic introductions, product introductions, etc., to ensure that the live broadcast content is closely connected to the current hotspots and attracts user attention.
[0049] To semantically associate and match product descriptions with materials such as 3D models and demonstration animations, a cross-modal similarity calculation model is constructed. This model uses text embedding vectors and visual embedding vectors to convert product descriptions and materials into vector representations in a high-dimensional space, and evaluates the degree of semantic association between them by calculating the similarity scores between these vectors. During the live broadcast, when a certain product needs to be displayed, the multi-modal material matching unit selects the most matching 3D model, demonstration animation and other materials from the material library according to the product description. This can not only improve the visual effect of the live broadcast, but also help users better understand the features and usage methods of the product.
[0050] To detect sensitive content in the live broadcast script, a sensitivity scoring model is constructed. This model combines a rule matching method based on a sensitive word library and a semantic risk prediction method based on a neural network to comprehensively score each statement in the live broadcast script. The higher the score, the greater the possibility that the statement contains sensitive content. During the live broadcast, the crisis handling decision-making unit will monitor the live broadcast script in real time and, according to the results of the sensitivity scoring model, give early warnings or replace statements containing sensitive content. This helps to avoid inappropriate remarks or sensitive topics during the live broadcast and protects the reputation of the live broadcast and the rights and interests of users.
[0051] In this embodiment, through the real-time script generation unit, a live broadcast script can be dynamically generated according to current hot topics, making the live broadcast content more in line with user needs and interests, and improving the timeliness and attractiveness of the live broadcast. The multi-modal material matching unit can automatically select materials that match the product description, improving the visual effect and user experience of the live broadcast. At the same time, this also reduces the burden on the host to search for and select materials during the live broadcast. The crisis handling decision-making unit can monitor sensitive content in the live broadcast script in real time and take corresponding handling measures to avoid inappropriate remarks or sensitive topics during the live broadcast, ensuring the compliance and security of the live broadcast. The entire intelligent live broadcast content generation module realizes the intelligent generation and optimization of live broadcast content by integrating multiple intelligent units, improving the intelligence and automation level of the live broadcast. This helps to reduce the labor cost and time cost of the live broadcast, and improve the efficiency and effect of the live broadcast.
[0052] In some embodiments, the system further includes: A live broadcast interaction decision-making module 106, including a dynamic reward calculation unit 1061, an interaction strategy network unit 1062, and a meta-learning adaptation unit 1063; The dynamic reward calculation unit 1061 is used to construct a multi-objective reward function based on real-time user behavior data. Based on the multi-objective reward function, a reward signal is calculated through a sliding window mechanism and compared with a historical baseline value to generate a normalized reward value. The real-time user behavior data includes the user's stay duration, conversion rate, and interaction frequency; The interaction strategy network unit 1062 is used to generate live broadcast rhythm instructions using an LSTM network, and decompose the live broadcast rhythm instructions into micro-expression parameters, voice synthesis parameters, and body movement parameters; The meta-learning adaptation unit 1063 is used to adapt the micro-expression parameters, the voice synthesis parameters, and the body movement parameters to the image model of the AI host using a parameter space mapping matrix.
[0053] Furthermore, the multi-objective reward function satisfies , where represents the comprehensive reward value at time point p, represents the weight coefficient of the kth objective, ( ) represents the Sigmoid function, represents the current objective value, represents the historical baseline mean, represents the standard deviation, m represents the total number of objectives, and objective k includes the user's stay duration, conversion rate, and interaction frequency.
[0054] In this embodiment, a multi-objective reward function is constructed based on real-time user behavior data (including user dwell time, conversion rate, and interaction frequency). This function aims to comprehensively evaluate the user's behavior performance during the live broadcast and provide a basis for subsequent reward calculation. Through a sliding window mechanism, the user's behavior data within each time window is calculated in real-time, and corresponding reward signals are generated according to the multi-objective reward function. These reward signals reflect the user's real-time behavior performance during the live broadcast. The calculated reward signals are compared with the historical baseline values to generate normalized reward values. The normalized reward values are used to measure the degree of improvement or decline of the user during the live broadcast relative to the historical baseline values, providing a reference for subsequent interaction strategy adjustment.
[0055] An LSTM (Long Short-Term Memory) network is used to train a network model for generating live broadcast rhythm instructions based on historical live broadcast data and user behavior data. This model can capture the time dependence during the live broadcast and the trend of user behavior changes. During the live broadcast, the interaction strategy network unit will generate corresponding live broadcast rhythm instructions according to the current user behavior data and the prediction results of the LSTM network model. These instructions include micro-expression parameters, speech synthesis parameters, and body movement parameters, etc., which are used to guide the performance of the AI host during the live broadcast.
[0056] To adapt the generated live broadcast rhythm instructions to the AI host's image model, a parameter space mapping matrix is constructed. This matrix can convert micro-expression parameters, speech synthesis parameters, and body movement parameters, etc., into a parameter format recognizable by the AI host's image model. Using the parameter space mapping matrix, the generated live broadcast rhythm instructions are adapted to the AI host's image model. This enables the AI host to adjust its micro-expressions, speech, and body movements, etc., according to the live broadcast rhythm instructions, so as to interact with users more naturally.
[0057] In this embodiment, through the cooperation of the dynamic reward calculation unit and the interaction strategy network unit, the real-time analysis of user behavior and the intelligent adjustment of the live broadcast rhythm are realized, which helps to improve the intelligent level of live broadcast interaction, enabling the AI host to better adapt to user needs and behavior changes. By using the meta-learning adaptation unit to adapt the live broadcast rhythm instructions to the AI host's image model, the AI host can show a more natural and vivid performance during the live broadcast, which helps to enhance the attractiveness and user participation of the live broadcast, and improve the conversion rate and interaction frequency of the live broadcast. The entire live broadcast interaction decision module realizes the comprehensive optimization of the live broadcast interaction process by integrating multiple intelligent units, which helps to reduce the labor cost and time cost of the live broadcast, improve the operation efficiency and effect of the live broadcast. At the same time, by real-time analyzing and adjusting the live broadcast rhythm instructions, problems in the live broadcast can be discovered and solved in a timely manner, improving the quality and user experience of the live broadcast.
[0058] In some embodiments, the system further includes: The cognitive graph construction module 107 includes a dynamic user portrait modeling unit 1071, a multi-modal knowledge graph unit 1072, and a cognitive bias correction unit 1073; The dynamic user portrait modeling unit 1071 is used to integrate the user's historical behavior data, cross-platform social relationship data, and consumption preference data in real time, and dynamically generate user feature vectors through a temporal graph neural network; The multi-modal knowledge graph unit 1072 is used to analyze commodity attributes, extract live broadcast topics, and mine user interest points to construct a heterogeneous data association graph; The cognitive bias correction unit 1073 is used to identify the user's understanding bias according to the user feature vector and the heterogeneous data association graph, and generate a corrected content expression scheme.
[0059] In this embodiment, a temporal graph neural network is used to obtain the user's historical behavior data, cross-platform social relationship data, and consumption preference data, and generate dynamic user feature vectors. The multi-modal knowledge graph technology is used to analyze commodity attributes, extract live broadcast topics, mine user interest points, and construct a heterogeneous data association graph. According to the dynamic user feature vector and the heterogeneous data association graph, the attention mechanism is used to identify the user's understanding bias. If the user's understanding bias is identified, a corrected content expression scheme is generated through natural language processing technology. The corrected content expression scheme is semantically matched with the live broadcast topic to determine the content suitability. According to the content suitability, a collaborative filtering algorithm is used to optimize the corrected content expression scheme. The optimized corrected content expression scheme is associated with the user interest points to obtain the final content recommendation scheme.
[0060] The dynamic user portrait modeling unit is used to collect historical behavior data through behaviors such as user login, browsing, purchasing, and evaluation; at the same time, obtain the user's relationship data on social media through a third-party API, such as following, liking, commenting, etc.; in addition, the user's consumption preference data is also collected, such as purchase frequency, brand preference, price sensitivity, etc. These data are processed using a temporal graph neural network (TGN). TGN can capture the changing trends of user behavior over time, such as the transfer of interest points and the changes in social circles. Through training, TGN generates the feature vectors of each user, and these vectors contain both static attributes (such as age, gender) and dynamic attributes (such as recent shopping tendencies, social activity levels).
[0061] The multi-modal knowledge graph unit is used to analyze the key attributes in product titles and descriptions, such as brand, material, function, etc., by using natural language processing (NLP) technology. It performs text analysis on live content to identify and extract popular topics and keywords, such as "new autumn and winter styles", "limited-time discounts", etc. Combining the user's historical behavior and social data, it uses clustering algorithms to identify the common interests of user groups, such as "technology product enthusiasts", "fashion trend followers". The above information is integrated into a heterogeneous knowledge graph, where the nodes in the graph include users, products, topics, etc., and the edges represent the relationships between them, such as "user A purchased product B", "topic C attracted user D".
[0062] The cognitive bias correction unit is used to analyze the possible understanding biases of users towards product information based on user feature vectors and the heterogeneous data association graph. For example, users may have cognitive biases due to information asymmetry (such as not knowing a new brand) or misunderstanding (such as exaggerating the function of a certain product). For the identified biases, a correction content expression plan is generated. This may include adjusting the language style of product descriptions, increasing the transparency of key information (such as detailed material descriptions), providing user evaluation comparisons, etc., to help users understand product information more accurately.
[0063] In this embodiment, through the dynamic user profile and the multi-modal knowledge graph, the AI anchor can understand users more deeply, achieve more accurate personalized recommendations and content displays, and improve user satisfaction and stickiness. The cognitive bias correction mechanism helps to reduce information asymmetry and misunderstandings, enhance users' trust in the AI anchor, and promote transaction completion.
[0064] Exemplarily, the dynamic user profile modeling unit generates user feature vectors in real time by integrating multi-source data. For example, this unit may analyze a user's activity on social media, browsing records on shopping platforms, and credit card consumption data. Through the temporal graph neural network, it can capture the changing trends of user interests. For example, the system may find that the user's interest in outdoor sports products has increased sharply recently, which may be related to their recent joining of a hiking group.
[0065] The multi-modal knowledge graph unit is committed to building a comprehensive information network. It not only analyzes the basic attributes of products, such as price, brand, material, etc., but also extracts key topics from live content. For example, in a beauty live broadcast, the system may identify popular topics such as "sun protection", "moisturizing", "concealer". At the same time, it can also discover the potential interests of users, such as finding that users are particularly interested in cosmetics with organic ingredients by analyzing browsing history. This information is integrated into a heterogeneous data association graph, laying the foundation for subsequent personalized recommendations.
[0066] The cognitive bias correction unit acts as an "error corrector". It comprehensively analyzes the user feature vector and the heterogeneous data association graph to identify possible understanding biases of the user. For example, if the system discovers that the user frequently searches for certain products claiming to "quickly lose weight", but these products are marked as "dubious efficacy" or "high potential risk" in the professional medical knowledge graph, the system will generate corresponding correction content expression schemes. This may include pushing some relevant articles on scientific weight loss, or inserting expert opinions as reminders when the user browses similar products.
[0067] These three units work together, not only being able to accurately capture the user's interests and needs, but also effectively preventing the formation of information cocoons. Through dynamic adjustment and cognitive correction, the system can provide users with more comprehensive and objective information, thereby helping users make more informed decisions. This method not only improves the user experience, but also assumes a certain degree of social responsibility to promote the healthy spread of information.
[0068] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0069] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0070] In the embodiments disclosed in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the device or unit can be in electrical, mechanical or other forms.
[0071] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
Claims
1. A live broadcast system based on AI interaction, characterized in that: The system specifically comprises: The data collection module is used to obtain the real-time bullet screen data and interactive behavior data of users during the live broadcast; An emotion judgment module is used to determine the emotion tendency state judgment result of the user during the live broadcast process through the real-time barrage data and the interactive behavior data; A parameter adjustment module, used to dynamically adjust the AI anchor's emotional expression parameters based on the emotional tendency state judgment result and the live content theme and product recommendation key information; The style transfer module is used to input the emotion expression parameters into the style transfer model, generate dynamic expression image data, and map the dynamic expression image data into the image model of the AI anchor.
2. The system according to claim 1, characterized in that Determining the result of judging the emotional tendency state of the user during the live broadcast process through the real-time bullet screen data and the interactive behavior data specifically includes: Use text analysis algorithms to perform sentiment analysis on the bullet comment data at each time point to obtain the sentiment tendency value at each time point; Calculate the user's comprehensive emotional index at a specific time point based on the type and frequency of the interactive behavior data and the emotional tendency value; The specific time point when the comprehensive emotion index exceeds the preset emotion index threshold is marked as a key emotion node, and each key emotion node is associated with the live broadcast content to determine the key factors that affect the user's emotion changes.
3. The system according to claim 2, characterized in that The text analysis algorithm is used to perform sentiment analysis on the bullet screen data at each time point to obtain the sentiment tendency value at each time point, specifically including: Based on the preset sentiment classification template, sentiment words are extracted from the bullet comment data at each time point; Obtaining user features corresponding to each bullet comment, associating and matching the emotional vocabulary with the corresponding user features to form a preliminary emotional tendency; The preliminary emotional tendency is classified by an emotional recognition algorithm model to obtain the emotional tendency value at each time point.
4. The system according to claim 1, characterized in that Based on the emotional tendency state judgment result, dynamically adjusting the emotional expression parameters of the AI anchor according to the live content theme and product recommendation key information specifically includes: Generate an emotion vector representation based on the emotion tendency state judgment result, and the emotion vector representation satisfies , where v represents the sentiment vector representation, represents several sentiment feature vectors, MLP represents multi-layer perceptron; The live content theme vector is determined by a theme classification model, and the theme classification model satisfies , where t represents the live content theme vector, The live content theme vector represents the content theme category y and the live text content x, where y represents the content theme category and x represents the live text content. express model, W and b represent learnable parameters; The commodity recommendation weight vector is determined by a commodity recommendation weight calculation model, and the commodity recommendation weight calculation model satisfies ,in, represents the recommendation weight of product i, Indicates the click rate, represents the conversion rate, Indicates the relevance score to the live content topic. , and Respectively represent the weights of click-through rate, conversion rate and relevance score to the live content theme; Generate the AI anchor's emotional expression parameters based on the emotional vector representation, the live content theme vector and the product recommendation weight vector. The emotional expression parameters satisfy ,in represents the emotion expression parameter, v represents the emotion vector, t represents the live content theme vector, and w represents the product recommendation weight vector.
5. The system according to claim 1, characterized in that The step of inputting the emotion expression parameters into the style transfer model, generating dynamic expression image data, and mapping the dynamic expression image data to the image model of the AI anchor specifically includes: A generative adversarial network model is used to train a style transfer model, and the emotion expression parameters are converted into dynamic expression image data through the style transfer model. The style transfer model includes a generator, a discriminator and a loss function. The generator is used to convert the emotion expression parameters into dynamic expression image data, and the discriminator is used to detect the dynamic expression image data and output an image authenticity score. The loss function satisfies ,in, represents the total loss, Represents resistance to loss, represents the perceived loss, represents the style loss, and Represent the weights of perceptual loss and style loss respectively; A key point detection algorithm is used to determine the key point coordinates of the dynamic expression image data, and the key point coordinates are mapped to the image model of the AI anchor through a deformation algorithm.
6. The system according to any one of claims 1 to 5, characterized in that The system further comprises: Live content intelligent generation module, including real-time script generation unit, multi-modal material matching unit and crisis handling decision unit; The real-time script generation unit is used to analyze multi-source hot topics according to the heat evaluation model and dynamically generate live broadcast scripts. The heat evaluation model satisfies ,in, Indicates the current time point The comprehensive popularity value of the topic, Indicates the current time point. represents the real-time weight of the jth topic, represents the time attenuation coefficient, represents the latest update timestamp of topic j, and n represents the total number of topics; The multimodal material matching unit is used to perform semantic association matching between the product description and the three-dimensional model and the demonstration animation through a cross-modal similarity calculation model, and the cross-modal similarity calculation model satisfies ,in, Indicates the matching score between the query content Q and the material M. represents the text embedding vector, represents the visual embedding vector, and represents the modal weight coefficient; The crisis handling decision unit is used to detect sensitive content in the live broadcast script according to a sensitivity scoring model, and the sensitivity scoring model satisfies ,in, represents the comprehensive risk score of statement s, Indicates the rule matching score of sentence s based on the sensitive vocabulary. represents the neural network semantic risk prediction value of sentence s, Represents the mixing weight coefficient.
7. The system according to any one of claims 1 to 5, characterized in that The system further comprises: Live interactive decision module, including dynamic reward calculation unit, interactive strategy network unit and meta-learning adaptation unit; The dynamic reward calculation unit is used to construct a multi-objective reward function based on the user's real-time behavior data, calculate the reward signal through a sliding window mechanism based on the multi-objective reward function, and generate a standardized reward value by comparing it with the historical baseline value, wherein the user's real-time behavior data includes the user's stay time, conversion rate and interaction frequency; The interactive strategy network unit is used to generate live broadcast rhythm instructions using an LSTM network, and decompose the live broadcast rhythm instructions into micro-expression parameters, speech synthesis parameters and body movement parameters; The meta-learning adaptation unit is used to adapt the micro-expression parameters, the speech synthesis parameters and the body movement parameters to the image model of the AI anchor using a parameter space mapping matrix.
8. The system according to claim 7, characterized in that The multi-objective reward function satisfies ,in, represents the comprehensive reward value at time point p, represents the weight coefficient of the kth target, ( ) represents the Sigmoid function, Indicates the current target value. represents the historical baseline mean, represents the standard deviation, m represents the total number of targets, and target k includes the user's stay time, conversion rate, and interaction frequency.
9. The system according to any one of claims 1 to 5, characterized in that The system further comprises: Cognitive graph building module, including dynamic user portrait modeling unit, multimodal knowledge graph unit and cognitive bias correction unit; The dynamic user portrait modeling unit is used to integrate the user's historical behavior data, cross-platform social relationship data and consumption preference data in real time, and dynamically generate user feature vectors through a time-series graph neural network; The multimodal knowledge graph unit is used to analyze product attributes, extract live broadcast topics, and mine user interests, and construct a heterogeneous data association graph; The cognitive bias correction unit is used to identify the user's understanding bias based on the user feature vector and the heterogeneous data association graph, and generate a corrected content expression solution.
Citation Information
Cited By
Digital human generation method and system based on AI interaction information
CN120388115A
A digital human generation method and system based on AI interactive information
CN120388115B
Live broadcast bullet screen real-time feedback method and system based on interactive semantic matching
CN120416569A
A method and system for real-time feedback of live broadcast bullet screen based on interactive semantic matching
CN120416569B
Intelligent live broadcast interaction method and system based on AI virtual human
CN120786085A