Multi-style intelligent live broadcast prompting system and method based on AI
Through the AI-based multi-style intelligent live broadcast teleprompt system, multi-source data is collected and analyzed in real time, and the display speed and emotional expression of teleprompt content are dynamically adjusted, the problem of single functions of the existing teleprompt system is solved, personalization and scene adaptation are achieved, and the fluency of live broadcast and audience participation are improved.
Patent Information
- Application Number
- CN202510413534.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing teleprompt system has a single function and cannot fully adapt to the changes in the personalized styles and real-time live broadcast scenes of different anchors, resulting in the teleprompt content that does not match the actual expression rhythm and emotional state of the anchor, affecting the fluency of the live broadcast and the viewing experience of the audience.
Using a multi-style intelligent live broadcast teleprompt system based on AI, we use real-time collection and analysis of multi-source data, combined with voice recognition and sentiment analysis, dynamically adjust the display speed and emotional expression of teleprompt content to generate personalized and scene-adapted teleprompt content.
The personalization and scene adaptation of the teleprompt content is achieved, the attractiveness and audience participation of the live content is improved, the mistakes in the live broadcast are reduced, and the fluency of the live broadcast and the resonance of the audience is improved.
Smart Images

Figure CN120281859A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and specifically to a multi-style intelligent live teleprompter system and method based on AI. Background Art
[0002] In the current live broadcast industry, the role of the teleprompter system is crucial. It can not only help the host express smoothly during the live broadcast, but also effectively improve the quality and attractiveness of the live broadcast content.
[0003] Most of the existing teleprompter systems have single functions and cannot fully adapt to the personalized styles of different hosts and the changes in real-time live broadcast scenarios. They usually can only provide fixed teleprompter content and lack real-time analysis and feedback on the host's emotional state and audience interaction data. This results in the teleprompter content often not matching the actual expression rhythm and emotional state of the host, affecting the fluency of the live broadcast and the viewing experience of the audience. Summary of the Invention
[0004] (I) Technical Problems to be Solved Aiming at the deficiencies of the existing technology, the present invention provides a multi-style intelligent live teleprompter system and method based on AI, which solves the problems of single function, lack of personalized adaptation and insufficient real-time performance of the existing teleprompter system.
[0005] (II) Technical Solutions To achieve the above objectives, the present invention is realized through the following technical solutions: A multi-style intelligent live teleprompter system based on AI, including: Data Acquisition and Analysis Module: Real-time collect live broadcast-related data, and construct an intelligent teleprompter generation network to fuse and analyze multi-source data to generate initial teleprompter content; Teleprompter Adaptation Module: Based on real-time voice analysis technology, dynamically identify the host's live broadcast performance, and combine the recognition results with the initial teleprompter content to adjust the teleprompter content and display speed; Content Generation and Optimization Module: Combine the intelligent teleprompter generation network and the real-time collected data to generate diverse teleprompter content; Optimize the teleprompter content according to the real-time reactions of the host and the audience to make it adapt to the live broadcast style; Output and Feedback Module: Push the optimized teleprompter content to the host in real-time, and at the same time collect real-time feedback data and feedback it to the intelligent teleprompter generation network for continuous optimization.
[0006] In the data collection and analysis module, by integrating user behavior data and market trends, the user interest points are accurately pinpointed. For example, when the market focus is on a certain emerging smart wearable device, the system deeply analyzes the browsing, collection and other behavior data of users on various online platforms for this product series, combined with the capture of trend data such as the popularity of related topics on social media and the growth trend of sales on e-commerce platforms, and quickly identifies users' strong interest in its unique health monitoring function and fashionable appearance design. The historical data of the anchor is analyzed by a deep learning model to extract style features. It records the vocabulary frequency and sentence structure preferences used by the anchor in past live broadcasts. After preprocessing such as data cleaning and dimensionality reduction, the extracted features are transformed into a multi-dimensional feature vector group, mapping out the anchor's style profile. For example, an anchor's language style tends to be vivid and lively, and is good at using catchphrases and metaphorical techniques, and the proportion of such expressions in past live broadcasts far exceeds that of similar anchors. The feature vector group accurately presents this trait. The dynamic capture of audience comments and question data reflects the emotional fluctuations of the live broadcast and activates the system's interactive adjustment mechanism. It classifies the emotions and identifies the intentions of interrogative sentences such as how long can this product be used and exclamatory sentences such as great! I'll place an order right away in real time, and adjusts the teleprompter dynamically based on this. If a large number of users consult the product after-sales policy, relevant detailed answer scripts are quickly pushed, resulting in a more than 30% significant increase in the interaction rate, achieving an accurate match between the information flow and the user feedback flow, enhancing the pertinence of the teleprompter and the harmony of the interaction, enabling the system to break through the complex and changeable live broadcast scenarios and open up a new realm of precise adaptation.
[0007] The system first collects various types of data in real time, including user behavior data such as the product browsing path, stay duration, purchase records of users on the e-commerce platform, market trends such as industry analysis reports, hot news, hot search keywords, and the historical live data of the anchor such as the sales data, audience interaction data, and marketing terms used by the anchor in past live broadcasts, as well as the real-time comments and question data of the audience; taking user behavior data as an example, when a user browses smart wearable devices on the e-commerce platform, the system records in real time the user's browsing order, the stay time on different product pages, and whether the product is added to the favorites or shopping cart, etc. These data are transmitted to the data collection module through the system's API interface and are timestamped; the collected data enters the preprocessing stage. For example, records of users staying on a product page for less than 1 second are removed. These records are generated by user misoperations and have no practical analysis value; for market trend data, if it contains text information, the system performs format conversion through natural language processing technology, such as extracting the key information from a news report and converting it into a semi-structured data format; the multi-branch architecture of the intelligent teleprompter generation network is the key to feature extraction; for time series data, a time-aware branch is used for processing. Taking the product sales time series data in market trends as an example, the temporal convolutional layer uses convolutional kernels of different sizes, such as 3x3 and 5x5, to perform convolutional operations on the sales data to capture the short-term fluctuations and long-term trends of sales over time; a time window mechanism is used, such as dividing the sales data within a quarter into weekly small windows to analyze the changing patterns of weekly sales, so as to reflect the seasonal characteristics of the market in the teleprompter content; for structured text data, self-attention mechanisms are used to extract semantic features. Taking the historical live broadcast script of the anchor as an example, the system tokenizes the script to obtain lexical units, and then converts these lexical units into word vectors through an encoder; the self-attention mechanism calculates the correlation weights between each word vector and the context words during this process, making semantic-rich words, such as key product feature words and marketing words, play a more important role in feature extraction, thus obtaining a vector that can accurately reflect the semantic features of the script; in the face of unstructured text data such as audience comments, the emotion perception branch uses sentiment analysis technology to extract emotion features, and identifies emotion words in the comments by training an emotion dictionary; at the same time, combining the context information and sentence structure of the comments to capture the dynamic changes of emotions. For example, in a long comment, the emotion changes from negative at the beginning to neutral at the end. The finally generated emotion feature vector contains the emotion tendency and intensity of the audience towards the live broadcast content; after the feature vectors are extracted, they are input into the dynamic multi-modal fusion framework DMF; the semantic collection and scattering framework of DMF first performs dimensionality reduction processing on different types of feature vectors and unifies them into the semantic space. For example, the time series feature vector, semantic feature vector, and emotion feature vector are respectively compressed into a low-dimensional semantic space representation through an autoencoder;Next, a hypergraph-based cross-layer and cross-location representation network is constructed to establish connections between various features. For example, time series features and semantic features are associated through shared nodes in the hypergraph, enabling the time trend to influence the content of semantic expressions. Then, through message propagation algorithms, such as the high-level message passing algorithm in non-adjacent communication, multiple interactions are carried out between different types of features to capture potential associations. For instance, when analyzing the live broadcast of smart wearable devices, the information on the increasing sales volume in the market trend and the audience's attention to the health monitoring function will generate high-order associations in the hypergraph, indicating that the teleprompter content should emphasize the advantages of the product in health monitoring. The generation decoder combines the trained language model and the domain knowledge graph to generate the initial teleprompter content adapted to the live broadcast scenario. The language model provides the norms of grammar and semantics, while the domain knowledge graph supplements the professional information and industry terms of the product. For example, when generating the teleprompter content related to smart wearable devices, the language model ensures that the generated sentences are smooth and logical, and the domain knowledge graph provides accurate product function descriptions, such as the heart rate monitoring accuracy of up to ±5 bpm. During the generation process, the decoder will fuse the feature vectors after the high-order message propagation of the previously generated hypergraph, and select the features most relevant to the current live broadcast scenario through the attention mechanism to generate the teleprompter content that conforms to the context. The system dynamically corrects the fusion ratio through adaptive weights, and adjusts the weights of each feature vector during fusion according to the real-time collected feedback from the anchor and the audience, such as whether the anchor is satisfied or dissatisfied with the teleprompter content through voice feedback, and the purchase intention of the audience for the products related to the teleprompter content in the comments.
[0008] In the teleprompter adaptation module, the adaptive algorithm integrates speech recognition, sentiment analysis, and the initial teleprompter content into a dynamic computational flow. It recognizes the real-time speech of the anchor, converts it into text information, and simultaneously extracts feature parameters such as the speech rate and intonation of the speech. Based on the emotional characteristics of the speech and the anchor's historical live broadcast emotion data, it judges the current emotional state of the anchor and generates an emotional state vector E, whose value range is [-1, 1], where -1 represents depression, 1 represents excitement, and 0 represents calmness. Define the semantic feature vector as S. The dynamic computational flow performs semantic matching between the text information obtained from speech recognition and the initial teleprompter content, and calculates the matching degree M, where M = the number of similar words / the total number of words in the initial teleprompter content. The higher the matching degree M, the higher the consistency between the current speech and the initial teleprompter content. Take the emotional state vector E and the matching degree M as the inputs of the adaptive algorithm. The adaptive algorithm dynamically adjusts the display speed V and the emotional expression method W of the teleprompter content according to these two parameters. The adjustment formula for the display speed V is V = V0×(1 + α×E + β×M), where V0 is the initial display speed, and α and β are adaptive weight coefficients, which are dynamically corrected according to the real-time feedback of the anchor and the audience. The initial values of α and β are set to a small positive value, for example, α = 0.2 and β = 0.3, to ensure a moderate response to changes in the emotional state and semantic matching degree while maintaining the stability of the system. Assume that the initial display speed V0 is 3 words per second. During the live broadcast, when the system detects that the emotional state vector E of the anchor is 0.5 (0.5 represents a moderate degree of excitement) and the semantic matching degree M is 0.6, according to the formula V = 3×(1 + 0.2×0.5 + 0.3×0.6) = 3.84 words per second, the display speed of the teleprompter content will be adjusted to 3.84 words per second, which is faster than the initial speed to better match the anchor's excited emotion and the coherence of the semantic topic. If the audience feedback data shows that the audience's acceptance of the teleprompter content has decreased, such as a decrease in the number of likes or an increase in confused comments; the system analyzes these feedbacks and judges that the audience has difficulty keeping up because the display rhythm of the teleprompter content is too fast. At this time, the values of α and β are automatically adjusted downwards, for example, adjusted to α = 0.15 and β = 0.25; and the formula V = 3×(1 + 0.15×0.5 + 0.25×0.6) = 3.675 words per second is reapplied; the teleprompter display speed is appropriately slowed down to adapt to the audience's acceptance. When the anchor is excited, E is positive and M is positive, the display speed V will increase, and vice versa, it will decrease. The adjustment of the emotional expression method W is based on the emotional state vector E. When E is positive, the emotional expression method of the teleprompter content tends to be positive and optimistic, such as using more positive words and emojis; when E is negative, the emotional expression method of the teleprompter content will tend to be negative and cautious, such as using more cautious wording and less emotional expression.
[0009] The content generation and optimization module generates diverse teleprompter content in real time and optimizes it through dynamic semantic flow generation technology, combining initial prompting content, anchor style features, audience interaction data, and real-time market hotspots. The initial prompting content is first deeply analyzed to extract the semantic structure, sentiment tendency, and logical relationship of the text, generating semantic vectors. For example, for the sentence "This product has stable performance and reasonable price", it is recognized that stable performance and reasonable price are two key attributes of the product, and the sentiment tendency and logical relationship of these attributes are extracted. This information is converted into semantic vectors. The anchor style features are combined with the real-time speech analysis results to generate style weights, which are used to adjust the expression mode of the teleprompter content. For example, if the anchor's style is humorous, according to the real-time speech analysis results such as a faster speaking speed and larger intonation fluctuations, corresponding style weights are generated to make the teleprompter content more vivid and interesting when expressed. For example, changing "This product has stable performance" to "This product is like an old friend, stable and reliable, making you feel at ease to use". The audience interaction data generates interaction weights by extracting audience questions, comments, and interaction frequencies, which are used to optimize the interactivity of the teleprompter content. For example, if the audience frequently asks questions about the product's usage method during the live broadcast, the system generates a higher interaction weight, making the teleprompter content pay more attention to the product's usage instructions and operation skills, thereby improving the audience's participation and satisfaction. The above multi-source data are fused, and the weight distribution of each data is adjusted through a dynamic attention mechanism to generate diverse teleprompter content. For example, the dynamic attention mechanism dynamically adjusts the weight distribution of semantic vectors, style weights, interaction weights, and hotspot weights according to real-time feedback data, enabling the teleprompter content to achieve the best effect in different scenarios. For example, when the audience interaction is frequent, the interaction weight increases, making the teleprompter content pay more attention to the audience's questions and comments; when the market hotspots change, the hotspot weight increases, making the teleprompter content more prominent in hot information.
[0010] The output and feedback module captures the emotional change data of the host and the audience in real time through the dynamic emotion feedback unit and optimizes it through an adaptive algorithm. The optimization process is as follows: First, capture the emotional change data of the host in real time, such as the pitch and speaking speed of the host's voice, and judge whether the host is in an excited, calm or depressed mood state by analyzing these data. At the same time, for the emotional change data of the audience, capture it based on whether the comments of the audience are positive, negative or neutral, and the interaction frequency of the audience, etc. When these data are captured, they are immediately transmitted to the intelligent teleprompter generation network. The adaptive algorithm in the intelligent teleprompter generation network plays a role. Taking the host's low mood as an example, if this situation is captured, the adaptive algorithm combines the emotional state vector E and the semantic matching degree M. Assuming that E is negative and the semantic matching degree M is low at this time, the algorithm reduces the display speed V according to the formula V = V0×(1 + α×E + β×M), and at the same time adjusts the emotional expression method W to present the teleprompter content in a more soothing way, helping the host adjust the mood and better adapt to the live broadcast scenario, so as to optimize the teleprompter content, enable the teleprompter system to adapt to the emotional changes of the host and the audience in real time, and improve the live broadcast effect.
[0011] (III) Beneficial effects The present invention provides an AI-based multi-style intelligent live teleprompter system and method, which has the following beneficial effects: 1. Through multi-source data fusion and intelligent generation technology, the present invention realizes the personalization and scene adaptation of the teleprompter content, improves the attractiveness of the live broadcast content and the audience participation, and reduces the live broadcast mistakes caused by improper teleprompters.
[0012] 2. Through the adaptive algorithm, the present invention realizes the real-time and precise optimization of the teleprompter content, improves the fluency of the live broadcast and the resonance of the audience, and reduces the live broadcast mistakes and audience loss caused by the mismatch between the teleprompter content and the host's speaking speed and mood.
[0013] 3. Through the intelligent teleprompter generation network and the dynamic semantic flow, the present invention realizes the diversified generation of the teleprompter content, and improves the interactivity and timeliness of the teleprompter content. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0016] In the data collection and analysis module, the system first collects multi-source data in real time. User behavior data comes from the public data of the e-commerce platform, is transmitted to this module through the API interface, and is timestamped for subsequent analysis. For example, the average stay time of users on a smart wearable device page is 3 minutes, and the browsing path includes viewing product details, reading user reviews, and comparing similar products. These data are obtained through the user behavior tracking system of the e-commerce platform. Market trend data is obtained through web scraping technology and natural language processing technology. For example, hot search keywords, topic popularity, industry reports, etc. related to smart wearable devices are crawled from platforms such as social media and news websites, and their formats are converted. The historical live data of the anchor includes sales data, audience interaction data, marketing terms used by the anchor, etc. in past live broadcasts. These data are obtained through natural language processing and speech analysis of the anchor's past live videos. For example, when introducing electronic products, the frequency of using professional terms by the anchor is about 5 times per minute, and the speaking speed is about 180 words per minute. These data are extracted through speech recognition and text analysis. The real-time comments and questions of the audience are obtained by monitoring the comment area of the live broadcast room in real time. The system will record each comment and question of the audience, including comment content, comment time, commenter ID, etc., and conduct preliminary sentiment analysis and semantic classification on them. For example, the proportion of audience questions about product battery life is 30%, and the proportion of questions about health monitoring functions is 40%, and the rest are other questions.
[0017] For time series data, such as the change of page views over time in user behavior data and the product sales time series in market trend data, the data collection and analysis module uses a time-aware branch for processing. The time-aware branch uses a temporal convolutional network to extract dynamic features. For example, for product sales data, the convolutional kernel size is set to 3 or 5, and the fluctuations of sales data in different time windows are captured through multiple convolutional operations. Assuming that the current sales data has shown an upward trend in the past week, this dynamic feature is captured and used for the generation of prompting content. The time window mechanism divides the data into different time windows. For example, the sales data for a quarter is divided into weekly small windows to analyze the change rules of weekly sales, so as to reflect the seasonal characteristics of the market in the prompting content.
[0018] Structured text data, such as hot search keywords in market trend data and product parameters in industry reports, is processed through a text parsing branch combined with a self-attention mechanism. The number of layers and heads of the self-attention mechanism are adjusted according to the complexity of the data. For example, they are set to 6 layers and 8 heads to fully explore the long-range dependencies between words in the text. For example, for the description of product parameters in an industry report, the heart rate monitoring accuracy is as high as ±5 bpm. Through the self-attention mechanism, the association between heart rate monitoring and accuracy is captured, thus providing more accurate features for the generation of prompting content.
[0019] Unstructured text data, such as audience comments, undergoes sentiment analysis through the sentiment perception branch; the sentiment perception branch utilizes a sentiment dictionary and a deep learning model to identify sentiment words in the comments, and combines the context information and sentence structure of the comments to capture the dynamic changes in sentiment; for example, by training a sentiment dictionary, words are classified into categories such as positive, negative, and neutral, and corresponding weights are assigned to each category; for a comment, the sentiment perception branch identifies the positive sentiment words in it and, through context analysis, determines that its sentiment intensity is high.
[0020] In the data preprocessing stage, the system cleans and converts the format of the collected data; for example, denoising the user behavior data to remove duplicate data and outliers; performing operations such as word segmentation, part-of-speech tagging, and keyword extraction on the text information in the market trend data to convert it into a semi-structured data format; the extracted feature vectors are input into the Dynamic Multi-modal Fusion Framework (DMF), and the semantic collection and scattering framework of the DMF performs dimensionality reduction on different types of feature vectors, such as using an autoencoder to compress the time series feature vector, semantic feature vector, and sentiment feature vector into a low-dimensional semantic space representation respectively; then, a cross-layer and cross-position representation network based on a hypergraph constructs connections between various features, and through message propagation, such as the high-level message passing algorithm in non-adjacent communication, allows different types of features to interact multiple times to capture potential associations.
[0021] In the teleprompter adaptation module, an adaptive algorithm integrates speech recognition, sentiment analysis, and the initial teleprompter content into a dynamic computational flow. Define an emotion state vector E, whose value range is [-1, 1], where -1 represents depression, 1 represents excitement, and 0 represents calm; define the semantic feature vector as S. The dynamic computational flow performs semantic matching between the text information obtained from speech recognition and the initial teleprompter content, calculates the matching degree M, where M = the number of similar words / the total number of words in the initial teleprompter content; takes the emotion state vector E and the matching degree M as the inputs of the adaptive algorithm, and the adaptive algorithm dynamically adjusts the display speed V and the emotional expression method W of the teleprompter content according to these two parameters; the adjustment formula for the display speed V is V = V0×(1 + α×E + β×M), where V0 is the initial display speed, and α and β are adaptive weight coefficients, which are dynamically corrected according to the real-time feedback of the anchor and the audience.
[0022] The content generation and optimization module generates diverse teleprompter content through dynamic semantic flow generation technology, combining the initial teleprompter content, host style features, audience interaction data, and real-time market hotspots. The initial teleprompter content undergoes in-depth analysis to extract the semantic structure, sentiment tendency, and logical relationships of the text, generating semantic vectors. For example, a semantic segmentation model is used to segment the teleprompter content into multiple semantic units, and pre-trained models such as BERT are used to extract the semantic features of each semantic unit to form semantic vectors. The host style features generate style weights by combining the real-time speech analysis results, which are used to adjust the expression of the teleprompter content. For example, if the host's style is inclined to be humorous, according to the real-time speech analysis results, such as a faster speaking speed and larger intonation fluctuations, corresponding style weights are generated to make the teleprompter content more vivid and interesting when expressed.
[0023] The audience interaction data is processed in real time through interaction analysis to extract audience questions, comments, and interaction frequencies, generating interaction weights to optimize the interactivity of the teleprompter content. For example, information such as the number of questions asked by the audience during the live broadcast and the comment frequency is statistically analyzed, and interaction weights are generated based on these data. When the audience interaction is frequent, the interaction weight increases, and the teleprompter content will pay more attention to answering audience questions and guiding interactions. The market hotspot data is parsed in real time to extract hotspot keywords and associated semantics, generating hotspot weights to enhance the timeliness and attractiveness of the teleprompter content. For example, when it is detected that the discussion heat of a certain smart wearable device on social media has increased, relevant hotspot keywords such as health monitoring and long battery life are extracted, and hotspot weights are generated based on the heat of these keywords.
[0024] The weight distribution of each data is adjusted through a dynamic attention mechanism to generate diverse teleprompter content. The dynamic attention mechanism will dynamically adjust the weight distribution of semantic vectors, style weights, interaction weights, and hotspot weights according to real-time feedback data, enabling the teleprompter content to achieve the best effect in different scenarios. For example, when introducing a product, if the discussion heat of health monitoring in the market hotspot data is relatively high, the hotspot weight will increase, and the teleprompter content will emphasize the health monitoring function more.
[0025] Real-time capture the emotional change data of the host and the audience, and optimize it through an adaptive algorithm. The optimization process is as follows: First, capture the emotional change data of the host in real time, such as the pitch and speech rate of the host's voice, and judge whether the host is in an excited, calm or depressed mood state by analyzing these data. At the same time, for the emotional change data of the audience, capture it based on whether the audience's comment content is positive, negative or neutral, and the interaction frequency of the audience, etc. When these data are captured, immediately transmit them to the intelligent teleprompter generation network. The adaptive algorithm in the intelligent teleprompter generation network plays a role. Taking the host's depression as an example, if this situation is captured, the adaptive algorithm combines the emotional state vector E and the semantic matching degree M. Assuming that E is negative and the semantic matching degree M is low at this time, the algorithm uses the formula V = V0×(1 + α×E + β×M) to reduce the display speed V and at the same time adjust the emotional expression method W to present the teleprompter content in a more soothing way, helping the host adjust the mood and better adapt to the live broadcast scenario, so as to realize the optimization of the teleprompter content, make the teleprompter system can adapt to the emotional changes of the host and the audience in real time, and improve the live broadcast effect.
[0026] The system presents the optimized teleprompter content to the host in a new multimodal form, including multimedia elements such as text, images, and sounds. For example, the teleprompter content includes text descriptions, as well as high-definition pictures of products, audio samples such as advertising video clips, etc., to enhance the attractiveness and expressiveness of the teleprompter content. At the same time, the system will display the teleprompter content to the host in a suitable way through the output interface, such as displaying it on the host's prompt screen in the form of scrolling subtitles, or playing it to the host in the form of voice synthesis.
[0027] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI-based multi-style intelligent live teleprompter system, characterized in that, Including: Data collection and analysis module: Collects live broadcast-related data in real time and constructs an intelligent teleprompter generation network to fuse and analyze multi-source data to generate initial teleprompter content; Teleprompter adaptation module: Based on real-time speech analysis technology, dynamically recognizes the live broadcast performance of the anchor and combines the recognition results with the initial teleprompter content to adjust the teleprompter content and display speed; Content generation and optimization module: Combines the intelligent teleprompter generation network and the data collected in real time to generate diverse teleprompter content; Optimizes the teleprompter content according to the real-time reactions of the anchor and the audience to make it adapt to the live broadcast style; Output and feedback module: Pushes the optimized teleprompter content to the anchor in real time, simultaneously collects real-time feedback data, and feeds it back to the intelligent teleprompter generation network for continuous optimization.
2. The multi-style intelligent live teleprompter system based on AI according to claim 1, wherein: The live broadcast-related data includes collected user behavior data, product features, market trends, the anchor's historical live broadcast data, and real-time comments and question data from the audience.
3. An AI-based multi-style intelligent live teleprompter system according to claim 1, characterized in that: The intelligent teleprompter generation network adopts a multi-branch structure to extract feature vectors for different types of data. The extracted feature vectors are input into the dynamic multi-modal fusion framework DMF. DMF uses a semantic collection and scattering framework to transform feature vectors of different structures into the semantic space and constructs a hypergraph for high-order message propagation to capture complex high-order associations between features. DMF is based on a hypergraph-based cross-layer and cross-position representation network for high-order interaction between different layers and positions; The fused feature vectors are further input into the generation decoder, and combined with a language model and a domain knowledge graph to generate initial teleprompter content adapted to the live broadcast scenario; According to the real-time feedback of the anchor and the audience, dynamically corrects the fusion ratio through adaptive weights, enabling the model to adapt to changes in the live broadcast scenario in real time.
4. The multi-style intelligent live teleprompter system based on AI according to claim 1, characterized in that: The teleprompter adaptation module analyzes the emotion of the anchor's speech to judge their current emotional state, transmits the judgment result to the intelligent teleprompter generation network, combines it with the initial teleprompter content, and dynamically adjusts the display speed and emotional expression of the teleprompter content through an adaptive algorithm.
5. The multi-style intelligent live teleprompter system based on AI according to claim 4, characterized in that: The adaptive algorithm integrates speech recognition, emotion analysis, and the initial teleprompter content into a dynamic computational flow. Defines an emotional state vector E, whose value range is [-1, 1], where -1 represents depression, 1 represents extreme excitement, and 0 represents calmness; Defines the semantic feature vector as S. The dynamic computational flow performs semantic matching between the text information obtained from speech recognition and the initial teleprompter content to calculate the matching degree M, M = number of similar words / total number of words in the initial teleprompter content; Uses the emotional state vector E and the matching degree M as the input of the adaptive algorithm. The adaptive algorithm dynamically adjusts the display speed V and emotional expression W of the teleprompter content according to these two parameters; The adjustment formula for the display speed V is V = V0×(1 + α×E + β×M), where V0 is the initial display speed, and α and β are adaptive weight coefficients, which are dynamically corrected according to the real-time feedback of the anchor and the audience.
6. The multi-style intelligent live teleprompter system based on AI according to claim 1, characterized in that: The content generation and optimization module generates diverse teleprompter content in combination with the initial prompt content and live broadcast-related data through dynamic semantic flow generation technology and optimizes it in real time; fuses the above multi-source data, adjusts the weight distribution of each data through a dynamic attention mechanism, and generates diverse teleprompter content.
7. The multi-style intelligent live teleprompter system based on AI according to claim 1, characterized in that: The output and feedback module constructs a dynamic emotional feedback unit, captures the emotional change data of the anchor and the audience in real time, transmits the captured data to the intelligent teleprompter generation network, and optimizes it through an adaptive algorithm.
8. An AI-based multi-style intelligent live teleprompter method, characterized in that, The method is applied to the AI-based multi-style intelligent teleprompter system as described in claim 1, including: Collect live broadcast-related data in real time and construct an intelligent teleprompter generation network to fuse and analyze multi-source data to generate initial teleprompter content; Dynamically identify the live broadcast performance of the anchor, combine the identification result with the initial teleprompter content, and adjust the teleprompter content and display speed; Combine the intelligent teleprompter generation network and the real-time collected data to generate diverse teleprompter content; optimize the teleprompter content according to the real-time reactions of the anchor and the audience to adapt it to the live broadcast style; Push the optimized teleprompter content to the anchor in real time, collect real-time feedback data at the same time, and feedback it to the intelligent teleprompter generation network for continuous optimization.