A method and system for intelligent control of television terminal configuration modes
Patent Information
- Application Number
- CN202611043561.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明的目的在于:针对现有技术中的智能电视配置模式,由于多数电视产品依赖预设固定模板,仅能基于时段、内容类型的单一特征触发简单的配置模式切换,对于不同用户群体的配置模式匹配度或需求适配精确度较低问题,提供一种电视终端配置模式的智能调控方法及配置模式智能调控系统
[0053]1、所述配置模式的智能调控方法,通过Transformer与BERT-CNN融合模型,实现用户特征、配置模式与推荐内容的深度语义对齐,配置模式的匹配准确率不低于92%;通过Apriori关联规则模型,挖掘的高频组合规则作为配置的解释依据,用户可通过电视界面查看配置提示的内容,便于用户查看和确认配置模式切换;对于不同用户群体的配置模式匹配度较高,实现用户需求的精确度更高的匹配度,提升用户使用体验;
Smart Images

Figure CN122845862A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart TV technology, and in particular to a smart control method and system for TV terminal configuration modes. Background Technology
[0002] With the rapid development of the smart TV industry, users' personalized needs for TV terminals are becoming increasingly diversified and dynamic. Currently, smart TV configurations and content recommendations on the market have the following shortcomings:
[0003] First, the configuration modes are simple and relatively fixed. Most products rely on preset fixed templates and can only trigger simple configuration mode switching based on a single feature, such as time period or content type. This cannot accurately adapt to different user groups, such as the complex needs of children, the elderly, and gamers. Users need to frequently manually adjust parameters such as brightness, sound effects, and subtitles, making the operation of smart TV control cumbersome.
[0004] Secondly, the accuracy of content recommendation is low. Existing recommendation systems are mostly based on collaborative filtering or single user profiles, which do not adequately explore the temporal dependence and scene association of users' viewing behavior. This leads to a disconnect between recommended content and users' real-time needs. For example, when a user is watching a suspense film late at night, the system may still recommend comedy content.
[0005] There is an urgent need for a smart TV terminal technology that can accurately match configuration modes and content, dynamically update, protect privacy, and be explainable, in order to meet users' growing personalized needs. Summary of the Invention
[0006] The purpose of this invention is to address the problem that existing smart TV configuration modes rely on preset fixed templates, which can only trigger simple configuration mode switching based on single features such as time period and content type. This results in low accuracy in matching configuration modes or adapting to the needs of different user groups. The invention provides an intelligent control method and system for TV terminal configuration modes.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0008] A method for intelligent control of television terminal configuration modes includes the following steps:
[0009] S1. Collect historical usage data of the TV terminal; the historical usage data includes: user characteristic data, viewing behavior data, environmental scene data, configuration mode description data, and content metadata;
[0010] S2. Through offline training, a multi-feature fusion matching model is obtained:
[0011] Using the Transformer model and the aforementioned movie-watching behavior data, the user's movie-watching behavior is modeled to generate a behavioral feature vector;
[0012] User feature data is converted into static feature vectors through static feature embedding encoding; and the behavioral features are concatenated and fused with the static feature vectors to generate user tags.
[0013] The user tags, description data, and content metadata are semantically aligned using the BERT-CNN fusion model, and a user-configuration-content ternary mapping relationship is generated.
[0014] By using the Apriori association rule model, setting minimum support and minimum confidence, and mining the description data and content metadata from the association rule base, a high-frequency rule base is established.
[0015] S3. Online matching of the configuration mode: Based on the triggering conditions, the multi-feature fusion matching model is invoked to obtain the preferred configuration mode, and the optimal configuration mode is output by filtering through the high-frequency rule base; the triggering conditions include: user input, scene recognition, and behavior monitoring.
[0016] The intelligent control method for TV terminal configuration modes described in this invention achieves deep semantic alignment of user features, configuration modes, and recommended content through a Transformer and BERT-CNN fusion model, with a configuration mode matching accuracy of no less than 92%. Using an Apriori association rule model, high-frequency combination rules are mined as the basis for configuration interpretation. Users can view configuration prompts on the TV interface, facilitating easy viewing and confirmation of configuration mode switching. The method also achieves high matching accuracy for different user groups, resulting in a more precise match to user needs and improved user experience.
[0017] Preferably, the intelligent control method for television terminal configuration modes described in this invention, which generates a three-element mapping relationship between user, configuration, and content, specifically includes the following steps:
[0018] The three branches are respectively inputted with user tag description data, configuration mode description data and meta content data, and then generate user tag vector, configuration mode vector and content metadata vector respectively;
[0019] Text semantic encoding is performed using BERT, and local features are extracted using CNN; the outputs of the three branches are projected into a shared high-dimensional semantic space.
[0020] Calculate the first cosine similarity, second cosine similarity, and third cosine similarity between any two of the user tag vector, configuration mode vector, and content metadata vector;
[0021] The first cosine similarity, the second cosine similarity, and the third cosine similarity are respectively filled into a ternary mapping matrix to generate a user-configuration-content ternary mapping relationship.
[0022] Preferably, in the intelligent control method for TV terminal configuration modes of the present invention, step S3 specifically includes the following steps: converting both structured and unstructured data in the triggering conditions into current feature vectors; calculating the cosine similarity between the current feature vector and the user tag using a cosine similarity algorithm, and selecting at least three preferred configuration modes with high cosine similarity.
[0023] As a preferred embodiment of the present invention, at least three preferred configuration modes with high cosine similarity are selected to improve the reliability of the filtered preferred configuration modes and further improve the accuracy and matching degree of the pre-selected configuration mode selection.
[0024] Preferably, the intelligent control method for the television terminal configuration mode of the present invention further includes S4, dynamic updating and evolution:
[0025] S41. Real-time collection of user interaction data; the user operation data includes: number of configuration adjustments, content operation behavior, and satisfaction rating.
[0026] S42. Based on the user interaction data, set reward calculation rules and calculate the reward value for each user interaction data.
[0027] S43. Based on the reward value, adjust the weights of user, configuration, and content in the ternary mapping relationship using the PPO reinforcement learning model;
[0028] As a preferred embodiment of the present invention, through the above-mentioned dynamic updates and evolution, the feedback-iteration closed loop based on the PPO reinforcement learning model can respond to long-term changes in user behavior and temporary scenario needs, and realize the self-evolution of configuration mode and recommended content; further improve the accuracy of adaptive adjustments to changes in user groups or preferences, and enhance user experience and intelligence.
[0029] Preferably, the intelligent control method for television terminal configuration modes of the present invention, wherein the dynamic update and evolution further includes the step of:
[0030] When the average of the reward values of at least three outputs of the optimal configuration mode is less than 0, the Transformer model and the BERT-CNN model are invoked to regenerate the ternary mapping relationship between the user-configuration-content based on the user interaction data.
[0031] Repeat the mining of descriptive data and content metadata, and update the high-frequency rule base.
[0032] As a preferred embodiment of the present invention, it can respond to long-term changes in user behavior and temporary scenario needs, and realize the self-evolution of configuration modes and recommended content; further improve the accuracy of adaptive adjustments to changes in user groups or preferences, and enhance the user experience and the intelligence of mode configuration or content recommendation.
[0033] Preferably, in the intelligent control method for television terminal configuration modes of the present invention, the historical usage data is stored in the local encrypted storage partition of the television terminal; a federated learning framework is used to process the historical usage data, and feature extraction and encoding are performed to generate an initial feature vector; the initial feature vector is transmitted to the cloud server after being encrypted with AES.
[0034] As a preferred embodiment of the present invention, a multi-level privacy protection mechanism is constructed through the above-mentioned encrypted storage and dynamic update and evolution methods, which runs through the entire process of data collection, processing, transmission and model training, ensuring user data security and complying with relevant regulatory requirements.
[0035] Preferably, the intelligent control method for the TV terminal configuration mode of the present invention further includes the steps of dynamic update and evolution: using a federated learning framework, and using the initial feature vector and user interaction data, executing steps S42 and S43.
[0036] The adjusted multi-feature fusion matching model is encrypted and uploaded to the cloud server. The global model is then optimized using the federated averaging algorithm to generate a global model.
[0037] The cloud server sends the parameters of the global model to the TV terminal and merges the global model with the local multi-feature fusion matching model.
[0038] As a preferred embodiment of the present invention, by retaining the initial feature vector for subsequent matching and model training, the risk of leakage of user privacy data during transmission and cloud storage is reduced. The federated learning framework and edge computing mechanism ensure that the user's original data does not leave the terminal, and only aggregated features are uploaded, effectively protecting user privacy.
[0039] Preferably, in the intelligent control method for the TV terminal configuration mode of the present invention, step S3 further includes the step of: obtaining recommended content through a multi-feature fusion matching model according to the triggering conditions, filtering it through an association rule base, and outputting preferred recommended content.
[0040] As a preferred embodiment of the present invention, based on achieving more precise matching and control of configuration modes, it can output recommended content that is more compatible with user characteristics. Furthermore, through filtering by an association rule base, the output of preferred recommended content is targeted at specific groups, which can filter out unsafe or inappropriate content, thereby improving the matching degree and security of content recommendations and increasing the click-through rate of content recommendations by more than 35%.
[0041] To achieve the objectives of this invention, another technical solution is provided:
[0042] A configuration mode intelligent control system, capable of realizing the intelligent control method for television terminal configuration modes described in this invention, includes: a data acquisition module, a privacy protection module, an offline training module, and an online matching module;
[0043] The data acquisition module is used to collect the historical usage data and the user interaction data;
[0044] The privacy protection module is used to set up a federated learning framework, realize the feature extraction and encoding of the historical usage data and user interaction data, and can encrypt the transmission of the initial feature vector.
[0045] The offline training module is used to train the multi-feature fusion matching model offline.
[0046] The online matching module is used to call the multi-feature fusion matching model to obtain the optimal configuration mode.
[0047] The multi-model fusion configuration mode control system for smart TVs described in this invention achieves deep semantic alignment of user characteristics, configuration modes, and recommended content, with a configuration mode matching accuracy of no less than 92%. Users can view configuration prompts through the TV interface, facilitating easy viewing and confirmation of configuration mode switching. It offers high matching accuracy for different user groups, achieving a higher degree of matching with user needs and improving the user experience. It is compatible with smart TV terminals of different brands and models, supports content integration with IPTV, OTT, and other platforms, and possesses excellent scalability.
[0048] Preferably, the intelligent configuration mode control system of the present invention further includes: a dynamic update module and an interactive display module;
[0049] The dynamic update module is used for the dynamic update and evolution of the multi-feature fusion matching model;
[0050] The interactive display module is used to explain the configuration mode and receive user control commands.
[0051] As a preferred embodiment of the present invention, the dynamic update module can trigger the generation of new modes and new content tags and update the database. In conjunction with the interactive display module, it provides users with an interactive interface for configuration and content, which can display the explanation of the configuration mode, show the basis for the generation of the current configuration and the reasons for the content recommendation. Through the interactive display module, users are provided with a manual adjustment entry, which supports users to manually switch configuration modes and provide feedback on content recommendation satisfaction, further improving the convenience and flexibility of the system in the use of smart TVs.
[0052] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0053] 1. The intelligent adjustment method for configuration modes achieves deep semantic alignment of user features, configuration modes, and recommended content through a Transformer and BERT-CNN fusion model, with a configuration mode matching accuracy of no less than 92%. The high-frequency combination rules mined by the Apriori association rule model serve as the interpretation basis for configuration. Users can view configuration prompts through the TV interface, facilitating easy viewing and confirmation of configuration mode switching. The method also achieves high matching accuracy for different user groups, resulting in a more precise match to user needs and improved user experience.
[0054] 2. The intelligent control system for the configuration mode achieves deep semantic alignment between user characteristics, configuration mode, and recommended content, with a matching accuracy of no less than 92%. Users can view configuration prompts through the TV interface, facilitating their viewing and confirmation of configuration mode switching. It offers high matching accuracy for different user groups, achieving a higher degree of matching with user needs and improving the user experience. It is compatible with smart TV terminals of different brands and models, supports content integration with IPTV, OTT, and other platforms, and possesses good scalability. Attached Figure Description
[0055] Figure 1 This is a flowchart illustrating the intelligent control method for the television terminal configuration mode of the present invention.
[0056] Figure 2 This is a schematic diagram of the module connection of the intelligent control system for configuration modes of the present invention;
[0057] Figure 3 This is a schematic diagram of the specific process of offline training in S2 of the present invention;
[0058] Figure 4 This is a schematic diagram illustrating the specific process of online matching and dynamic updating of the configuration mode as described in this invention;
[0059] Figure 5 This is a schematic diagram illustrating the specific process of the privacy protection mechanism of this invention. Detailed Implementation
[0060] The present invention will now be described in detail with reference to the accompanying drawings.
[0061] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0062] Example 1:
[0063] refer to Figure 1 As shown, this embodiment discloses an intelligent control method for television terminal configuration modes, including the following steps:
[0064] S1. Collect historical usage data of the TV terminal; the historical usage data includes: user characteristic data, viewing behavior data, environmental scene data, configuration mode description data, and content metadata;
[0065] S2. Through offline training, a multi-feature fusion matching model is obtained:
[0066] Using the Transformer model and the aforementioned movie-watching behavior data, the user's movie-watching behavior is modeled to generate a behavioral feature vector;
[0067] User feature data is converted into static feature vectors through static feature embedding encoding; and the behavioral features are concatenated and fused with the static feature vectors to generate user tags.
[0068] The user tags, description data, and content metadata are semantically aligned using the BERT-CNN fusion model, and a user-configuration-content ternary mapping relationship is generated.
[0069] By using the Apriori association rule model, setting minimum support and minimum confidence, and mining the description data and content metadata from the association rule base, a high-frequency rule base is established.
[0070] S3. Online matching of the configuration mode: Based on the triggering conditions, the multi-feature fusion matching model is invoked to obtain the preferred configuration mode, and the optimal configuration mode is output by filtering through the high-frequency rule base; the triggering conditions include: user input, scene recognition, and behavior monitoring.
[0071] The user characteristic data described in this invention includes: age, gender, device model, login identity, and historical configuration preferences; the viewing behavior data includes: playback duration, fast-forward / pause frequency, content type preference, operation records, and content ratings, where operation records are understood to include, for example, subtitle switching and sound effect adjustment; the environmental scene data includes: ambient light intensity, time period, and device connection status, such as whether a game controller is connected; the descriptive data includes: configuration mode templates, such as brightness, sound effects, subtitles, and application list; and the content metadata includes the type of video, director, actors, and duration.
[0072] Specifically, the steps for generating user tags in this invention include:
[0073] Step 1: The user's viewing behavior, such as playback duration, fast forward / pause frequency, and content switching rhythm, is sequentially fed into the Transformer sequence encoder. Through multi-layer self-attention, long-term behavioral dependencies are captured, and a 512-dimensional user behavior feature vector is output, representing the user's dynamic viewing habits.
[0074] Step 2: Static feature embedding encoding method, for example, embedding and mapping discrete structured data such as age, account, time period, device, etc., into static feature vectors with the same dimension as the behavior vector.
[0075] Step 3: Concatenate and fuse the static feature vector and the high-dimensional vector of the Transformer behavior, and send it to the clustering module for unsupervised grouping to generate standardized user labels: for example, children who watch movies at night, heavy game console users, and middle-aged and elderly users who prefer documentaries; bind a set of fixed fused feature vectors to each user label, store them in the database, and use them as the user-end primary key in the mapping relationship.
[0076] It should be noted that the behavioral feature vector described in this invention is understood as a highly condensed and mapped representation of a user's complex movie-watching behavior, interests, preferences, and temporal evolution into a multi-dimensional mathematical space. For example, a user's movie-watching behavior data includes: the content of 50 movies clicked on on TV in the past month, the duration of viewing, and the fast-forwarded positions. After the movie-watching behavior data is transformed into a low-dimensional vector, it is input into a Transformer. Through a self-attention mechanism, the internal connections and temporal order between these 50 movies can be analyzed. After calculation by multiple layers of networks, the movie-watching behavior data is finally compressed into a fixed-length high-dimensional vector through pooling operations, such as a 256-dimensional UserEmbedding.
[0077] The BERT-CNN fusion model described in this invention can be understood as a hybrid neural network model with a classic deep learning architecture in the field of natural language processing. The bidirectional encoder model BERT, through its self-attention mechanism, can deeply understand the context and global semantics of the text. Combined with the convolutional neural network, it scans the text through convolutional kernels to capture local, highly discriminative key phrases. The fusion can generate a more accurate and three-dimensional feature representation.
[0078] The semantic alignment and establishment of a user-configuration-content mapping relationship described in this invention may include: First, capturing global semantics and local details of the behavioral feature vector, the description data, and the content metadata respectively, generating local semantic features and local feature vectors; Second, mapping the heterogeneous features of the three types of data to the same low-dimensional shared latent space through a projection layer, and performing layer normalization; Third, automatically focusing on the relevant parts of the user features or pattern description through a cross-modal or cross-feature attention mechanism to achieve fine-grained semantic alignment and mutual enhancement, concatenating the global semantic features extracted by BERT with the local features extracted by CNN, and then passing through an optional fully connected layer to generate a hybrid feature vector, thereby establishing a ternary mapping relationship between user, configuration, and content.
[0079] In this preferred embodiment, a ternary mapping relationship between user, configuration, and content is generated, specifically including the following steps:
[0080] The three branches are respectively inputted with user tag description data, configuration mode description data and meta content data, and then generate user tag vector, configuration mode vector and content metadata vector respectively;
[0081] Text semantic encoding is performed using BERT, and local features are extracted using CNN; the outputs of the three branches are projected into a shared high-dimensional semantic space.
[0082] Calculate the first cosine similarity, second cosine similarity, and third cosine similarity between any two of the user tag vector, configuration mode vector, and content metadata vector;
[0083] The first cosine similarity, the second cosine similarity, and the third cosine similarity are respectively filled into a ternary mapping matrix to generate a user-configuration-content ternary mapping relationship.
[0084] More specifically, the BERT-CNN fusion model establishes a ternary mapping relationship, including the specific execution steps:
[0085] Step 1: BERT-CNN can unify semantic space alignment. It inputs three streams of text / vector data: Stream 1 includes: user labels + fused user feature vectors (Transformer output vector + static features); Stream 2 includes: text descriptions of various configuration modes (parameter descriptions for eye protection mode, game mode, children's mode, etc.); Stream 3 includes: video content metadata (type, theme, audience, duration).
[0086] Step 2: BERT performs text semantic encoding, CNN extracts local features, and the three outputs are projected onto the same shared high-dimensional semantic space to achieve cross-modal semantic alignment.
[0087] Step 3: Calculate ternary similarity and construct a mapping matrix: Within a unified semantic space, calculate pairwise cosine similarity between the three types of entities: User tag vector ↔ Configuration pattern vector; User tag vector ↔ Content metadata vector; Configuration pattern vector ↔ Content metadata vector;
[0088] Step 4: Fill the similarity values into a ternary mapping matrix with dimensions of [number of user tags × number of configuration templates × number of content libraries]; each cell value in the matrix represents "the matching score of the user tag with the corresponding configuration + content";
[0089] For example, user characteristic data indicates a 6-year-old child, and viewing behavior data shows a historical focus on films like *Paw Patrol* and *Frozen*. The configuration mode description is that when the TV enters children's mode, the system automatically activates eye protection and blue light filtering, and limits single viewing time to no more than 30 minutes. Content metadata is for the animated film *Coco*, with metadata tags including Pixar, family, music, and all-ages. A ternary mapping relationship is then established, strongly binding the vector for children's mode with the content vector for all-ages / animation content. Combined with Transformer encoding of the user's historical behavior sequence, this prioritizes mapping animated feature films that combine high quality with educational / family-oriented themes.
[0090] For example, the offline training process of the present invention includes the following specific steps:
[0091] The first step is data preprocessing: cleaning the collected data, removing missing values and outliers, normalizing the numerical data to the [0,1] interval, and then encoding the unstructured data, such as content types, into One-Hot vectors.
[0092] The second step is user feature extraction: user viewing time sequence behavior data is input into the Transformer sequence encoder, and long-range dependencies between behaviors are captured through a multi-layer self-attention mechanism to generate a user behavior feature vector with a dimension of 512.
[0093] The third step is semantic alignment training: Input the user behavior feature vector and the description data of the configuration mode into the BERT-CNN fusion model. The description data includes, for example, eye protection mode, brightness 30%, and blue light filter enabled. The BERT is responsible for deep semantic analysis, and the CNN captures local feature associations and outputs a user-configuration-content ternary mapping matrix.
[0094] The fourth step is association rule mining: The Apriori algorithm is used to mine the backend configuration and content data, with a minimum support of 0.2 and a minimum confidence of 0.85 to generate a high-frequency rule base, such as "Nighttime (22:00-6:00) ∧ Suspense film → Eye protection mode ∧ Suspense film recommendation";
[0095] Step 5, Model Validation and Optimization: The five-fold cross-validation method is used to evaluate the model performance, with configuration matching accuracy and content recommendation click-through rate as indicators. Model parameters are adjusted, such as the number of Transformer layers and the dimension of BERT hidden layers, to ensure that the configuration matching accuracy is ≥92% and the content recommendation click-through rate is improved by ≥35%.
[0096] Step 6, Database Initialization: Input the user tags, ternary mapping relationships, and association rule databases obtained from training into the database to complete system initialization.
[0097] Step 7, Database Implementation: After training, the matrix, user labels, corresponding feature vectors, and Apriori rule base are uniformly stored in the database module to form a permanent mapping relationship for use in the online matching stage.
[0098] The Apriori association rule model described in this invention can be understood as an association rule mining algorithm. It is a statistical tool that can automatically find hidden connections from massive amounts of data. It measures the association between things through three core indicators: support, confidence, and lift. Support refers to the frequency of a certain feature combination in the total data; confidence refers to the probability that B will occur when A occurs, representing the reliability of the rule; lift refers to how much the occurrence of A promotes the occurrence of B. A lift greater than 1 indicates that A and B are truly strongly associated, rather than B being recommended or configured because it is popular.
[0099] Specifically, the implementation of the Apriori association rule model described in this invention includes, for example, mining high-frequency "user tags - configuration - content" using the Apriori model and sharing rules such as "children's account tags → children's configuration templates + animation content", or "nighttime movie viewing tags → eye protection mode + suspense movie recommendations".
[0100] For example, the Apriori association rule model described in this embodiment mines high-frequency configuration content combinations in the background data to generate high-frequency rules with support ≥ 0.2 and confidence ≥ 0.85. In the online matching configuration mode, the matching results are filtered based on the Apriori rule base. For example, when a child's account is triggered, content containing violent elements and high-brightness configurations are filtered. Such as "Child's account → Children's mode + Animation content recommendation" and "Nighttime viewing → Eye protection mode + Suspense film recommendation".
[0101] In this preferred embodiment, step S3 specifically includes the following steps: converting both structured and unstructured data in the triggering conditions into the current feature vector; calculating the cosine similarity between the current feature vector and the user tag using a cosine similarity algorithm, and selecting at least three preferred configuration modes with high cosine similarity.
[0102] The online matching of the configuration mode described in this embodiment is specifically implemented through the following steps:
[0103] The first step is for the system to monitor trigger events in real time. These trigger events include: user-initiated triggers, such as voice commands (e.g., activating children's mode) or remote control operations (e.g., selecting a family viewing scene); scene-based triggers, such as those triggered by ambient light sensors detecting dimming light or facial recognition detecting a child user; and real-time behavior triggers, such as those triggered by monitoring if a user watches three consecutive suspense films, showing a preference for suspense scenes, or repeatedly fast-forwards through comedy content and rejecting comedy scenes.
[0104] The second step is online matching and execution, which converts the triggering conditions into feature vectors. For example, the feature vector corresponding to the children's scenario is [Age: 5, Identity: Child, Time of Day: Weekend].
[0105] The third step is similarity matching, which calculates the cosine similarity between the current feature vector and the user tags in the database, and selects the 3 configuration patterns and 10 content items with the highest matching degree; rule filtering: filters results that do not conform to the scenario based on the Apriori rule library, for example, filtering content containing violent elements and high brightness configuration in the children's scenario;
[0106] The fourth step is execution and display: the optimal configuration mode is applied to the TV terminal, the recommended content is displayed on the interface, and the explanation of the configuration and recommendation is displayed simultaneously.
[0107] In this preferred embodiment, it further includes S4, dynamic updating and evolution: S41, real-time collection of user interaction data; the user operation data includes: configuration adjustment times, content operation behavior, and satisfaction score; S42, based on the user interaction data, setting reward calculation rules and calculating the reward value for each user interaction data; S43, based on the reward value, adjusting the weights of user, configuration, and content in the ternary mapping relationship through the PPO reinforcement learning model.
[0108] The PPO model described in this invention is understood as a reinforcement learning model for proximal policy optimization. Through a truncation mechanism, it limits the update magnitude, allowing the model to try new policies while restricting the ratio of change between old and new policies. For example, the triple mapping of the PPO model includes the mapping relationship between the agent, environment, and reward. The agent represents the ranking strategy or reordering module of the recommendation system; the environment reflects the user's real feedback, including clicks, playback duration, bounces, and ratings; the reward reflects the user's positive feedback, such as high completion rates for high scores and bounces for negative scores. When the configuration mode control system wants to try a new content ranking logic, for example, interspersing high-quality behind-the-scenes documentaries in home theater mode, PPO can be safely tested in online or offline environments. If user feedback is positive, the model will steadily increase the weight of this configuration mode's ranking; if users dislike it, PPO's truncation mechanism will prevent the model from over-adjusting.
[0109] Specifically, in this embodiment, dynamic updating and evolution can be achieved through the following steps:
[0110] The first step is feedback collection: real-time collection of user operation data, including the number of configuration adjustments, content click / skip behavior, and satisfaction rating;
[0111] The second step is reward calculation: Calculate the reward value for each interaction according to the preset reward mechanism. For example, a user gets +0.5 points for clicking on recommended content and -0.8 points for manually adjusting the configuration.
[0112] The third step is model iteration: the PPO reinforcement learning model updates the policy network and adjusts the weights of the ternary mapping relationship based on the reward signal.
[0113] The fourth step is new pattern generation: When the average reward of three consecutive matching results is less than 0, the system calls the Transformer and BERT-CNN models to generate new configuration patterns and content tags based on the latest user behavior data and update them to the database; Rule update: The system performs offline mining on the association rule base once a month, updates high-frequency rules, and injects them into the main model.
[0114] In a preferred embodiment, the dynamic update and evolution further includes the following steps: when the average of the reward values of at least three outputs of the optimal configuration mode is less than 0, the Transformer model and the BERT-CNN model are invoked to regenerate the ternary mapping relationship of generating user-configuration-content based on the user interaction data; the mining of descriptive data and content metadata is performed again to update the high-frequency rule base.
[0115] For example, based on the PPO reinforcement learning model, the dynamic iteration of configuration patterns and recommended content can be implemented using the same steps:
[0116] The first step is feedback collection: collect user operation feedback in real time, including configuration adjustment behavior, content click / skip behavior, and satisfaction ratings from 1 to 5.
[0117] The second step is to establish a reward mechanism: Positive rewards are set up such that if a user uses the configuration mode continuously for 7 days without adjustment, the reward is +1; clicking on recommended content results in a reward of +0.5. Negative rewards are set up such that if a user manually adjusts the configuration, the reward is -0.8; skipping recommended content results in a reward of -0.3.
[0118] The third step is model update: adjust the PPO model weights based on the reward signal and optimize the user-configuration-content mapping relationship; when the average reward of the matching result is lower than the threshold, such as when it is lower than 0 for 3 consecutive times, trigger the generation of new modes and new content tags and update the database.
[0119] In this preferred embodiment, the historical usage data is stored in a local encrypted storage partition of the television terminal; a federated learning framework is used to process the historical usage data, and features are extracted and encoded to generate an initial feature vector; the initial feature vector is transmitted to the cloud server after being encrypted with AES.
[0120] This invention constructs a multi-layered privacy protection mechanism that runs through the entire process of data collection, processing, transmission and model training, ensuring user data security and complying with relevant regulations.
[0121] In a preferred embodiment, the dynamic update and evolution further includes the following steps: using a federated learning framework, executing steps S42 and S43 with the initial feature vector and user interaction data; encrypting the adjusted multi-feature fusion matching model and uploading it to the cloud server; optimizing the global model using a federated averaging algorithm to generate a global model; and having the cloud server distribute the parameters of the global model to the television terminal, fusing the global model with the local multi-feature fusion matching model.
[0122] For example, to achieve global optimization and self-evolution of the model while avoiding the collection of raw data, this invention employs a federated learning architecture for model training:
[0123] Local training on the terminal: The TV terminal acts as a node in the federated learning process, using locally generated feature vectors and user feedback to perform local iterative updates on the PPO reinforcement learning model in this invention.
[0124] Gradient Upload and Aggregation: The terminal only encrypts the updated gradients or changes in model weights (not user data or feature vectors) before uploading them to the cloud server.
[0125] Global model optimization: The cloud server aggregates encrypted gradient information from a large number of decentralized terminals and optimizes the global model through a federated averaging algorithm to generate a higher-performance "global model".
[0126] Model distribution and synchronization: The cloud distributes optimized global model parameters to each terminal, which then merges them with its local model to improve model performance. The entire process ensures that no user's personalized data is exposed to the cloud or other stakeholders.
[0127] In this embodiment, specifically, federated training can be achieved through a federated learning framework. Only the aggregated statistical information of the initial feature vectors and the model gradient are uploaded to the cloud server. After the cloud server performs global model optimization, the updated model parameters are sent to the smart TV terminal. Furthermore, data encryption is performed, and the transmission of the initial feature vectors and model parameters uses end-to-end AES encryption to prevent data leakage.
[0128] For example, all the raw user data described in this invention is collected locally on the TV terminal and immediately processed by a feature encoder deployed on the terminal side. This encoder converts the raw data into high-dimensional, irreversible feature vectors. The raw data is strictly stored in an encrypted storage partition on the terminal and is strictly prohibited from being directly uploaded to the cloud in any form. The system only retains the feature vectors for subsequent matching and model training, fundamentally eliminating the risk of leakage of user privacy data during transmission and cloud storage.
[0129] More specifically, all data to be transmitted in this invention (including but not limited to encryption gradients, aggregated feature statistics, system configuration updates, etc.) is end-to-end encrypted using AES-256 or equivalent high-strength encryption algorithms. Simultaneously, the transmitted data is anonymized using temporarily generated communication credentials not bound to user identity, ensuring that data transmission cannot be identified, intercepted, or traced during the transmission process.
[0130] In a preferred embodiment, step S3 further includes the step of: obtaining recommended content through a multi-feature fusion matching model based on the triggering condition, filtering it through an association rule base, and outputting preferred recommended content.
[0131] For example, privacy is protected through a differential privacy reward feedback mechanism. Differential privacy technology is introduced into the reward signal collection stage of the PPO reinforcement learning model. By adding controllable random noise to locally collected user action feedback, such as clicks, adjustments, and ratings, it becomes impossible to accurately infer a user's specific behavioral habits from a single user's feedback. This mechanism provides an additional layer of privacy protection for the federated learning framework, further enhancing the security of the model training process.
[0132] For example, privacy protection and data security are achieved through user-controllable privacy authorization and data management, specifically through the following mechanisms: The system provides a user interface that allows users to perform fine-grained privacy authorization management, including: data collection switch, local data cleaning, and anonymity mode; the data collection switch allows users to disable the collection of non-core viewing data at any time, retaining only basic configuration mode matching. Local data cleaning allows users to clear all raw data and feature vectors stored locally on the terminal with one click. Anonymity mode provides "guest mode" or "anonymous viewing mode." In this mode, the system only relies on the behavior of the current session for real-time matching, does not record any long-term user characteristics, and the data is automatically cleared after the viewing ends.
[0133] Through the above mechanisms, this invention constructs a closed-loop privacy protection system from the data source to model training and then to user management, which maximizes the security and privacy of user data while realizing the intelligent functions of the system.
[0134] Example 2:
[0135] This embodiment discloses an intelligent control system for configuration modes, which can realize the intelligent control method for TV terminal configuration modes as described in Embodiment 1, including: a data acquisition module, a privacy protection module, an offline training module, and an online matching module;
[0136] The data acquisition module is used to collect the historical usage data and the user interaction data;
[0137] The privacy protection module is used to set up a federated learning framework, realize the feature extraction and encoding of the historical usage data and user interaction data, and can encrypt the transmission of the initial feature vector.
[0138] The offline training module is used to train the multi-feature fusion matching model offline.
[0139] The online matching module is used to call the multi-feature fusion matching model to obtain the optimal configuration mode.
[0140] It should be noted that the control system described in this invention also includes conventional execution modules such as data storage modules and calling interfaces. Specifically, a database module is set up to store user tags, configuration mode templates, content metadata, ternary mapping relationships, association rule bases, and reinforcement learning model weights. A distributed storage architecture is adopted to support fast querying and updating.
[0141] Specifically, for example, the privacy protection module adopts a federated learning framework to achieve local processing of user data; it uses local feature encoding, where the user's original data completes feature extraction and encoding locally on the terminal, thereby generating a feature vector.
[0142] In a preferred embodiment, the system further includes a dynamic update module and an interactive display module. The dynamic update module is used for the dynamic updating and evolution of the multi-feature fusion matching model. The interactive display module is used to display the configuration mode explanation and receive user control commands. Specifically, it provides users with an interactive interface for configuration and content, including explaining the configuration mode and displaying the basis for generating the current configuration, such as automatically turning on eye protection mode based on your nighttime viewing habits; displaying the reasons for content recommendations and showing the matching points between the recommended content and user needs, such as recommending this film because you recently prefer the suspense genre; and also enabling a manual adjustment entry, supporting users to manually switch configuration modes and provide feedback on content recommendation satisfaction.
[0143] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for intelligent control of a television terminal configuration mode, characterized in that, Including the following steps: S1. Collect historical usage data of the TV terminal; the historical usage data includes: user characteristic data, viewing behavior data, environmental scene data, configuration mode description data, and content metadata; S2. Through offline training, a multi-feature fusion matching model is obtained: Using the Transformer model and the aforementioned movie-watching behavior data, the user's movie-watching behavior is modeled to generate a behavioral feature vector; User feature data is converted into static feature vectors through static feature embedding encoding; and the behavioral features are concatenated and fused with the static feature vectors to generate user tags. The user tags, description data, and content metadata are semantically aligned using the BERT-CNN fusion model, and a user-configuration-content ternary mapping relationship is generated. By using the Apriori association rule model, setting minimum support and minimum confidence, and mining the description data and content metadata from the association rule base, a high-frequency rule base is established. S3. Online matching of the configuration mode: Based on the triggering conditions, the multi-feature fusion matching model is invoked to obtain the preferred configuration mode, and the optimal configuration mode is output by filtering through the high-frequency rule base; the triggering conditions include: user input, scene recognition, and behavior monitoring.
2. The intelligent control method for television terminal configuration modes according to claim 1, characterized in that, It also generates a ternary mapping relationship between user, configuration, and content, specifically including the following steps: The three branches are respectively inputted with user tag description data, configuration mode description data and meta content data, and then generate user tag vector, configuration mode vector and content metadata vector respectively; Text semantic encoding is performed using BERT, and local features are extracted using CNN; the outputs of the three branches are projected into a shared high-dimensional semantic space. Calculate the first cosine similarity, second cosine similarity, and third cosine similarity between any two of the user tag vector, configuration mode vector, and content metadata vector; The first cosine similarity, the second cosine similarity, and the third cosine similarity are respectively filled into a ternary mapping matrix to generate a user-configuration-content ternary mapping relationship.
3. The intelligent control method for television terminal configuration modes according to claim 1, characterized in that, S3 specifically includes the following steps: converting both structured and unstructured data in the triggering conditions into the current feature vector; calculating the cosine similarity between the current feature vector and the user tag using a cosine similarity algorithm, and selecting at least three preferred configuration modes with high cosine similarity.
4. The intelligent control method for television terminal configuration modes according to claim 1, characterized in that, It also includes S4, dynamic updates and evolution: S41. Real-time collection of user interaction data; the user operation data includes: number of configuration adjustments, content operation behavior, and satisfaction rating. S42. Based on the user interaction data, set reward calculation rules and calculate the reward value for each user interaction data. S43. Based on the reward value, adjust the weights of user, configuration, and content in the ternary mapping relationship using the PPO reinforcement learning model.
5. The intelligent control method for television terminal configuration modes according to claim 4, characterized in that, The dynamic update and evolution also includes the following steps: When the average of the reward values of at least three outputs of the optimal configuration mode is less than 0, the Transformer model and the BERT-CNN model are invoked to regenerate the ternary mapping relationship between the user-configuration-content based on the user interaction data. Repeat the mining of descriptive data and content metadata, and update the high-frequency rule base.
6. The intelligent control method for television terminal configuration modes according to claim 4, characterized in that, The historical usage data is stored in the local encrypted storage partition of the TV terminal; the historical usage data is processed using a federated learning framework, and features are extracted and encoded to generate an initial feature vector; the initial feature vector is transmitted to the cloud server after being encrypted with AES.
7. The intelligent control method for television terminal configuration modes according to claim 6, characterized in that, The dynamic update and evolution also includes the following steps: Using a federated learning framework, S42 and S43 are executed with the initial feature vector and user interaction data. The adjusted multi-feature fusion matching model is encrypted and uploaded to the cloud server. The global model is then optimized using the federated averaging algorithm to generate a global model. The cloud server sends the parameters of the global model to the TV terminal and merges the global model with the local multi-feature fusion matching model.
8. The intelligent control method for television terminal configuration modes according to any one of claims 1-7, characterized in that, S3 further includes the step of: obtaining recommended content through a multi-feature fusion matching model based on the triggering condition, filtering it through an association rule base, and outputting the preferred recommended content.
9. A configuration mode intelligent control system, characterized in that, The intelligent control method for realizing the TV terminal configuration mode as described in any one of claims 1-8 includes: a data acquisition module, a privacy protection module, an offline training module, and an online matching module; The data acquisition module is used to collect the historical usage data and the user interaction data; The privacy protection module is used to set up a federated learning framework, realize the feature extraction and encoding of the historical usage data and user interaction data, and can encrypt the transmission of the initial feature vector. The offline training module is used to train the multi-feature fusion matching model offline. The online matching module is used to call the multi-feature fusion matching model to obtain the optimal configuration mode.
10. The intelligent control system for configuration modes according to claim 9, characterized in that, Also includes: Dynamic update module and interactive display module; The dynamic update module is used for the dynamic update and evolution of the multi-feature fusion matching model; The interactive display module is used to explain the configuration mode and receive user control commands.