IPTV interface element differentiation display method and system
By establishing a temporal correlation model between voice intent and interface operation and an intent-element probability matrix, combined with dynamic interface reconstruction and attention detection, the problems of inaccurate voice intent perception and insufficient intelligent interface adaptation in IPTV interface element display are solved, realizing personalized and real-time interface adjustment and optimization.
Patent Information
- Application Number
- CN202511463863.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing IPTV interface element display technologies suffer from insufficient voice intent perception, weak context perception, and simplified interface adjustment mechanisms in multimodal voice interaction. This leads to a discrepancy between the interface response and user expectations, making it unable to adapt to real-time changes in users' personalized needs.
By collecting the audio features and semantic components of user voice commands, a temporal correlation model between voice intent categories and subsequent interface operation sequences is established. An intent-element probability matrix is constructed, and a multi-level progressive adjustment mechanism is adopted to control element visibility and interaction response priority. The interface layout is dynamically reconstructed and the user's visual scanning path is predicted. Personalized display is achieved by combining an attention detection mechanism.
It achieves precise interface adjustments, improves the ability to perceive voice intent, adapts to real-time changes in user status, optimizes interface layout and element visibility, shortens user visual search time, provides adaptive visual guidance, and enhances user experience.
Smart Images

Figure CN120935418A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method and system for differentiating IPTV interface elements. Background Technology
[0002] With the rapid development of intelligent voice technology, IPTV interface element display is gradually evolving towards voice interaction. Existing methods for differentiating IPTV interface elements mainly rely on preset user classification templates and fixed interface configuration rules. They perform simple interface personalization adjustments by analyzing users' basic attributes, viewing history, and device information. Traditional methods employ rule-based interface element control mechanisms, matching corresponding interface templates based on user profiles. Differentiation is achieved by adjusting basic attributes such as font size, color theme, and layout style. Simultaneously, basic voice recognition functionality supports basic voice navigation operations. These technologies have a certain application foundation in static personalized display.
[0003] However, existing technologies have significant limitations in adapting interfaces for multimodal voice interaction. First, they lack the ability to perceive voice intent. Traditional speech recognition mainly focuses on the literal meaning of instructions and lacks the ability to predict the user's deeper intent and subsequent operational needs, resulting in a deviation between the interface response and the user's expectations. Second, they have weak context awareness. Existing methods struggle to comprehensively consider multidimensional factors such as educational scenario characteristics, user cognitive load, and attention state, making it impossible to accurately adjust the interface according to the user's current specific context. Third, the interface element adjustment mechanism is too simplistic and lacks the ability to dynamically optimize based on the user's real-time state. The adjustment of interface layout and element visibility often adopts a fixed pattern, making it difficult to adapt to the real-time changes in the user's personalized needs. Summary of the Invention
[0004] This application provides a method and system for differentiated display of IPTV interface elements, which solves the problems of inaccurate voice intent perception and insufficient intelligent interface adaptation in differentiated display of IPTV interface elements.
[0005] Firstly, this application provides a method for differentiated display of IPTV interface elements, the method comprising: Step S1: Collect the audio features and semantic components of the user's voice commands, establish a temporal correlation model between the voice intent category and the subsequent interface operation sequence, and train and generate a voice interaction prediction dataset based on historical interaction data. Step S2: Weight the voice interaction prediction dataset according to the educational scenario and the user's cognitive load to construct an intent-element probability matrix; Step S3: Calculate the real-time importance score of the interface elements based on the intent-element probability matrix, and use a multi-level progressive adjustment mechanism to synchronously control the element visibility, position weight and interaction response priority; Step S4: Based on the interface space utilization and visual balance constraints, dynamically reconstruct the element layout and predict the user's visual scanning path to generate an adaptive interface space allocation scheme. Step S5: Detect the user's distracted state. When the user's concentration is below the preset threshold, activate the enhanced guidance mode. Otherwise, adopt a lightweight prompt strategy and output personalized and differentiated display results.
[0006] Secondly, this application provides a differentiated display system for IPTV interface elements, the IPTV interface element differentiated display system comprising: The acquisition module is used to acquire the audio features and semantic components of user voice commands, establish a temporal correlation model between voice intent category and subsequent interface operation sequence, and train and generate a voice interaction prediction dataset based on historical interaction data. The encoding module is used to weight the voice interaction prediction dataset according to the educational scenario and the user's cognitive load, and construct an intent-element probability matrix; The synchronization module is used to calculate the real-time importance score of interface elements based on the intent-element probability matrix, and adopts a multi-level progressive adjustment mechanism to synchronously control the visibility of elements, position weight and interaction response priority. The generation module is used to dynamically reconstruct the element layout and predict the user's visual scanning path based on the interface space utilization and visual balance constraints, and generate an adaptive interface space allocation scheme. The output module is used to detect the user's distracted state. When the user's concentration is below a preset threshold, an enhanced guidance mode is activated. Otherwise, a lightweight prompt strategy is adopted to output personalized and differentiated display results.
[0007] Thirdly, an IPTV interface element differentiation display device is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the IPTV interface element differentiation display device to execute the above-described IPTV interface element differentiation display method.
[0008] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the above-described method for differentiating IPTV interface elements.
[0009] The technical solution provided in this application overcomes the technical bottleneck of insufficient voice intent perception capability in traditional IPTV interface element display technology by innovatively collecting audio features and semantic components of user voice commands and establishing a temporal correlation model. This method can simultaneously extract multi-dimensional features such as spectrum, pitch, and speech rate from audio signals, and accurately identify the user's true intent by combining natural language processing algorithms. More importantly, it establishes a temporal correlation between voice intent and subsequent interface operation sequences through Markov chain modeling technology, enabling interface elements to predictively respond to user needs rather than passively waiting for operation commands. The voice interaction prediction dataset generated based on historical interaction data provides a reliable data foundation for subsequent intelligent decision-making, solving the core problem of the semantic gap between voice input and interface response in traditional methods. The technical solution of constructing an intent-element probability matrix by weighting the voice interaction prediction dataset according to educational scenarios and user cognitive load effectively solves the limitation of existing technologies that cannot comprehensively consider multi-dimensional contextual factors. This matrix provides a scientific basis for accurate interface adjustment decisions by quantifying the influence of different voice intents on each interface element. In particular, the introduction of educational scenario weights and cognitive load weights enables interface adjustments to simultaneously take into account the application characteristics of the IPTV education platform and the user's real-time cognitive state. This innovative technology, based on the intent-element probability matrix to calculate the real-time importance score of interface elements and adopting a multi-level progressive adjustment mechanism, completely changes the coarse mode of traditional interface element adjustment. By establishing five refined visibility levels—fully visible, highlighted, normal, faded, and hidden—and combining them with progressive transparency and size adjustment algorithms, it ensures the smoothness of interface changes and the continuity of user experience. At the same time, the design of the three-dimensional adjustment control vector realizes the coordinated optimization of visibility, position weight, and interaction response priority.
[0010] The algorithm design, which dynamically reconstructs and predicts the user's visual scanning path based on interface space utilization and visual balance constraints, demonstrates significant technical advantages in the field of differentiated display of IPTV interface elements. By establishing a mathematical constraint model for the uniformity of element density distribution and the stability of the visual center of gravity, the algorithm ensures that the reconstructed interface meets both functional requirements and visual aesthetic principles. In particular, the application of the user visual scanning path prediction model, based on the research results of cognitive psychology and eye tracking, can accurately predict the user's visual movement trajectory and time cost from the current point of attention to the target element, providing scientific theoretical guidance for interface layout optimization and significantly shortening the user's visual search time and operation path length. The intelligent mechanism that detects user distraction and selects a guidance mode based on attention concentration threshold represents a significant advancement in IPTV interface technology towards cognitive perception. This mechanism accurately judges the user's cognitive load by monitoring objective indicators such as user operation pause time and error frequency in real time. When attention concentration is below a preset threshold, it automatically activates an enhanced guidance mode including highlighting and arrow indication. When attention is good, it adopts a lightweight prompt strategy with color changes and transparency adjustment. This adaptive guidance mechanism not only avoids user interference caused by over-guidance but also ensures that effective visual guidance can be provided in a timely manner when the user needs help. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of an embodiment of the method for differentiated display of IPTV interface elements in this application. Figure 2 This is a schematic diagram of an embodiment of the IPTV interface element differentiation display system in this application. Figure 3 This is a schematic block diagram of the structure of the IPTV interface element differentiation display device in an embodiment of the present invention. Detailed Implementation
[0013] This application provides a method and system for differentiated display of IPTV interface elements. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0014] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the method for differentiated display of IPTV interface elements in this application includes: Step S1: Collect the audio features and semantic components of the user's voice commands, establish a temporal correlation model between the voice intent category and the subsequent interface operation sequence, and train and generate a voice interaction prediction dataset based on historical interaction data. Step S2: Weight the voice interaction prediction dataset according to the educational scenario and the user's cognitive load to construct an intent-element probability matrix; Step S3: Calculate the real-time importance score of the interface elements based on the intent-element probability matrix, and use a multi-level progressive adjustment mechanism to synchronously control the element visibility, position weight and interaction response priority; Step S4: Based on the interface space utilization and visual balance constraints, dynamically reconstruct the element layout and predict the user's visual scanning path to generate an adaptive interface space allocation scheme. Step S5: Detect the user's distracted state. When the user's concentration is below the preset threshold, activate the enhanced guidance mode. Otherwise, adopt a lightweight prompt strategy and output personalized and differentiated display results.
[0015] It is understood that the executing entity of this application can be an IPTV interface element differentiation display system, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiment uses a server as an example for illustration.
[0016] Specifically, the voice command data acquisition process captures the user's audio input through the iFlytek voice assistant interface, extracting audio features such as spectrum, pitch, and speech rate. Then, natural language processing algorithms are used to perform semantic analysis on the audio features, identifying semantic components such as keywords, action words, and target objects. These semantic components are categorized into navigation, operation, query, and setting categories, forming voice intent category labels with timestamps and confidence levels. The construction of the temporal correlation model is based on Markov chain theory, using voice intent as state nodes and subsequent user interface operation sequences as state transitions. By statistically analyzing the frequency of transitions from different intent categories to various operation sequences in historical data, a transition probability matrix is calculated. This matrix, after normalization and Laplace smoothing optimization, forms the voice intent temporal correlation model. Historical user interaction records are input into this model for parameter training and weight adjustment, ultimately generating a voice interaction prediction dataset with predictive capabilities.
[0017] The weighted encoding process categorizes and labels the voice interaction prediction dataset according to the specific application scenarios of the IPTV education platform. Different educational scenarios, such as tutorials, function browsing, content search, and module switching, are assigned corresponding label weights. Simultaneously, a cognitive load index is calculated by monitoring user pause times and error rates. Cognitive load is quantified into three levels: low, medium, and high, each corresponding to a different weight coefficient. The generation of the dual-weighted encoding vector combines the educational scenario weights and cognitive load weights to form a multi-dimensional vector representation. This vector is used to calculate the probability of each voice intent's influence on different interface elements, ultimately constructing a two-dimensional probability matrix structure with behavioral intents as rows and interface elements as columns. Real-time importance scoring extracts corresponding probability values from the intent-element probability matrix and performs real-time matching with the current user's specific voice intent to derive the importance value of each interface element. This value is used as the basis for visibility level classification. Interface elements are divided into five levels: fully visible, highlighted, normal, faded, and hidden. A progressive adjustment algorithm controls the smooth transition of elements between different visibility levels. By setting the transparency change step size and size scaling ratio, abrupt changes in interface elements are avoided to prevent them from affecting the user experience.
[0018] The dynamic reconstruction process is based on interface space utilization data and visual balance constraints. Space utilization is calculated by the ratio of occupied area to total display area. Visual balance constraints include two dimensions: uniformity of element density distribution and stability of visual center of gravity. The dynamic reconstruction algorithm allocates visual focus areas to high-importance elements and arranges edge positions for low-importance elements according to these constraints. During the reconstruction process, a user visual scanning path prediction model is also established. This model calculates the shortest visual distance and expected scanning time from the current point of focus to the target element, ensuring that the interface layout conforms to the user's visual habits. Attention concentration is judged by monitoring the pause time and error frequency in user operations. When the concentration index is below a preset threshold, an enhanced guidance mode is triggered, using prominent visual guidance methods such as highlighting and flashing, and arrow indicators. When the concentration is normal, a lightweight prompt mode is used, providing gentle prompts through color changes and transparency adjustments. The selection of personalized guidance strategies is dynamically adjusted based on the user's current cognitive state and attention level.
[0019] In one specific embodiment, step S1 includes: The audio features of user voice commands are extracted through the iFlytek voice assistant interface to obtain audio feature data including spectrum, pitch, and speech rate. Semantic component analysis of the audio feature data is performed based on natural language processing algorithms to obtain the semantic component structure of keywords, action words, and target objects. The semantic component structure is classified into navigation, operation, query, and setting categories to obtain voice intent category labels with timestamps and confidence levels; Collect the user interface operation sequence after the occurrence of the voice intent category label, record the interaction behaviors of clicking, swiping, and dwell time, and obtain the intent-operation behavior correspondence data; Based on the intent-operation behavior correspondence data, a temporal association model is constructed, and the transition probability matrix between different intent categories and subsequent operation sequences is calculated to obtain the speech intent temporal association model. Historical user interaction records are input into the speech intent temporal association model for training and optimization. The transition probability parameters and association weight coefficients are updated to obtain the speech interaction prediction dataset.
[0020] Specifically, the audio feature extraction process of the iFlytek voice assistant interface digitizes the user's voice input at a sampling frequency of 16kHz. Then, it converts the time-domain audio signal into a frequency-domain representation using a Fast Fourier Transform (FFT) algorithm, extracting spectral distribution features within the range of 0Hz to 8kHz. Simultaneously, it uses an autocorrelation function to calculate the fundamental frequency variation trajectory of the audio to form pitch features, and uses frame length analysis and energy detection algorithms to calculate the rhythmic changes of the speech to obtain speech rate features. These three types of feature data are organized and stored in a multi-dimensional vector format. The natural language processing algorithm analyzes the semantic components of the audio feature data using a combination of part-of-speech tagging and syntactic analysis. First, the text converted from speech recognition is segmented into words. Then, a Hidden Markov Model (HMM) is used for part-of-speech tagging to identify nouns, verbs, adjectives, and other part-of-speech categories. Next, dependency parsing is used to determine the grammatical relationships between words, extracting keywords as the core concepts of the speech content, action words as the user's operational intentions, and target objects as the specific targets of the operations, forming a semantic component representation with a triple structure.
[0021] The voice intent classification process is based on a predefined intent classification system. Navigation intents include navigation-related voice commands such as page navigation and module switching; operation intents cover specific operation commands such as clicking, selecting, and confirming; query intents include query requests such as searching, finding, and retrieving information; and setting intents involve setting operations such as configuration modification and parameter adjustment. The classification algorithm uses the Support Vector Machine (SVM) method, inputting the semantic component structure as feature vectors into the classifier. The classifier outputs the probability distribution of the four categories, selecting the category with the highest probability as the final intent classification result. Simultaneously, the confidence score and the timestamp information of the voice input are recorded to form a complete voice intent category label data structure. The collection of intent-operation behavior correspondence data is achieved through the interface interaction monitoring module. This module continuously monitors the user's subsequent operations after issuing a voice command. Click operations record mouse or remote control button events and target element identifiers; swipe operations record the starting position, movement direction, and distance parameters of the gesture trajectory; and dwell time is calculated by measuring the duration of the mouse pointer or focus on a specific interface element. A mapping relationship is established between various interaction behavior data and the corresponding voice intent category labels.
[0022] The temporal correlation model is constructed based on Markov chain theory, using the voice intent category as the initial state in the state space and the user's interface operation sequence as the subsequent state transition path. The model construction process first statistically analyzes the frequency of transitions from each intent category to different operation sequences in historical data, establishing a state transition frequency statistics table. Then, the frequency data is normalized, and the probability distribution of various transition paths is calculated to form a transition probability matrix. Each element in the matrix represents the probability value of transitioning from a specific intent state to a specific operation state. To avoid the zero-probability problem caused by data sparsity, the probability matrix is optimized using the Laplace smoothing algorithm. A smoothing parameter is added to the original frequency, and the probability distribution is recalculated to ensure that all possible state transitions have non-zero probability values. The training and optimization process of historical user interaction records employs the maximum likelihood estimation method. Model parameters are continuously adjusted through iterative calculations to achieve the optimal fit between the model and historical data. During training, the transition probability parameters and the correlation weight coefficients between each state are updated simultaneously. The weight coefficients reflect the importance and reliability level of different state transition paths.
[0023] In one specific embodiment, the process of constructing a temporal correlation model based on the intent-operation behavior correspondence data can specifically include the following steps: The intent-operation behavior correspondence data is sorted in chronological order to establish a directed graph structure of intent state nodes and operation behavior state nodes, resulting in a temporal state transition graph. Based on the temporal state transition graph, the frequency of transitions from each intent category to different operation sequences is statistically analyzed, and the state transition frequency matrix is calculated to obtain the original transition statistics. The original transfer statistics are normalized, and the transfer frequency percentage is calculated by row to obtain the operation sequence transfer probability distribution corresponding to each intent category. Based on the transition probability distribution, a Markov chain model structure is constructed, with the intention state as the starting node and the operation sequence state as the target node, thus obtaining the temporal correlation model framework. The transition probability parameters in the aforementioned temporal association model framework are smoothed, and the Laplace smoothing algorithm is used to avoid the zero probability problem, thus obtaining the speech intent temporal association model.
[0024] Specifically, the temporal sorting of the intent-action correspondence data is performed by arranging the timestamp fields in the data records in ascending order. This ensures that each record is organized according to the chronological order of events. During the sorting process, a data index mapping table is simultaneously established to record the correspondence between the original data positions and the sorted positions, facilitating subsequent data traceability and verification. The directed graph structure is built by using voice intent categories as starting nodes and specific user interface actions as target nodes. The connections between nodes are represented by directed edges, where the direction of the edge indicates the causal relationship from the voice intent to the action. The initial weight of each edge is set to the number of times that relationship occurs in the data, forming a complete temporal state transition graph data structure.
[0025] The statistical analysis of the temporal state transition graph involves traversing all directed edges in the graph and counting the specific number of transitions from each intent category to different operation sequences. During the statistical process, calculations are performed in groups according to intent categories, with each intent category corresponding to a statistical subset. This subset records the frequency data of transitions from that intent to various operation behaviors. The calculation of the state transition frequency matrix organizes the statistical results into a two-dimensional matrix. Rows in the matrix represent different speech intent categories, columns represent different operation sequence types, and the value of each element in the matrix represents the corresponding intent-operation transition frequency. Once the matrix is constructed, it forms a complete representation of the original transition statistics. The normalization process standardizes the original transition statistics row by row. The normalization operation for each row is to divide each element in the row by the sum of all elements in that row. The formula is that the normalized value equals the original frequency divided by the total frequency of the row. Normalization ensures that the sum of each row equals 1, forming a standard probability distribution. The operation sequence transition probability distribution corresponding to each intent category reflects the relative probability of that intent transitioning to different operations.
[0026] The Markov chain model structure is constructed based on transition probability distribution data. A Markov chain is a stochastic process model whose core characteristic is that the next state of the system depends only on the current state and is independent of historical states. In this invention, the current state refers to the user's voice intent category, and the next state refers to the user's interface operation behavior. The model structure is set up by using the voice intent state as the starting node of the Markov chain and the operation sequence state as the target node. The transition probabilities between nodes are directly obtained using normalized probability distribution values. The completed temporal correlation model framework includes core components such as state space definition, transition probability matrix, and initial state distribution. This framework can predict the most likely interface operation sequence executed by the user based on the current voice intent state. The application of the Laplace smoothing algorithm optimizes the zero-probability problem in the transition probability matrix. The zero-probability problem refers to situations where certain intent-operation transition combinations have never occurred in historical data, resulting in a probability of zero. Such situations can lead to calculation errors or unreasonable results during model prediction.
[0027] The Laplace smoothing algorithm adds a small smoothing parameter, usually set to 1, to each transition path based on the original frequency data. Then, it recalculates the transition probability. The smoothed probability is calculated as follows: the smoothed probability equals the original frequency plus the smoothing parameter divided by the total row frequency plus the smoothing parameter multiplied by the total number of columns. This approach ensures that all possible state transitions have non-zero probability values, preventing computational anomalies when the model encounters unseen intention-operation combinations, while maintaining the basic characteristics of the probability distribution without significant changes.
[0028] In one specific embodiment, step S2 includes: The intent categories in the voice interaction prediction dataset are labeled with educational scenarios, and classified into beginner tutorials, function browsing, content search, and module switching to obtain scenario-labeled data. The cognitive load index is calculated based on user operation pause time and error rate. The cognitive load is then quantified and classified into three levels: low, medium, and high, to obtain user cognitive load level data. The scene annotation data and user cognitive load level data are weighted and encoded, and the weight coefficients for educational scenarios and cognitive load are set respectively to obtain a double-weighted encoding vector. The probability of each speech intent affecting the interface elements is calculated based on the dual-weighted encoding vector, and a two-dimensional matrix structure of behavioral intent column and interface element column is constructed to obtain the intent-element probability matrix.
[0029] Specifically, the educational scenario annotation process categorizes and labels the intent categories in the voice interaction prediction dataset. The beginner tutorial scenario annotation covers users seeking basic operational guidance, such as introductory queries like "How to use it?" and "How to operate it?". The function browsing scenario annotation includes users exploring platform functions, such as discovery-type operations like "What functions are available?" and "Show all options." The content search scenario annotation targets users searching for specific content, such as targeted queries like "Find courses" and "Search videos." The module switching scenario annotation corresponds to users navigating between different functional modules, such as navigation-type operations like "Return to homepage" and "Enter settings." The annotation process employs a combination of keyword matching and semantic analysis to establish a scenario keyword library. Through word matching and contextual analysis, the scenario affiliation of each voice intent is determined, forming structured data containing scenario labels.
[0030] The cognitive load index is calculated based on objective indicators of user behavior for quantitative evaluation. The pause time is determined by monitoring the time interval between issuing a voice command and executing the specific operation; a longer pause time indicates more thinking time and a heavier cognitive load. The error rate is calculated by statistically analyzing the ratio of the number of erroneous operations within a specific time window to the total number of operations. Erroneous operations include clicking invalid buttons, selecting incorrect menu items, and repeatedly performing the same operation. The cognitive load index is calculated using a weighted summation method, combining standardized pause time and error rate according to preset weights. The pause time weight is set to 0.6, and the error rate weight is set to 0.4. The calculated cognitive load index is mapped to three levels: low (0-0.3), medium (0.3-0.7), and high (0.7-1.0). The level classification results form the user cognitive load level data.
[0031] The dual-weighted encoding process vectorizes and weights scene annotation data and cognitive load level data. Scene annotation data is converted into one-hot encoded vectors: the beginner tutorial scene is encoded as [1,0,0,0], the function browsing scene as [0,1,0,0], the content search scene as [0,0,1,0], and the module switching scene as [0,0,0,1]. Cognitive load level data also uses one-hot encoding: low level is encoded as [1,0,0], medium level as [0,1,0], and high level as [0,0,1]. The weight coefficients for education scenes and cognitive load are optimized based on the application characteristics of the IPTV education platform. The weight coefficient for education scenes is set to 0.7 to reflect the dominant role of scene type in influencing interface elements, and the weight coefficient for cognitive load is set to 0.3 to reflect the moderating role of user cognitive state. The dual-weighted encoding vector is formed by multiplying the scene encoding vector by the scene weight coefficient and the cognitive load encoding vector by the cognitive load weight coefficient, and then concatenating the two weighted results into a 7-dimensional comprehensive encoding vector.
[0032] The construction of the intent-element probability matrix is based on calculating the probability of each voice intent's influence on different interface elements using a double-weighted encoded vector. The probability of influence is calculated using a similarity measurement method, obtaining the influence strength value by calculating the cosine similarity between the double-weighted encoded vector of the voice intent and the feature vector of the interface element. Interface element feature vectors are pre-encoded based on factors such as the element's functional attributes, location features, and interaction frequency. The feature vectors of menu buttons emphasize navigation function attributes, the feature vectors of content display areas emphasize information presentation attributes, and the feature vectors of operation controls emphasize interaction function attributes. Each interface element corresponds to a fixed-dimensional feature vector representation. The construction of the two-dimensional matrix structure uses all voice intents as the row dimension and all interface elements as the column dimension. The value at each position in the matrix represents the probability of the corresponding row intent's influence on the corresponding column element. The probability values are obtained by normalizing the cosine similarity calculation results to ensure that the sum of the probability values in each row equals 1.
[0033] In one specific embodiment, step S3 includes: Based on the intent-element probability matrix, the probability values corresponding to each interface element are extracted, and real-time matching calculations are performed in conjunction with the current user's voice intent category to obtain a real-time importance score for the interface element. Based on the real-time importance score of the interface elements, a visibility level threshold is set, and the interface elements are divided into five visibility levels: fully visible, highlighted, normally displayed, faded, and hidden, to obtain the element visibility classification result. A progressive adjustment algorithm is applied to the visibility grading results of the elements, and the step size of the transparency change and the size scaling ratio are set to control the smooth transition of the visual attributes of the interface elements, so as to obtain the element visibility adjustment parameters. The element visibility adjustment parameters are updated synchronously with the position weights and interaction response priorities of the interface elements to establish a three-dimensional adjustment control vector, thereby obtaining a multi-level progressive adjustment mechanism configuration.
[0034] Specifically, the calculation of the real-time importance score of interface elements is based on real-time data extraction and matching operations using an intent-element probability matrix. When the user's current voice intent category is detected, all probability values corresponding to that intent are extracted from the probability matrix. Each value represents the intensity of the intent's influence on a specific interface element. The real-time matching calculation process weights and sums these probability values with the basic importance weight of the current interface element. The weighted summation method is to multiply the probability value by the basic weight of the corresponding element and then sum them up. The basic weight reflects the inherent importance of the interface element in the IPTV education platform. For example, the basic weight of the navigation menu is higher, while the basic weight of decorative icons is lower. The final weighted summation result is normalized to form a real-time importance score value between 0 and 1.
[0035] The determination process for element visibility grading results is based on threshold division and level mapping using real-time importance scores. The visibility level thresholds are set using an equal-interval division method: fully visible level corresponds to a score range of 0.8 to 1.0, highlighted level corresponds to a score range of 0.6 to 0.8, normally visible level corresponds to a score range of 0.4 to 0.6, faded level corresponds to a score range of 0.2 to 0.4, and hidden level corresponds to a score range of 0 to 0.2. Each interface element is automatically assigned to the corresponding visibility level based on its real-time importance score. A hard threshold judgment method is used during the grading process. When the score value falls within a specific range, it is directly mapped to the corresponding level, avoiding uncertainty caused by fuzzy judgment. The grading results are recorded in the form of level identifiers: fully visible is recorded as V5, highlighted as V4, normally visible as V3, faded as V2, and hidden as V1.
[0036] The progressive adjustment algorithm controls the smooth transition of element visibility grading results. Its core purpose is to avoid abrupt visual jumps when interface elements switch between different visibility levels. The step size for transparency change is calculated based on the difference between the target level and the current level; the larger the level difference, the smaller the step size, ensuring a smoother transition. Specifically, the absolute value of the level difference is used as the denominator, and a fixed value of 0.1 is used as the numerator to calculate the transparency change per frame. A similar method is used for setting the size scaling ratio: the target size for highlighted levels is 1.2 times the original size, normal display is 1.0 times, and faded display is 0.8 times. Size changes also use a progressive approach; the scaling change per frame is calculated by dividing the difference between the target ratio and the current ratio by the preset number of transition frames. Element visibility adjustment parameters include key data such as the target transparency value, target size scaling value, step size, and transition time.
[0037] The process of establishing a three-dimensional adjustment control vector synchronously integrates visibility adjustment parameters with the position weights and interaction response priorities of interface elements. Position weights reflect the spatial importance of interface elements in the current layout. The calculation of position weights is based on factors such as the element's coordinate position on the interface, its distance from the visual focus, and the density of surrounding elements. The visual focus is typically located at the golden ratio point of the interface; elements closer to the visual focus have higher position weights. Interaction response priorities reflect the importance of an element's response to user actions. Response priorities are set based on the element's functional attributes and the matching degree with the current voice intent. Elements with strong functionality and a high degree of matching with the current intent have higher priorities, while decorative elements or elements unrelated to the current intent have lower priorities. The construction of the three-dimensional adjustment control vector uses visibility parameters as the first dimension, position weights as the second dimension, and interaction response priorities as the third dimension, forming a triplet data structure. The values of each dimension in the vector are normalized to ensure the uniformity of the numerical range. A multi-level progressive adjustment mechanism is configured to manage the three-dimensional adjustment control vectors of all interface elements collectively, establishing a vector index table to quickly locate the adjustment parameters of specific elements.
[0038] In one specific embodiment, step S4 includes: Based on the multi-level progressive adjustment mechanism, the space utilization rate of the current interface is calculated, and the ratio of the occupied area to the total display area is statistically analyzed to obtain the interface space utilization rate data. Based on the interface space utilization data and the importance scores of interface elements, visual balance constraints are set, and constraint functions for element density distribution uniformity and visual center of gravity stability are established to obtain spatial layout constraint parameters. The spatial layout constraint parameters are input into a dynamic reconstruction algorithm to redistribute the element positions. Visual focus areas are assigned to high-importance scoring elements, and edge areas are assigned to low-importance scoring elements, thus obtaining an element position reconstruction scheme. Based on the element position reconstruction scheme, a user visual scanning path prediction model is constructed to calculate the shortest visual distance and scanning time from the current point of interest to the target element, thereby obtaining an adaptive interface space allocation scheme.
[0039] Specifically, the calculation of interface space utilization data is based on the area statistics of each element in the multi-level progressive adjustment mechanism configuration. The area occupied by each interface element is obtained by geometric calculation through its current size and position coordinates. The size data comes from the width and height attributes of the element, and the position coordinates include the horizontal and vertical coordinate values of the upper left corner of the element. The area calculation adopts the rectangle area formula, that is, width multiplied by height. The total area of the occupied area is obtained by accumulating the occupied areas of all visible elements. The total display area is determined by the screen resolution of the IPTV interface. The common 1920×1080 resolution corresponds to a total area of 2,073,600 pixels. The space utilization rate is calculated by dividing the occupied area by the total display area to obtain the ratio data. This ratio reflects the current degree of congestion and space utilization efficiency of the interface.
[0040] The process of setting spatial layout constraint parameters combines interface space utilization data and element importance scores to establish multiple constraints. The element density distribution uniformity constraint aims to prevent excessive concentration of interface elements in a certain area, causing visual crowding. This constraint is achieved by dividing the interface into multiple grid regions and counting the number of elements in each region. The ideal density distribution requires that the difference in the number of elements in each region does not exceed a preset threshold. The threshold is calculated based on the average of the total number of elements divided by the number of grids. The visual center of gravity stability constraint ensures the visual balance of the interface. The visual center of gravity is calculated using a weighted average method. The position coordinates of each element are multiplied by its importance score, summed, and then divided by the sum of all element importance scores to obtain the visual center of gravity coordinates. The stability constraint requires that the visual center of gravity be close to the geometric center of the interface. The degree of deviation is quantified by calculating the Euclidean distance between the center of gravity coordinates and the center coordinates; the smaller the distance, the better the visual balance. The constraint function combines the density uniformity index and the center of gravity stability index in a weighted combination to form a comprehensive evaluation function.
[0041] The dynamic reconstruction algorithm redistributes element positions based on spatial layout constraints. It employs a heuristic search method to find the optimal layout that satisfies the constraints. The search process first divides the interface into three levels: a visual focus area, a secondary focus area, and edge areas. The visual focus area is located near the golden ratio point of the interface, covering one-third of the total area. The secondary focus areas surround the visual focus area, covering half of the total area. The edge areas are located at the four edges of the interface. Elements with high importance scores are preferentially assigned to the visual focus area, following a descending order of importance. The position of each element is determined by finding the best available location within the target area. The criteria for the best location include distance from the area center, overlap with already assigned elements, and compliance with constraints. Elements with low importance scores are assigned to the edge areas with a similar but lower priority. The reconstruction process also considers the logical relationships between elements, assigning elements with related functions to adjacent positions for easy user operation.
[0042] The user visual scanning path prediction model is built upon eye-tracking research and cognitive psychology principles. This model models user visual scanning behavior as a path planning problem involving movement from the current point of attention to a target element. Scanning paths typically follow left-to-right and top-to-bottom reading habits and are attracted to elements with high contrast, large size, and dynamic changes. The model uses an attention map to represent the visual attractiveness of each location on the interface, with attention values calculated by weighting the element's visual attributes and importance scores. The shortest visual distance is calculated using the Manhattan distance formula, which is the sum of the absolute values of the differences in the horizontal and vertical axes. This distance measurement method aligns with the actual path characteristics of human visual scanning. Scanning time is calculated by dividing the visual distance by the average scanning speed, which is determined based on statistical analysis of user group eye-tracking data. The model also considers the impact of visual obstacles on the scanning path; when a target element is occluded by other elements, it increases scanning time and path complexity.
[0043] In one specific embodiment, step S5 includes: Based on the adaptive interface space allocation scheme, monitor the user's operation pause time and error operation frequency, calculate the attention concentration index, and obtain user attention distraction state data; The user's distraction status data is compared with a preset attention concentration threshold. When the attention concentration is lower than the threshold, an enhanced guidance branch is triggered; otherwise, a lightweight prompt branch is triggered, thus obtaining the guidance mode selection result. Based on the guidance mode selection result, the visual guidance intensity parameter is set. The enhanced guidance mode uses bright flashing and arrow indication, while the lightweight prompt mode uses color change and transparency adjustment to obtain a personalized guidance strategy configuration. The personalized guidance strategy configuration is applied to the final display state of interface elements, and the visual attributes and interactive response mechanisms of the elements are adjusted simultaneously to obtain personalized and differentiated display results.
[0044] Specifically, user attention distraction data is acquired based on user behavior monitoring after the implementation of the adaptive interface space allocation scheme. Operation pause time is measured by recording the time interval between the completion of interface reconstruction and the execution of the next operation. The measurement start point is the end of the interface element position adjustment animation, and the end point is the detection of a user mouse click, key input, or voice command. Measurement accuracy reaches the millisecond level to ensure data accuracy. Error operation frequency statistics include the number of times unexpected behaviors such as clicking invalid areas, repeatedly clicking the same element, and selecting the wrong menu item occur. The statistical time window is set to the first 30 seconds after interface reconstruction. During this period, users typically complete the adaptation to the new layout and the target operation. Error operation identification is based on comparison with a predefined correct operation path; behaviors deviating from the expected path are recorded as errors. The attention concentration index is calculated using an inverse correlation method. The longer the pause time, the more scattered the attention; the more frequent the error operation, the heavier the cognitive load. The calculation process converts the pause time inversely by dividing the preset standard time by the actual pause time. The error operation frequency is also converted inversely. The two conversion results are then weighted and averaged to obtain the final attention concentration value. In the weight allocation, the pause time accounts for 0.6 and the error operation frequency accounts for 0.4.
[0045] The process of determining the guidance mode selection result is based on a comparison between the attention concentration index and a preset threshold. The attention concentration threshold is set based on statistical analysis of historical behavioral data of IPTV education platform users. The median is calculated using attention performance data from a large sample of users as the benchmark threshold, typically set to 0.6. When a user's real-time attention concentration is below this threshold, it indicates that the user is in a state of distraction and needs stronger visual guidance to help them quickly locate the target element. At this time, the enhanced guidance branch is triggered, and its activation is recorded as an "Enhanced" status in the guidance mode selection result. When the attention concentration is higher than or equal to the threshold, it indicates that the user's attention is relatively focused, and excessive visual guidance may cause interference. At this time, the lightweight prompt branch is triggered, and it is recorded as a "Subtle" status in the selection result. The judgment process uses a hard threshold method to avoid uncertainty in the guidance strategy caused by ambiguous states. The selection result simultaneously records the attention concentration value and threshold data at the time of judgment for subsequent strategy optimization.
[0046] The process of configuring personalized guidance strategies determines specific visual guidance parameters based on the guidance mode selection results. The visual guidance intensity parameters of the enhanced guidance mode include control variables such as the frequency of highlight flashing, brightness gain, and flashing duration. The frequency of highlight flashing is set to 2 times per second to ensure that it attracts the user's attention without causing visual fatigue. The brightness gain is set to 150% of the original brightness to highlight the target element. The flashing duration is set to 3 seconds to give the user sufficient reaction time. The arrow indication function is implemented by generating dynamic arrow graphics near the target element. The arrow color uses high-contrast red or orange, and the size is adaptively adjusted according to the element size. The animation effect adopts a gradual pointing method from far to near. The lightweight prompt mode uses relatively mild visual guidance intensity parameters. Color changes are achieved by adjusting the border color or background color of the target element. Color selection avoids overly bright colors and usually uses relatively soft tones such as blue or green. The transparency adjustment changes the transparency of the target element from the current value to complete opacity. The adjustment process is gradual to avoid abrupt changes. The adjustment time is set to 1 second to ensure that users can perceive the change but will not be overly disturbed. The strategy configuration data is stored in the form of parameter groups, which contain complete information such as guidance type, intensity level, duration, color value, and transparency value.
[0047] The application process of the final display state of interface elements integrates and updates the personalized guidance strategy configuration with the existing visual attributes of the elements. Synchronous adjustments to visual attributes include coordinated changes across multiple dimensions such as element color, transparency, size, position, and animation effects. During these adjustments, priority is given to ensuring the prominence of the guidance effect while also considering the overall aesthetic harmony of the interface. When multiple elements require guidance simultaneously, a priority ranking mechanism is used. Priority is calculated based on the element's importance score and its match with the current voice intent. Higher-priority elements receive stronger guidance effects, while lower-priority elements receive reduced guidance intensity to avoid visual confusion. Synchronous updates to the interaction response mechanism include adjustments to element click response speed, hover effects, and selected states. The click response priority of guided elements is temporarily increased, response time is reduced to half the normal value, hover effects are enhanced for more obvious visual feedback, and the visual indicator of the selected state is more prominent. These adjustments ensure a smoother and more natural user experience under guidance prompts.
[0048] The above describes the method for differentiated display of IPTV interface elements in the embodiments of this application. The following describes the system for differentiated display of IPTV interface elements in the embodiments of this application. Please refer to [link / reference]. Figure 2 One embodiment of the IPTV interface element differentiation display system in this application includes: The acquisition module is used to acquire the audio features and semantic components of user voice commands, establish a temporal correlation model between voice intent category and subsequent interface operation sequence, and train and generate a voice interaction prediction dataset based on historical interaction data. The encoding module is used to weight the voice interaction prediction dataset according to the educational scenario and the user's cognitive load, and construct an intent-element probability matrix; The synchronization module is used to calculate the real-time importance score of interface elements based on the intent-element probability matrix, and adopts a multi-level progressive adjustment mechanism to synchronously control the visibility of elements, position weight and interaction response priority. The generation module is used to dynamically reconstruct the element layout and predict the user's visual scanning path based on the interface space utilization and visual balance constraints, and generate an adaptive interface space allocation scheme. The output module is used to detect the user's distracted state. When the user's concentration is below a preset threshold, an enhanced guidance mode is activated. Otherwise, a lightweight prompt strategy is adopted to output personalized and differentiated display results.
[0049] above Figure 2 The IPTV interface element differentiation display system in this embodiment of the invention is described in detail from the perspective of modular functional entities. The IPTV interface element differentiation display device in this embodiment of the invention is described in detail from the perspective of hardware processing.
[0050] Reference Figure 3 This invention also provides an IPTV interface element differentiation display device, which can be a server, and its internal structure can be as follows: Figure 3 As shown, the IPTV interface element differentiation display device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor, designed as a computer, provides computing and control capabilities. The memory of the IPTV interface element differentiation display device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the IPTV interface element differentiation display device stores the data corresponding to this embodiment. The network interface of the IPTV interface element differentiation display device is used for communication with external terminals via network connection. When the computer program is executed by the processor, it implements the above-described method.
[0051] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the IPTV interface element differentiation display device to which the present invention is applied.
[0052] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the IPTV interface element differentiation display method.
[0053] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0054] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an IPTV interface element differentiation display device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0055] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for differentiated display of IPTV interface elements, characterized in that, The method includes: Step S1: Collect the audio features and semantic components of the user's voice commands, establish a temporal correlation model between the voice intent category and the subsequent interface operation sequence, and train and generate a voice interaction prediction dataset based on historical interaction data. Step S2: Weight the voice interaction prediction dataset according to the educational scenario and the user's cognitive load to construct an intent-element probability matrix; Step S3: Calculate the real-time importance score of the interface elements based on the intent-element probability matrix, and use a multi-level progressive adjustment mechanism to synchronously control the element visibility, position weight and interaction response priority; Step S4: Based on the interface space utilization and visual balance constraints, dynamically reconstruct the element layout and predict the user's visual scanning path to generate an adaptive interface space allocation scheme. Step S5: Detect the user's distracted state. When the user's concentration is below the preset threshold, activate the enhanced guidance mode. Otherwise, adopt a lightweight prompt strategy and output personalized and differentiated display results.
2. The method for differentiated display of IPTV interface elements according to claim 1, characterized in that, Step S1 includes: The audio features of user voice commands are extracted through the iFlytek voice assistant interface to obtain audio feature data including spectrum, pitch, and speech rate. Semantic component analysis of the audio feature data is performed based on natural language processing algorithms to obtain the semantic component structure of keywords, action words, and target objects. The semantic component structure is classified into navigation, operation, query and setting categories to obtain voice intent category labels with timestamps and confidence levels; Collect the user interface operation sequence after the occurrence of the voice intent category label, record the interaction behaviors of clicking, swiping, and dwell time, and obtain the intention-operation behavior correspondence data; Based on the intent-operation behavior correspondence data, a temporal association model is constructed, and the transition probability matrix between different intent categories and subsequent operation sequences is calculated to obtain the speech intent temporal association model. Historical user interaction records are input into the speech intent temporal association model for training and optimization. The transition probability parameters and association weight coefficients are updated to obtain the speech interaction prediction dataset.
3. The method for differentiated display of IPTV interface elements according to claim 2, characterized in that, The step of constructing a temporal correlation model based on the intent-operation correspondence data, calculating the transition probability matrix between different intent categories and subsequent operation sequences, and obtaining the speech intent temporal correlation model includes: The intent-operation behavior correspondence data is sorted in chronological order to establish a directed graph structure of intent state nodes and operation behavior state nodes, resulting in a temporal state transition graph. Based on the temporal state transition graph, the frequency of transitions from each intent category to different operation sequences is statistically analyzed, and the state transition frequency matrix is calculated to obtain the original transition statistics. The original transfer statistics are normalized, and the transfer frequency percentage is calculated by row to obtain the operation sequence transfer probability distribution corresponding to each intent category. Based on the transition probability distribution, a Markov chain model structure is constructed, with the intention state as the starting node and the operation sequence state as the target node, thus obtaining the temporal correlation model framework. The transition probability parameters in the aforementioned temporal association model framework are smoothed, and the Laplace smoothing algorithm is used to avoid the zero probability problem, thus obtaining the speech intent temporal association model.
4. The method for differentiated display of IPTV interface elements according to claim 1, characterized in that, Step S2 includes: The intent categories in the voice interaction prediction dataset are labeled with educational scenarios, and classified into beginner tutorials, function browsing, content search, and module switching to obtain scenario-labeled data. The cognitive load index is calculated based on user operation pause time and error rate. The cognitive load is then quantified and classified into three levels: low, medium, and high, to obtain user cognitive load level data. The scene annotation data and user cognitive load level data are weighted and encoded, and the weight coefficients for educational scenarios and cognitive load are set respectively to obtain a double-weighted encoding vector. The probability of each speech intent affecting the interface elements is calculated based on the dual-weighted encoding vector, and a two-dimensional matrix structure of behavioral intent column and interface element column is constructed to obtain the intent-element probability matrix.
5. The method for differentiated display of IPTV interface elements according to claim 1, characterized in that, Step S3 includes: Based on the intent-element probability matrix, the probability values corresponding to each interface element are extracted, and real-time matching calculations are performed in conjunction with the current user's voice intent category to obtain a real-time importance score for the interface element. Based on the real-time importance score of the interface elements, a visibility level threshold is set, and the interface elements are divided into five visibility levels: fully visible, highlighted, normally displayed, faded, and hidden, to obtain the element visibility classification result. A progressive adjustment algorithm is applied to the visibility grading results of the elements, and the step size of the transparency change and the size scaling ratio are set to control the smooth transition of the visual attributes of the interface elements, so as to obtain the element visibility adjustment parameters. The element visibility adjustment parameters are updated synchronously with the position weights and interaction response priorities of the interface elements to establish a three-dimensional adjustment control vector, thereby obtaining a multi-level progressive adjustment mechanism configuration.
6. The method for differentiated display of IPTV interface elements according to claim 1, characterized in that, Step S4 includes: Based on the multi-level progressive adjustment mechanism, the space utilization rate of the current interface is calculated, and the ratio of the occupied area to the total display area is statistically analyzed to obtain the interface space utilization rate data. Based on the interface space utilization data and the importance scores of interface elements, visual balance constraints are set, and constraint functions for element density distribution uniformity and visual center of gravity stability are established to obtain spatial layout constraint parameters. The spatial layout constraint parameters are input into a dynamic reconstruction algorithm to redistribute the element positions. Visual focus areas are assigned to high-importance scoring elements, and edge areas are assigned to low-importance scoring elements, thus obtaining an element position reconstruction scheme. Based on the element position reconstruction scheme, a user visual scanning path prediction model is constructed to calculate the shortest visual distance and scanning time from the current point of interest to the target element, thereby obtaining an adaptive interface space allocation scheme.
7. The method for differentiated display of IPTV interface elements according to claim 1, characterized in that, Step S5 includes: Based on the adaptive interface space allocation scheme, monitor the user's operation pause time and error operation frequency, calculate the attention concentration index, and obtain user attention distraction state data; The user's distraction status data is compared with a preset attention concentration threshold. When the attention concentration is lower than the threshold, an enhanced guidance branch is triggered; otherwise, a lightweight prompt branch is triggered, thus obtaining the guidance mode selection result. Based on the guidance mode selection result, the visual guidance intensity parameter is set. The enhanced guidance mode uses bright flashing and arrow indication, while the lightweight prompt mode uses color change and transparency adjustment to obtain a personalized guidance strategy configuration. The personalized guidance strategy configuration is applied to the final display state of interface elements, and the visual attributes and interactive response mechanisms of the elements are adjusted simultaneously to obtain personalized and differentiated display results.
8. A differentiated display system for IPTV interface elements, characterized in that, For implementing the IPTV interface element differentiation display method as described in any one of claims 1-7, the IPTV interface element differentiation display system comprises: The acquisition module is used to acquire the audio features and semantic components of user voice commands, establish a temporal correlation model between voice intent category and subsequent interface operation sequence, and train and generate a voice interaction prediction dataset based on historical interaction data. The encoding module is used to weight the voice interaction prediction dataset according to the educational scenario and the user's cognitive load, and construct an intent-element probability matrix; The synchronization module is used to calculate the real-time importance score of interface elements based on the intent-element probability matrix, and adopts a multi-level progressive adjustment mechanism to synchronously control the visibility of elements, position weight and interaction response priority. The generation module is used to dynamically reconstruct the element layout and predict the user's visual scanning path based on the interface space utilization and visual balance constraints, and generate an adaptive interface space allocation scheme. The output module is used to detect the user's distracted state. When the user's concentration is below a preset threshold, an enhanced guidance mode is activated. Otherwise, a lightweight prompt strategy is adopted to output personalized and differentiated display results.
9. A device for differentiating IPTV interface elements, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the IPTV interface element differentiation display method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it causes the processor to execute the IPTV interface element differentiation display method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-mode interaction control method and system for smart blackboard
CN118655979A
Multi-modal user intention understanding and personalized shopping guide generation method and system
CN120106942A
Vision generation method and device based on semantic association modeling, equipment and medium
CN120542428A
AI-based digital media interface design optimization method
CN120723235A
Intention classification method and apparatus, electronic device, and computer-readable storage medium
WO2023065544A1
Cited By
Dynamic information position determination method
CN121636784A
Exhibition and exhibition voice drive generation method and system
CN122266359A