Method and system for differentiating display of iptv interface elements
By establishing a temporal correlation model between voice intent and interface operation and an intent-element probability matrix, combined with dynamic interface reconstruction and attention detection, the problem of insufficient voice intent perception in IPTV interface element display is solved, and personalized, real-time interface adjustment and optimization are realized.
Patent Information
- Application Number
- CN202511463863.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Existing IPTV interface element display technologies suffer from insufficient voice intent perception, weak context perception, and simplified interface adjustment mechanisms in multimodal voice interaction. This leads to a discrepancy between the interface response and user expectations, making it unable to adapt to real-time changes in users' personalized needs.
By collecting the audio features and semantic components of user voice commands, a temporal correlation model between voice intent categories and subsequent interface operation sequences is established. An intent-element probability matrix is constructed, and a multi-level progressive adjustment mechanism is adopted to control element visibility and interaction response priority. The interface layout is dynamically reconstructed and the user's visual scanning path is predicted. Personalized display is achieved by combining an attention detection mechanism.
It achieves precise interface response and dynamic adjustment, improves the ability to perceive voice intent, adapts to real-time changes in user status, optimizes interface layout and element visibility, and enhances user experience.
Smart Images

Figure CN120935418B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to an IPTV interface element differential display method and system. BACKGROUND
[0002] With the rapid development of intelligent voice technology, IPTV interface element display gradually evolves towards voice interaction, and the existing IPTV interface element differential display method mainly relies on preset user classification templates and fixed interface configuration rules, and performs simple interface personalization adjustment by analyzing the basic attributes, viewing history and device information of the user. The traditional method adopts a rule-based interface element control mechanism, matches the corresponding interface template according to the user portrait, and realizes differential display by adjusting the font size, color theme, layout style and other basic attributes, and simultaneously supports basic voice navigation operation in combination with a simple voice recognition function. These technologies have a certain application basis in static personalized display.
[0003] However, the existing technology has significant limitations in handling interface adaptation of multi-modal voice interaction. Firstly, the voice intention perception ability is insufficient. Traditional voice recognition mainly focuses on the literal meaning of the instruction, and lacks the prediction ability of the user's deep intention and subsequent operation demand, resulting in deviation between the interface response and the user's expectation. Secondly, the context perception ability is weak. The existing method is difficult to comprehensively consider multi-dimensional factors such as education scene characteristics, user cognitive load, attention state and the like, and cannot accurately adjust the interface according to the specific situation of the user. Thirdly, the interface element adjustment mechanism is too simple, and lacks dynamic optimization ability based on the real-time state of the user. The interface layout and element visibility adjustment often adopts a fixed mode, and is difficult to adapt to the real-time changes of the user's personalized demand. SUMMARY
[0004] The present application provides an IPTV interface element differential display method and system, which solves the problems of inaccurate voice intention perception and insufficient intelligent degree of interface adaptation in IPTV interface element differential display.
[0005] In a first aspect, the present application provides an IPTV interface element differential display method, which comprises:
[0006] Step S1: collecting the audio features and semantic components of the user voice instruction, establishing a time sequence association model of the voice intention category and the subsequent interface operation sequence, and training a voice interaction prediction data set based on historical interaction data;
[0007] Step S2: weighting and coding the voice interaction prediction data set according to the education scene and the user cognitive load, and constructing an intention-element probability matrix;
[0008] Step S3: calculating the real-time importance score of the interface element based on the intention-element probability matrix, and synchronously controlling the element visibility, position weight and interaction response priority using a multi-level progressive adjustment mechanism;
[0009] Step S4: dynamically reconstructing the element layout and predicting the user visual scanning path according to the interface space utilization and visual balance constraint conditions, and generating an adaptive interface space allocation scheme;
[0010] Step S5: detecting the user attention dispersion state, starting the intensive guidance mode when the attention concentration degree is lower than a preset threshold, otherwise adopting a lightweight prompting strategy, and outputting a personalized differentiated display result.
[0011] In a second aspect, the present application provides an IPTV interface element differentiated display system, which comprises:
[0012] The acquisition module is configured to acquire the audio features and semantic components of the user voice instruction, establish a time sequence association model of the voice intention category and the subsequent interface operation sequence, and train a voice interaction prediction dataset based on historical interaction data;
[0013] The encoding module is configured to weight and encode the voice interaction prediction dataset according to the education scene and the user cognitive load, and construct an intention-element probability matrix;
[0014] The synchronization module is configured to calculate the real-time importance score of the interface element based on the intention-element probability matrix, and synchronously control the element visibility, position weight and interaction response priority using a multi-level progressive adjustment mechanism;
[0015] The generation module is configured to dynamically reconstruct the element layout and predict the user visual scanning path according to the interface space utilization and visual balance constraint conditions, and generate an adaptive interface space allocation scheme;
[0016] The output module is configured to detect the user attention dispersion state, start the intensive guidance mode when the attention concentration degree is lower than a preset threshold, otherwise adopt a lightweight prompting strategy, and output a personalized differentiated display result.
[0017] In a third aspect, an IPTV interface element differentiated display device is provided, which comprises a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to enable the IPTV interface element differentiated display device to perform the above-mentioned IPTV interface element differentiated display method.
[0018] In a fourth aspect, a computer readable storage medium is provided, in which instructions are stored, when executed on a computer, cause the computer to perform the IPTV interface element differentiation display method described above.
[0019] In the technical solution provided in the present application, the audio features and semantic components of the user voice instructions are collected and a timing correlation model is established, which breaks through the technical bottleneck of insufficient voice intention perception in the traditional IPTV interface element display technology. This method can extract multi-dimensional features such as frequency spectrum, pitch, and speech rate from audio signals, accurately identify the real intention of the user by combining natural language processing algorithms, and more importantly, establish the timing correlation between voice intention and subsequent interface operation sequence through Markov chain modeling technology, so that the interface elements can predictively respond to user needs rather than passively wait for operation instructions. The voice interaction prediction dataset generated based on historical interaction data provides a reliable data basis for subsequent intelligent decision-making, solving the core problem of the semantic gap between voice input and interface response in traditional methods. The technical solution of weighting and coding the voice interaction prediction dataset according to the education scene and user cognitive load to build an intention-element probability matrix effectively solves the limitation of existing technologies that cannot consider multiple situational factors comprehensively. This matrix quantifies the influence of different voice intentions on each interface element, providing a scientific basis for precise interface adjustment decisions. In particular, the introduction of education scene weight and cognitive load weight allows the interface adjustment to take into account both the application characteristics of the IPTV education platform and the real-time cognitive state of the user. The technical innovation of calculating the real-time importance score of the interface elements based on the intention-element probability matrix and using a multi-level progressive adjustment mechanism completely changes the rough mode of traditional interface element adjustment. By establishing five refined visibility levels: completely visible, highlighted, normally displayed, faded display, and hidden, and cooperating with progressive transparency and size adjustment algorithms, the smoothness of interface changes and the continuity of user experience are ensured. At the same time, the design of the three-dimensional adjustment control vector realizes the collaborative optimization of visibility, position weight, and interaction response priority.
[0020] The algorithm design of dynamic reconstruction and prediction of user visual scanning path according to interface space utilization and visual balance constraint conditions has significant technical advantages in the field of IPTV interface element differentiated display. The algorithm establishes mathematical constraint models of element density distribution uniformity and visual barycenter stability, ensures that the interface reconstruction meets functional requirements and conforms to visual aesthetic principles, especially the application of user visual scanning path prediction model based on cognitive psychology and eye tracking research results can accurately predict the visual movement trajectory and time cost of users from the current focus point to the target element, providing scientific theoretical guidance for interface layout optimization and significantly shortening the visual search time and operation path length of users. The intelligent mechanism of detecting user attention dispersion state and selecting guide mode based on attention concentration threshold represents an important progress of IPTV interface technology towards cognitive perception. The mechanism accurately judges the cognitive load state of users by monitoring objective indicators such as operation pause time and error frequency of users in real time, automatically starts the intensive guide mode including highlight flashing and arrow indication when the attention concentration is lower than the preset threshold, and adopts the light-weight prompting strategy of color change and transparency adjustment when the attention state is good. This adaptive guide mechanism not only avoids user interference caused by excessive guidance, but also ensures that effective visual guidance can be provided in time when users need help. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.
[0022] Figure 1 An embodiment schematic diagram of the IPTV interface element differentiated display method in the embodiments of the present application;
[0023] Figure 2 An embodiment schematic diagram of the IPTV interface element differentiated display system in the embodiments of the present application;
[0024] Figure 3 A structural schematic block diagram of the IPTV interface element differentiated display device in the embodiments of the present application. DETAILED DESCRIPTION
[0025] The embodiments of the present application provide an IPTV interface element differential display method and system. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] For ease of understanding, the specific flow of the embodiments of the present application is described below. Please refer to Figure 1 One embodiment of the IPTV interface element differential display method in the embodiments of the present application includes:
[0027] Step S1: Collecting audio features and semantic components of user voice instructions, establishing a time sequence association model of voice intent categories and subsequent interface operation sequences, training and generating a voice interaction prediction data set based on historical interaction data;
[0028] Step S2: Weighted coding the voice interaction prediction data set according to education scenarios and user cognitive load, constructing an intent-element probability matrix;
[0029] Step S3: Calculating real-time importance scores of interface elements based on the intent-element probability matrix, and synchronously controlling element visibility, position weight and interaction response priority using a multi-level progressive adjustment mechanism;
[0030] Step S4: Dynamically reconstructing element layout and predicting user visual scanning path according to interface space utilization and visual balance constraint conditions, and generating an adaptive interface space allocation scheme;
[0031] Step S5: Detecting user attention dispersion state, starting a reinforcement guidance mode when the attention concentration degree is lower than a preset threshold, otherwise using a lightweight prompting strategy, and outputting a personalized differential display result.
[0032] It can be understood that the execution subject of the present application can be an IPTV interface element differential display system, and can also be a terminal or a server, which is not limited here. The embodiments of the present application take the server as the execution subject for example.
[0033] Specifically, the voice instruction data collection process captures the user's audio input through the Xunfei voice assistant interface, extracts audio features such as frequency spectrum, pitch, and speech rate, and then uses natural language processing algorithms to analyze the audio features semantically, identify semantic components such as keywords, action words, and target objects, and classify these semantic components into navigation, operation, query, and setting categories to form voice intent category labels with timestamps and confidence levels. The construction of the timing association model is based on Markov chain theory, taking voice intent as a state node and user's subsequent interface operation sequence as a state transition. By counting the frequency of different intent category transitions to various operation sequences in historical data, a transition probability matrix is calculated. After normalization and Laplace smoothing optimization, the matrix forms a voice intent timing association model. Historical user interaction records are input into the model for parameter training and weight adjustment, and finally a voice interaction prediction dataset with predictive ability is generated.
[0034] The weighted encoding process classifies and labels the voice interaction prediction dataset according to the specific application scenarios of the IPTV education platform. Different education scenarios such as novice tutorial, function browsing, content searching, and module switching are assigned corresponding label weights. At the same time, the cognitive load index is calculated by monitoring the user's operation pause time and error rate. Cognitive load is quantified into three levels: low, medium, and high, each level corresponding to a different weight coefficient. The generation of the dual-weighted encoding vector combines the education scenario weight and the cognitive load weight to form a multi-dimensional vector representation. This vector is used to calculate the influence probability of each voice intent on different interface elements, and finally a two-dimensional probability matrix structure is constructed with behavior intent as rows and interface elements as columns. Real-time importance score calculation extracts the corresponding probability values from the intent-element probability matrix and matches them in real time with the user's specific voice intent to obtain the importance value of each interface element. This value is used as the basis for dividing the visibility level, and interface elements are divided into five levels: completely visible, highlighted, normally displayed, faded, and hidden. The gradual adjustment algorithm controls the smooth transition of elements between different visibility levels, and by setting the transparency change step and size scaling ratio, it avoids the sudden change of interface elements affecting user experience.
[0035] The dynamic reconstruction process is based on interface space utilization data and visual balance constraints. The space utilization is obtained by calculating the ratio of occupied area to total display area. The visual balance constraints include two dimensions of element density distribution uniformity and visual barycenter stability. The dynamic reconstruction algorithm allocates a visual focus area for high importance elements and arranges an edge position for low importance elements. A user visual scanning path prediction model is established during the reconstruction process. The model calculates the shortest visual distance and expected scanning time from the current focus point to the target element, ensuring that the interface layout conforms to the user's visual habits. The attention concentration degree is determined by monitoring the pause time and error frequency in the user's operation behavior. When the concentration degree indicator is lower than the preset threshold, the intensive guidance mode is triggered, using significant visual guidance methods such as highlight flashing and arrow indication. When the concentration degree is normal, the light-weight prompting mode is used, with gentle prompting through color change and transparency adjustment. The selection of personalized guidance strategy is dynamically adjusted based on the user's current cognitive state and attention level.
[0036] In a specific embodiment, step S1 comprises:
[0037] Audio feature extraction is performed on the user voice command through the Xunfei voice assistant interface to obtain audio feature data containing frequency spectrum, pitch, and speech rate;
[0038] Semantic component analysis is performed on the audio feature data based on a natural language processing algorithm to obtain a semantic component structure of keywords, action words, and target objects;
[0039] The semantic component structure is classified by intent according to navigation, operation, query, and setting to obtain voice intent category labels with timestamps and confidence levels;
[0040] The user interface operation sequence after the voice intent category labels are collected, and the click, swipe, and dwell time interaction behaviors are recorded to obtain intent-operation behavior correspondence data;
[0041] A timing association model is constructed based on the intent-operation behavior correspondence data, and a transition probability matrix of different intent categories and subsequent operation sequences is calculated to obtain a voice intent timing association model;
[0042] Historical user interaction records are input into the voice intent timing association model for training and optimization to update transition probability parameters and association weight coefficients, and a voice interaction prediction data set is obtained.
[0043] Specifically, the audio feature extraction process of the Xunfei voice assistant interface digitally samples the user's voice input with a sampling frequency of 16 kHz, and then converts the time-domain audio signal into a frequency-domain representation through a fast Fourier transform (FFT) algorithm, extracts the spectral distribution features in the range of 0 Hz to 8 kHz, and simultaneously calculates the pitch variation trajectory of the audio using an autocorrelation function to form the pitch feature. The speech rhythm feature is obtained by calculating the rhythm variation of the speech through frame length analysis and energy detection algorithms. The three types of feature data are organized into a multi-dimensional vector form for storage. The semantic component analysis of the audio feature data by the natural language processing algorithm uses a combination of part-of-speech tagging and syntactic analysis. First, the text obtained through speech recognition is processed for word segmentation, and then the hidden Markov model (HMM) is used for part-of-speech tagging to identify noun, verb, adjective, and other part-of-speech categories. Next, the syntactic relationship between words is determined through dependency syntactic analysis, from which the key words are extracted as the core concepts of the voice content, the action words are extracted as the user's operation intent, and the target objects are extracted as the specific targets of the operation, forming a semantic component representation in the form of a triple structure.
[0044] The voice intent classification process is based on a pre-defined intent classification system. The navigation intent includes page jumping, module switching, and other navigation-related voice commands. The operation intent covers clicking, selecting, confirming, and other specific operation commands. The query intent includes searching, finding, and obtaining information, and other query requests. The setting intent involves configuration modification, parameter adjustment, and other setting operations. The classification algorithm uses the support vector machine (SVM) method, with the semantic component structure as the feature vector input to the classifier. The classifier outputs the probability distribution of the four categories, and selects the category with the highest probability as the final intent classification result. At the same time, the confidence value of the classification and the timestamp information of the voice input are recorded to form a complete voice intent category label data structure. The collection of intent-operation behavior correspondence data is achieved through the interface interaction monitoring module. This module continuously monitors the user's subsequent operation behavior after the user issues a voice command. The click operation records the key event and target element identifier of the mouse or remote control. The sliding operation records the starting position, moving direction, and distance parameters of the gesture track. The dwell time is obtained by calculating the duration of the mouse pointer or focus on a specific interface element. Various interaction behavior data are associated with the corresponding voice intent category label to establish a mapping relationship.
[0045] The construction of the time sequence correlation model is based on Markov chain theory, taking the speech intent category as the starting state in the state space and the user interface operation sequence as the subsequent state transition path. The model construction process first counts the frequency of each intent category in the historical data transferring to different operation sequences, establishes a state transition frequency statistics table, then normalizes the frequency data, calculates the probability distribution of various transition paths, forms a transition probability matrix, and each element in the matrix represents the probability value of transferring from a specific intent state to a specific operation state. To avoid the zero probability problem caused by data sparsity, the Laplace smoothing algorithm is used to optimize the probability matrix, add a smoothing parameter to the original frequency, then recalculate the probability distribution to ensure that all possible state transitions have non-zero probability values. The training and optimization process of the historical user interaction records uses the maximum likelihood estimation method, which continuously adjusts the model parameters through iterative calculation to make the fitting degree of the model to the historical data optimal. The transition probability parameters and the correlation weight coefficients between states are updated during the training process, and the weight coefficients reflect the importance and credibility level of different state transition paths.
[0046] In a specific embodiment, the process of constructing a time sequence correlation model based on the intent-operation behavior correspondence data can specifically include the following steps:
[0047] The intent-operation behavior correspondence data is sorted in time sequence, a directed graph structure of intent state nodes and operation behavior state nodes is established, and a time sequence state transition graph is obtained;
[0048] Based on the time sequence state transition graph, the frequency of each intent category transferring to different operation sequences is counted, a state transition frequency matrix is calculated, and original transition statistical data is obtained;
[0049] The original transition statistical data is normalized, the transition frequency proportion is calculated by row, and the operation sequence transition probability distribution corresponding to each intent category is obtained;
[0050] A Markov chain model structure is constructed according to the transition probability distribution, the intent state is set as the starting node and the operation sequence state is set as the target node, and a time sequence correlation model framework is obtained;
[0051] The transition probability parameters in the time sequence correlation model framework are smoothed, the Laplace smoothing algorithm is used to avoid the zero probability problem, and a speech intent time sequence correlation model is obtained.
[0052] Specifically, the time sequence ordering processing of the intention-operation behavior correspondence data is arranged in ascending order according to the timestamp field in the data record, ensuring that each record is organized in the order of event occurrence, and a data index mapping table is established at the same time to record the correspondence between the original data position and the position after ordering, facilitating subsequent data tracing and verification. The establishment process of the directed graph structure takes the voice intention category as the starting node and the specific interface operation behavior of the user as the target node, and the connection relationship between the nodes is represented by a directed edge. The direction of the directed edge represents the causal relationship from the voice intention to the operation behavior, and the initial value of the weight of the edge is set as the number of occurrences of the association relationship in the data, forming a complete time sequence state transition graph data structure.
[0053] The statistical analysis process of the time sequence state transition graph traverses all the directed edges in the graph, and counts the specific number of transitions of each intention category to different operation sequences. During the counting process, grouping calculation is performed according to the intention category, and each intention category corresponds to a statistical subset, which records the frequency data of the transition of the intention to various operation behaviors. The calculation of the state transition frequency matrix organizes the statistical results into a two-dimensional matrix form, where the rows represent different voice intention categories, the columns represent different operation sequence types, and the numerical value of each element in the matrix represents the corresponding intention-operation transition frequency. After the matrix is constructed, a complete representation of the original transition statistical data is formed. The normalization process performs standardization calculation on the original transition statistical data by row. The normalization operation of each row of data is to divide each element in the row by the sum of all elements in the row. The calculation formula is: the normalized value equals the original frequency divided by the row total frequency. The normalization process ensures that the sum of each row of data is equal to 1, forming a standard probability distribution form. The operation sequence transition probability distribution corresponding to each intention category reflects the relative possibility of the transition of the intention to different operations.
[0054] The construction of the Markov chain model structure is based on the transition probability distribution data. Markov chain is a stochastic process model, and its core feature is that the next state of the system only depends on the current state and is independent of the historical state. In this invention, the current state refers to the voice intention category of the user, and the next state refers to the interface operation behavior of the user. The model structure setting process takes the voice intention state as the starting node of the Markov chain and the operation sequence state as the target node. The transition probability between the nodes directly uses the probability distribution values obtained by normalization. The completed time sequence association model framework includes core components such as state space definition, transition probability matrix, and initial state distribution. This framework can predict the most likely interface operation sequence of the user according to the current voice intention state. The application of Laplace smoothing algorithm optimizes the zero probability problem in the transition probability matrix. The zero probability problem refers to the case where some intention-operation transition combinations never occur in historical data, resulting in a probability of zero. This situation can cause calculation errors or unreasonable results when the model is predicted.
[0055] The specific implementation process of the Laplace smoothing algorithm adds a small smoothing parameter, usually set to 1, to each transition path based on the original frequency data, and then recalculates the transition probability. The formula for calculating the smoothed probability is equal to the original frequency plus the smoothing parameter divided by the total frequency of the row plus the smoothing parameter multiplied by the total number of columns. This processing method ensures that all possible state transitions have non-zero probability values, avoids computational abnormalities when encountering unseen intent-operation combinations, and maintains the basic characteristics of the probability distribution without significant changes.
[0056] In a specific embodiment, step S2 comprises:
[0057] The intent categories in the voice interaction prediction dataset are annotated with education scenarios, classified as novice tutorial, function browsing, content searching, and module switching, to obtain scenario annotated data.
[0058] The cognitive load index is calculated based on user operation pause time and error rate, and the cognitive load is quantified into three levels of low, medium and high, to obtain user cognitive load level data.
[0059] The scenario annotated data and user cognitive load level data are weighted and encoded, and education scenario weight coefficients and cognitive load weight coefficients are set respectively to obtain a double weighted encoding vector.
[0060] The influence probability of each voice intent on the interface element is calculated based on the double weighted encoding vector, a two-dimensional matrix structure of behavior intent column and interface element column is constructed, and an intent-element probability matrix is obtained.
[0061] Specifically, the education scenario annotation process classifies and labels the intent categories in the voice interaction prediction dataset. The novice tutorial scenario annotation covers voice intents for seeking basic operation guidance, such as "how to use", "how to operate", and other introductory queries. The function browsing scenario annotation includes intents for exploring platform functions, such as "what functions are there", "display all options", and other discovery operations. The content searching scenario annotation targets intents for searching specific content, such as "find courses", "search videos", and other specific queries. The module switching scenario annotation corresponds to intents for jumping between different functional modules, such as "back to home page", "enter settings", and other navigation operations. The annotation process uses a combination of keyword matching and semantic analysis to establish a scenario keyword library. The context analysis determines the scene attribution of each voice intent, forming structured data containing scenario labels.
[0062] The calculation of the cognitive load index is based on quantitative evaluation of objective indicators of user operation behavior. The operation pause time is obtained by monitoring the time interval between the user issuing a voice instruction and performing a specific operation. The longer the pause time, the more time the user thinks, the heavier the cognitive load, and the error rate is calculated by counting the ratio of the number of error operations to the total number of operations within a certain time window. Error operations include clicking invalid buttons, selecting incorrect menus, and repeating the same operation. The cognitive load index is calculated using a weighted summation method, combining the standardized pause time and error rate according to the preset weights. The pause time weight is set to 0.6, and the error rate weight is set to 0.4. The calculated cognitive load index is mapped to three levels: low, medium, and high. The low level corresponds to an index range of 0 to 0.3, the medium level corresponds to 0.3 to 0.7, and the high level corresponds to 0.7 to 1.0. The level division result forms the user cognitive load level data.
[0063] The dual-weighted encoding process vectorizes and weights the scene annotation data and cognitive load level data. The scene annotation data is converted into a one-hot encoding vector. The novice tutorial scene is encoded as [1, 0, 0, 0], the function browsing scene is encoded as [0, 1, 0, 0], the content searching scene is encoded as [0, 0, 1, 0], and the module switching scene is encoded as [0, 0, 0, 1]. The cognitive load level data is also one-hot encoded. The low level is encoded as [1, 0, 0], the medium level is encoded as [0, 1, 0], and the high level is encoded as [0, 0, 1]. The education scene weight coefficient and the cognitive load weight coefficient are set based on the application characteristics of the IPTV education platform. The education scene weight coefficient is set to 0.7, reflecting the dominant role of the physical scene type in interface element influence. The cognitive load weight coefficient is set to 0.3, reflecting the adjusting role of the user's cognitive state. The dual-weighted encoding vector is obtained by multiplying the scene encoding vector by the scene weight coefficient and the cognitive load encoding vector by the cognitive load weight coefficient. Then, the two weighted results are vector spliced to form a 7-dimensional comprehensive encoding vector.
[0064] The construction process of the intent-element probability matrix calculates the influence probability of each voice intent on different interface elements based on the double-weighted encoding vector. The influence probability is calculated using a similarity measurement method. The influence strength value is obtained by calculating the cosine similarity between the double-weighted encoding vector of the voice intent and the feature vector of the interface element. The feature vector of the interface element is pre-encoded according to factors such as functional attributes, location characteristics, and interaction frequency. The feature vector of the menu button focuses on navigation function attributes, the feature vector of the content display area focuses on information presentation attributes, and the feature vector of the operation control focuses on interaction function attributes. Each interface element corresponds to a fixed-dimensional feature vector representation. The construction of the two-dimensional matrix structure takes all voice intents as the row dimension of the matrix and all interface elements as the column dimension of the matrix. The numerical value at each position in the matrix represents the influence probability of the corresponding row intent on the corresponding column element. The probability value is obtained by normalizing the cosine similarity calculation result, ensuring that the sum of the probability values in each row is equal to 1.
[0065] In a specific embodiment, step S3 comprises:
[0066] Based on the intent-element probability matrix, the probability values corresponding to each interface element are extracted, and real-time matching calculation is performed in combination with the current user voice intent category to obtain an interface element real-time importance score.
[0067] According to the interface element real-time importance score, a visibility level threshold is set to divide the interface elements into five visibility levels: completely visible, highlighted, normally displayed, faded, and hidden, to obtain an element visibility grading result.
[0068] A progressive adjustment algorithm is applied to the element visibility grading result, a transparency change step and a size scaling ratio are set, and the visual properties of the interface elements are smoothly transitioned to obtain an element visibility adjustment parameter.
[0069] The element visibility adjustment parameter is synchronously updated with the position weight and the interaction response priority of the interface element, a three-dimensional adjustment control vector is established, and a multi-level progressive adjustment mechanism configuration is obtained.
[0070] Specifically, the calculation of the interface element real-time importance score is based on real-time data extraction and matching operation of the intent-element probability matrix. When the current voice intent category of the user is detected, all probability values in the corresponding row of the intent are extracted from the probability matrix, each value representing the influence strength of the intent on a specific interface element. The real-time matching calculation process weights and sums these probability values with the base importance weight of the current interface element. The weighting and summing method is to multiply the probability value by the base weight of the corresponding element and then accumulate. The base weight reflects the inherent importance of the interface element in the IPTV education platform. For example, the base weight of the navigation menu is higher, and the base weight of the decorative icon is lower. The final weighted sum result is normalized to form a real-time importance score value between 0 and 1.
[0071] The determination process of the element visibility classification result is based on threshold division and level mapping of the real-time importance score. The visibility level threshold is set using the equal interval division method. The fully visible level corresponds to a score range of 0.8 to 1.0, the highlight display level corresponds to a score range of 0.6 to 0.8, the normal display level corresponds to a score range of 0.4 to 0.6, the fade display level corresponds to a score range of 0.2 to 0.4, and the hidden level corresponds to a score range of 0 to 0.2. Each interface element is automatically assigned to the corresponding visibility level according to its real-time importance score value. In the classification process, a hard threshold judgment method is used. When the score value falls within a specific interval, it is directly mapped to the corresponding level to avoid uncertainty caused by ambiguous judgment. The classification result is recorded in the form of level identifier. Fully visible is recorded as V5, highlight display is recorded as V4, normal display is recorded as V3, fade display is recorded as V2, and hidden is recorded as V1.
[0072] The application process of the gradual adjustment algorithm controls the smooth transition of the element visibility classification result. The core purpose of this algorithm is to avoid abrupt visual jump effects when interface elements switch between different visibility levels. The step size of the transparency change is set according to the difference between the target level and the current level. The larger the level difference, the smaller the step size, ensuring a smoother transition process. The specific calculation method is to use the absolute value of the level difference as the denominator and a fixed value of 0.1 as the numerator to calculate the transparency change amount of each frame. A similar method is used to set the size scaling ratio. The target size of the highlight display level is 1.2 times the original size, the normal display is 1.0 times, and the fade display is 0.8 times. The size change also adopts a gradual approach. The scaling change amount of each frame is calculated by dividing the difference between the target ratio and the current ratio by the preset transition frame number. The element visibility adjustment parameters include the target value of the transparency, the target value of the size scaling, the change step, the transition time, and other key data.
[0073] The establishment process of the three-dimensional adjustment control vector synchronously integrates the visibility adjustment parameter with the position weight of the interface element and the interaction response priority. The position weight reflects the spatial importance of the interface element in the current layout. The calculation of the position weight is based on factors such as the coordinate position of the element in the interface, the distance from the visual focus, and the density of surrounding elements. The visual focus is usually located at the golden section point position of the interface. The closer the element is to the visual focus, the higher the position weight. The interaction response priority reflects the importance of the response of the element to the user operation. The setting of the response priority is based on the matching degree of the functional attribute of the element and the current voice intent. The element with strong functionality and high matching degree with the current intent has a higher priority, and the decorative element or the element irrelevant to the current intent has a lower priority. The construction of the three-dimensional adjustment control vector takes the visibility parameter as the first dimension, the position weight as the second dimension, and the interaction response priority as the third dimension, forming a three-tuple data structure. The values of each dimension in the vector are normalized to ensure the uniformity of the value range. The multi-level progressive adjustment mechanism configuration performs collection management on the three-dimensional adjustment control vectors of all interface elements, and establishes a vector index table to quickly locate the adjustment parameters of specific elements.
[0074] In a specific embodiment, step S4 comprises:
[0075] Based on the multi-level progressive adjustment mechanism configuration, the spatial utilization rate of the current interface is calculated. The ratio of the occupied area to the total display area is calculated to obtain interface space utilization data.
[0076] According to the interface space utilization data and the importance score of the interface element, a visual balance constraint condition is set. A constraint function of element density distribution uniformity and visual barycenter stability is established to obtain a spatial layout constraint parameter.
[0077] The spatial layout constraint parameter is input into a dynamic reconstruction algorithm for element position redistribution. The visual focus area is allocated to high importance score elements, and the edge area is allocated to low importance score elements to obtain an element position reconstruction scheme.
[0078] Based on the element position reconstruction scheme, a user visual scanning path prediction model is constructed. The shortest visual distance and scanning time from the current focus point to the target element are calculated to obtain an adaptive interface space allocation scheme.
[0079] Specifically, the calculation of the interface space utilization data is based on the statistics of the occupied area of each element in the multi-level progressive adjustment mechanism configuration. The occupied area of each interface element is calculated geometrically based on its current size and position coordinates. The size data is derived from the width and height attributes of the element, and the position coordinates include the horizontal and vertical coordinate values of the top-left corner of the element. The area calculation uses the rectangular area formula, i.e., width multiplied by height. The total occupied area is obtained by adding up the occupied areas of all visible elements. The total display area is determined by the screen resolution of the IPTV interface. For a common resolution of 1920x1080, the total area is 2073600 pixel units. The space utilization is calculated by dividing the occupied area by the total display area to obtain a ratio data, which reflects the degree of congestion and the space usage efficiency of the current interface.
[0080] The setting process of the space layout constraint parameters establishes multiple constraint conditions in combination with the interface space utilization data and the element importance score. The element density distribution uniformity constraint aims to prevent the interface elements from being excessively concentrated in a certain area, causing visual congestion. This constraint is achieved by dividing the interface into multiple grid areas and counting the number of elements in each area. The ideal density distribution requires that the number of elements in each area differ by no more than a pre-set threshold. The threshold is set based on the average value calculated by dividing the total number of elements by the number of grids. The visual barycenter stability constraint ensures the visual balance of the interface. The visual barycenter is calculated by a weighted average method. The position coordinates of each element are multiplied by its importance score and then added up. Then, the sum is divided by the total of all element importance scores to obtain the visual barycenter coordinates of the interface. The stability constraint requires that the visual barycenter position be close to the geometric center of the interface. The deviation is quantified by calculating the Euclidean distance between the barycenter coordinates and the center coordinates. The smaller the distance, the better the visual balance. The constraint function combines the density uniformity index and the barycenter stability index by weighting to form a comprehensive evaluation function.
[0081] The execution process of the dynamic reconstruction algorithm is based on the re-allocation of element positions based on spatial layout constraint parameters. The algorithm uses a heuristic search method to find the optimal layout scheme that meets the constraint conditions. The search process first divides the interface into three levels: the visual focus area, the secondary focus area, and the edge area. The visual focus area is located near the golden section point of the interface and covers one-third of the total area of the interface. The secondary focus area is distributed around the visual focus area and covers half of the total area of the interface. The edge area is located at the edge of the interface. High-importance-score elements are preferentially allocated to the visual focus area. The allocation process is performed in order of importance from high to low. The allocation position of each element is determined by searching for the best free position in the target area. The best position is judged based on multiple factors, including the distance from the center of the region, the degree of overlap with allocated elements, and the degree of compliance with constraint conditions. Low-importance-score elements are allocated to the edge area with similar but lower priority allocation strategies. The reconstruction process also considers the logical association between elements, and elements of related functions are allocated to adjacent positions to facilitate user operation.
[0082] The construction of the user visual scanning path prediction model is based on eye tracking research and cognitive psychology principles. The model models the user's visual scanning behavior as a path planning problem from the current focus point to the target element. The scanning path usually follows the reading habit of from left to right and from top to bottom, and is attracted by elements with high contrast, large size, and dynamic changes. The model uses an attention map to represent the visual attraction of each position on the interface. The attention value is calculated by weighting the visual properties and importance score of the element. The shortest visual distance is calculated using the Manhattan distance formula, which is the absolute value of the horizontal coordinate difference plus the absolute value of the vertical coordinate difference. This distance measurement method conforms to the actual path characteristics of human eye visual scanning. The scanning time is calculated by dividing the visual distance by the average scanning speed. The average scanning speed is determined based on statistical analysis of eye movement data of user groups. The model also considers the influence of visual obstacles on the scanning path. When the target element is blocked by other elements, it will increase the additional scanning time and path complexity.
[0083] In a specific embodiment, step S5 comprises:
[0084] Based on the adaptive interface space allocation scheme, the user operation pause time and error operation frequency are monitored, the attention concentration index is calculated, and the user attention dispersion state data is obtained.
[0085] The user attention dispersion state data is compared with the preset attention concentration threshold value to determine whether to trigger the intensive guidance branch or the lightweight prompt branch. If the attention concentration is lower than the threshold value, the intensive guidance branch is triggered. Otherwise, the lightweight prompt branch is triggered. The guidance mode selection result is obtained.
[0086] According to the guidance mode selection result, a visual guidance intensity parameter is set, the reinforcement guidance mode adopts highlight flashing and arrow indication, the lightweight prompt mode adopts color change and transparency adjustment, and a personalized guidance strategy configuration is obtained;
[0087] The personalized guidance strategy configuration is applied to the final display state of the interface element, the visual properties and the interactive response mechanism of the element are synchronously adjusted, and a personalized differentiated display result is obtained.
[0088] Specifically, the acquisition of the user attention dispersion state data is based on the user behavior monitoring after the adaptive interface space allocation scheme is implemented, the operation pause time is measured by recording the time interval between the completion of the interface reconstruction and the execution of the next operation, the starting point of the measurement of the time interval is the time when the interface element position adjustment animation ends, and the endpoint is the time when the user mouse click, key input or voice instruction is detected, and the measurement accuracy reaches the millisecond level to ensure data accuracy. The error operation frequency statistics include the number of times of non-expected behaviors such as user click on invalid area, repeated click on the same element, selection of wrong menu item, etc., and the statistical time window is set to the first 30 seconds after the interface reconstruction. In this time period, the user usually completes the adaptation to the new layout and the target operation, the identification of the error operation is based on the comparison and judgment of the pre-defined correct operation path, and the behavior deviating from the expected path is recorded as the error operation. The calculation of the attention concentration degree index adopts the reverse correlation method, the longer the pause time, the more dispersed the attention, the more frequent the error operation, the heavier the cognitive load, the pause time is inversely proportional converted in the calculation process, that is, the actual pause time is divided by the pre-set standard time, the error operation frequency is also inversely proportional processed, and then the two conversion results are weighted and averaged to obtain the final attention concentration degree value. The weight distribution is that the pause time accounts for 0.6 and the error operation frequency accounts for 0.4.
[0089] The determination process of the guide mode selection result is based on the comparison between the attention concentration index and the preset threshold value. The setting of the attention concentration threshold value is based on the statistical analysis of the historical behavior data of the user group of the IPTV education platform. The median is calculated as the benchmark threshold value through the attention performance data of a large sample of users. Generally, it is set to 0.6. When the real-time attention concentration of the user is lower than the threshold value, it indicates that the user is in a state of attention dispersion, and stronger visual guidance is needed to help him quickly locate the target element. At this time, the enhanced guidance branch is triggered. The activation of the enhanced guidance branch is recorded as the "Enhanced" state identifier in the guide mode selection result. When the attention concentration is higher than or equal to the threshold value, it indicates that the user's attention is relatively concentrated. Strong visual guidance may cause interference. At this time, the light prompt branch is triggered, and the "Subtle" state identifier is recorded in the selection result. The judgment process uses a hard threshold method to avoid the uncertainty of the guidance strategy caused by the fuzzy state. The selection result records the attention concentration value and the threshold value data at the time of judgment, which facilitates the subsequent strategy optimization.
[0090] The setting process of the personalized guidance strategy configuration determines the specific visual guidance parameters according to the guide mode selection result. The visual guidance intensity parameters of the enhanced guidance mode include the frequency of highlight flashing, brightness gain, flashing duration, etc. The frequency of highlight flashing is set to 2 times per second to ensure that the user's attention is attracted without causing visual fatigue. The brightness gain is set to 150% of the original brightness to highlight the target element. The flashing duration is set to 3 seconds to give the user sufficient reaction time. The arrow indication function is realized by generating a dynamic arrow pattern near the target element. The color of the arrow is high-contrast red or orange. The size is adjusted adaptively according to the element size. The animation effect adopts a gradual pointing method from far to near. The visual guidance intensity parameters of the light prompt mode are relatively mild. Color change is realized by adjusting the border color or background color of the target element. The color selection avoids too bright colors. Usually, blue or green and other relatively soft tones are used. The transparency adjustment adjusts the transparency of the target element from the current value to complete opacity. The adjustment process adopts a gradual method to avoid abrupt changes. The adjustment time is set to 1 second to ensure that the user can perceive the change but will not be disturbed too much. The strategy configuration data is stored in the form of parameter groups, containing complete information such as guide type, intensity level, duration, color value, and transparency value.
[0091] The application process of the final display state of the interface element fuses and updates the personalized guidance strategy configuration with the existing visual properties of the element, the synchronous adjustment of the visual properties including the coordinated changes of multiple dimensions of the element such as color, transparency, size, position, animation effect, and the like, the adjustment process prioritizing the saliency of the guidance effect while taking into account the overall aesthetic coordination of the interface, when multiple elements need to be guided at the same time, a priority ranking mechanism is adopted, the priority being calculated based on the importance score of the element and the matching degree with the current voice intent, the high-priority element obtaining stronger guidance effect, and the guidance strength of the low-priority element being correspondingly reduced to avoid visual confusion. The synchronous update of the interactive response mechanism includes the adjustment of the interactive behaviors of the element such as click response speed, hovering effect, selected state, and the like, the element being guided having its click response priority temporarily increased, the response time being shortened to half of the normal value, the hovering effect being enhanced to display more obvious visual feedback, and the visual identifier of the selected state being more prominent, these adjustments ensuring that the operation experience of the user under the guidance prompt is more smooth and natural.
[0092] The IPTV interface element differential display method in the embodiments of the present application is described above, and the IPTV interface element differential display system in the embodiments of the present application is described below, please refer to Figure 2 An embodiment of the IPTV interface element differential display system in the embodiments of the present application includes:
[0093] The collection module is configured to collect audio features and semantic components of a user voice instruction, establish a time sequence association model of a voice intent category and a subsequent interface operation sequence, and train a voice interaction prediction dataset based on historical interaction data;
[0094] The encoding module is configured to weight and encode the voice interaction prediction dataset according to an education scene and a user cognitive load, and construct an intent-element probability matrix;
[0095] The synchronization module is configured to calculate real-time importance scores of interface elements based on the intent-element probability matrix, and synchronously control element visibility, position weight, and interactive response priority by using a multi-level progressive adjustment mechanism;
[0096] The generation module is configured to dynamically reconstruct element layout and predict a user visual scanning path according to interface space utilization and visual balance constraint conditions, and generate an adaptive interface space allocation scheme;
[0097] The output module is configured to detect a user attention dispersion state, start a strengthened guidance mode when the attention concentration degree is lower than a preset threshold, otherwise, adopt a lightweight prompt strategy, and output a personalized differential display result.
[0098] The above Figure 2The IPTV interface element differential display system in the embodiment of the application is described in detail from the perspective of a modular functional entity, and the IPTV interface element differential display device in the embodiment of the application is described in detail from the perspective of hardware processing.
[0099] With reference to Figure 3 The embodiment of the application also provides an IPTV interface element differential display device, which can be a server, and the internal structure of the IPTV interface element differential display device can be as shown in Figure 3 The IPTV interface element differential display device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. The processor of the computer is used to provide computing and control capabilities. The memory of the IPTV interface element differential display device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the IPTV interface element differential display device is used to store corresponding data in the embodiment. The network interface of the IPTV interface element differential display device is used to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the method.
[0100] Those skilled in the art can understand that Figure 3 The structure shown in the embodiment is only a block diagram of part of the structure related to the scheme of the application, and does not constitute a limitation on the IPTV interface element differential display device to which the scheme of the application is applied.
[0101] The application also provides a computer readable storage medium, which can be a non-volatile computer readable storage medium or a volatile computer readable storage medium. The computer readable storage medium stores instructions, and when the instructions are run on a computer, the computer is caused to perform the steps of the IPTV interface element differential display method.
[0102] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, the system and the unit described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein.
[0103] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for making an IPTV interface element differentiation display device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0104] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for differentiating display of IPTV interface elements, characterized in that, The method comprises: Step S1: Collecting audio features and semantic components of user voice instructions, establishing a time sequence association model of voice intent categories and subsequent interface operation sequences, and training a voice interaction prediction dataset based on historical interaction data; Step S2: Weighted coding the voice interaction prediction dataset according to education scenarios and user cognitive load, and constructing an intent-element probability matrix; Step S3: Calculating the real-time importance score of interface elements based on the intent-element probability matrix, and synchronously controlling the visibility, position weight and interaction response priority of elements using a multi-level progressive adjustment mechanism; Step S4: Dynamically reconstructing element layout and predicting user visual scanning path according to interface space utilization and visual balance constraint conditions, and generating an adaptive interface space allocation scheme; Step S5: Detecting the state of user attention dispersion, starting the reinforcement guidance mode when the attention concentration degree is lower than the preset threshold, otherwise adopting the lightweight prompting strategy, and outputting the personalized differentiated display result.
2. The method of claim 1, wherein the IPTV interface element differentiation presentation method is characterized by, The step S1 comprises: Extracting audio features of user voice instructions through the Xunfei voice assistant interface to obtain audio feature data containing frequency spectrum, pitch and speech rate; Performing semantic component analysis on the audio feature data based on a natural language processing algorithm to obtain semantic component structures of keywords, action words and target objects; Classifying the semantic component structures according to navigation, operation, query and setting to obtain voice intent category labels with time stamps and confidence levels; Collecting user interface operation sequences after the voice intent category labels occur, recording click, slide and dwell time interaction behaviors, and obtaining intent-operation behavior correspondence data; Constructing a time sequence association model based on the intent-operation behavior correspondence data, calculating the transition probability matrix of different intent categories and subsequent operation sequences, and obtaining a voice intent time sequence association model; Inputting historical user interaction records into the voice intent time sequence association model for training and optimization, updating transition probability parameters and association weight coefficients, and obtaining a voice interaction prediction dataset.
3. The method of claim 2, wherein the IPTV interface element differentiation presentation method is characterized by, The step of constructing a time sequence association model based on the intent-operation behavior correspondence data, calculating the transition probability matrix of different intent categories and subsequent operation sequences, and obtaining a voice intent time sequence association model comprises: Sorting the intent-operation behavior correspondence data in time sequence, establishing a directed graph structure of intent state nodes and operation behavior state nodes, and obtaining a time sequence state transition graph; Statistically analyzing the frequency of transition of each intent category to different operation sequences based on the time sequence state transition graph, calculating a state transition frequency matrix, and obtaining original transition statistical data; Normalizing the original transition statistical data, calculating the transition frequency proportion by row, and obtaining the operation sequence transition probability distribution corresponding to each intent category; Constructing a Markov chain model structure according to the transition probability distribution, setting the intent state as the starting node and the operation sequence state as the target node, and obtaining a time sequence association model framework; Smoothing the transition probability parameters in the time sequence association model framework, avoiding zero probability problem by using Laplace smoothing algorithm, and obtaining a voice intent time sequence association model.
4. The method of claim 1, wherein the IPTV interface element differentiation presentation method is characterized by, The step S2 comprises: The scene annotation data of the voice interaction prediction data set is annotated according to the education scene, classified into new tutorial, function browsing, content searching and module switching, and scene annotation data is obtained. The cognitive load index is calculated based on the user operation pause time and error rate, the cognitive load is quantified into three levels of low, medium and high, and user cognitive load level data is obtained. The scene annotation data and user cognitive load level data are weighted and encoded, the education scene weight coefficient and cognitive load weight coefficient are set respectively, and a double weighted encoding vector is obtained. According to the double weighted encoding vector, the influence probability of each voice intention on the interface element is calculated, a two-dimensional matrix structure of the behavior intention column and the interface element column is constructed, and an intention-element probability matrix is obtained.
5. The method of claim 1, wherein the IPTV interface element differentiation presentation method is characterized by, The step S3 comprises: Based on the intention-element probability matrix, the probability value corresponding to each interface element is extracted, and real-time matching calculation is performed combined with the current user voice intention category, and an interface element real-time importance score is obtained. According to the interface element real-time importance score, a visibility level threshold is set, the interface elements are divided into five visibility levels of completely visible, highlighted, normally displayed, faded display and hidden, and an element visibility grading result is obtained. The progressive adjustment algorithm is applied to the element visibility grading result, the transparency change step and the size scaling ratio are set, the visual properties of the interface elements are controlled to smoothly transition, and an element visibility adjustment parameter is obtained. The element visibility adjustment parameter is updated synchronously with the position weight and the interaction response priority of the interface element, a three-dimensional adjustment control vector is established, and a multi-level progressive adjustment mechanism configuration is obtained.
6. The method of claim 1, wherein the IPTV interface element differentiation presentation method is characterized by, The step S4 comprises: Based on the multi-level progressive adjustment mechanism configuration, the space utilization rate of the current interface is calculated, the ratio of the occupied area to the total display area is counted, and interface space utilization data is obtained. According to the interface space utilization data and the importance score of the interface element, a visual balance constraint condition is set, a constraint function of element density distribution uniformity and visual barycenter stability is established, and a space layout constraint parameter is obtained. The space layout constraint parameter is input into the dynamic reconstruction algorithm to redistribute the element position, the visual focus area is allocated to the high importance score element, and the edge area is allocated to the low importance score element, and an element position reconstruction scheme is obtained. Based on the element position reconstruction scheme, a user visual scanning path prediction model is constructed, the shortest visual distance and scanning time from the current focus point to the target element are calculated, and an adaptive interface space allocation scheme is obtained.
7. The method of claim 1, wherein the IPTV interface element differentiation presentation method is characterized by, The step S5 comprises: Based on the adaptive interface space allocation scheme, the user operation pause time and the error operation frequency are monitored, the attention concentration index is calculated, and user attention dispersion state data is obtained. The user attention dispersion state data is compared with the preset attention concentration threshold, when the attention concentration is lower than the threshold, the intensive guidance branch is triggered, otherwise the light prompt branch is triggered, and a guidance mode selection result is obtained. According to the guidance mode selection result, a visual guidance intensity parameter is set, the intensive guidance mode adopts highlight flashing and arrow indication, the lightweight prompt mode adopts color change and transparency adjustment, and a personalized guidance strategy configuration is obtained; The personalized guidance strategy configuration is applied to the final display state of the interface element, the visual properties and interactive response mechanism of the element are synchronously adjusted, and a personalized differentiated display result is obtained.
8. An IPTV interface element differential presentation system, characterized in that, The IPTV interface element differentiated display system for implementing the IPTV interface element differentiated display method as claimed in any one of claims 1-7 comprises: The acquisition module is configured to acquire audio features and semantic components of a user voice instruction, establish a time sequence association model of a voice intention category and a subsequent interface operation sequence, and train a voice interaction prediction dataset based on historical interaction data; The encoding module is configured to weight and encode the voice interaction prediction dataset according to education scenarios and user cognitive load, and construct an intention-element probability matrix; The synchronization module is configured to calculate real-time importance scores of interface elements based on the intention-element probability matrix, and synchronously control element visibility, position weight, and interactive response priority by using a multi-level progressive adjustment mechanism; The generation module is configured to dynamically reconstruct element layout and predict a user visual scanning path according to interface space utilization and visual balance constraint conditions, and generate an adaptive interface space allocation scheme; The output module is configured to detect a user attention dispersion state, start an intensive guidance mode when the attention concentration degree is lower than a preset threshold, otherwise, adopt a lightweight prompt strategy, and output a personalized differentiated display result.
9. A device for differentiating IPTV interface elements, characterized in that, The computer program is stored in the memory and executable on the processor, and the processor implements the IPTV interface element differentiated display method as claimed in any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is stored in the memory and executable on the processor, and the processor implements the IPTV interface element differentiated display method as claimed in any one of claims 1-7 when executing the computer program.
Citation Information
Patent Citations
Multi-mode interaction control method and system for smart blackboard
CN118655979A
Multi-modal user intention understanding and personalized shopping guide generation method and system
CN120106942A