IPTV education content intelligent navigation method and system based on voice recognition

By constructing an educational knowledge graph neural network and user cognitive profiles, the problem of IPTV education platforms being unable to understand the knowledge level of educational content and user cognition has been solved, enabling personalized educational content recommendations and improving the accuracy of recommendations and the coherence of learning paths.

CN120935409AActive Publication Date: 2025-11-11CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD +1

Patent Information

Application Number
CN202511476994.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-11-11
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing IPTV education platforms are unable to understand the knowledge levels and logical relationships of educational content, resulting in a lack of systematic and personalized content recommendations. Furthermore, they cannot accurately assess users' cognitive levels and learning abilities, leading to a mismatch between content difficulty and user capabilities.

Method used

By constructing a three-level semantic tree of knowledge points, courses, and subjects through an educational knowledge graph neural network, and combining it with user cognitive profiles and voice data parsing, query representation vectors are generated. Semantic matching and cognitive level adaptation are then performed to assess the coherence of the learning context and achieve personalized content recommendations.

Benefits of technology

It improves the accuracy of educational content recommendations and the coherence of learning paths, ensuring that recommended content matches the user's cognitive level and reducing the fragmentation of learning content and cognitive load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935409A_ABST
    Figure CN120935409A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses an IPTV education content intelligent navigation method and system based on voice recognition. The method comprises the following steps: constructing a knowledge point-course-subject three-level semantic tree through an educational knowledge graph neural network; performing cognitive modeling on the learning behaviors of the user to generate a user cognitive file; semantic analysis is performed on the voice data, and a query vector is generated in combination with a historical track; matching the query vector with the semantic tree, and screening candidate contents in combination with the cognitive archive; and performing continuity evaluation based on the session state to generate a navigation result. According to the method and the device, the problem that the recommended content is lack of systematicness and individuation due to the fact that an existing IPTV education platform cannot understand the education content knowledge hierarchical structure and the user cognitive development law is solved. And the education content recommendation accuracy and the learning path continuity are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to an intelligent navigation method and system for IPTV educational content based on speech recognition. Background Technology

[0002] Existing IPTV education platforms primarily employ content recommendation technologies based on keyword matching and collaborative filtering. Users search for educational content via remote control input or simple voice commands, and the system makes recommendations based on content tags and the user's viewing history. These technologies are effective in handling basic user search needs and can provide relevant educational content recommendations based on users' explicit behavioral data, and have been widely used in commercial online education platforms.

[0003] However, existing technologies have significant shortcomings: First, traditional keyword matching cannot understand the knowledge levels and logical relationships of educational content, resulting in a lack of systematicity and coherence in recommended content; second, existing user profile construction is mainly based on viewing history and click behavior, which cannot accurately assess users' cognitive level and learning ability, and is prone to the problem of mismatch between content difficulty and user ability; third, traditional speech recognition technology lacks a deep understanding of the semantics of the education field and cannot handle complex educational query intentions and contextual relationships. Summary of the Invention

[0004] This application provides a speech recognition-based intelligent navigation method and system for IPTV educational content, addressing the problem that existing IPTV educational platforms cannot understand the hierarchical structure of educational content and the cognitive development patterns of users, resulting in a lack of systematic and personalized content recommendations. This improves the accuracy of educational content recommendations and the coherence of learning paths.

[0005] Firstly, this application provides a method for intelligent navigation of IPTV educational content based on speech recognition, the method comprising: Step S1: Use an educational knowledge graph neural network to perform hierarchical structure analysis on IPTV educational content and construct a three-level semantic tree of knowledge points, courses, and subjects; Step S2: Model the cognitive level of users' learning behavior on the IPTV platform and generate user cognitive profiles; Step S3: Perform semantic hierarchical perceptual analysis on the voice data collected by the IPTV remote control, construct the educational conversation state by combining the historical learning trajectory, and generate a query representation vector through the educational knowledge graph neural network; Step S4: Perform semantic matching calculation between the query representation vector and the three-level semantic tree of knowledge point-course-subject, combine it with the user cognitive profile to adapt the cognitive level, and filter out candidate educational content; Step S5: Based on the educational session state, evaluate the cross-round learning context coherence of the candidate educational content and generate intelligent navigation results for IPTV educational content.

[0006] Secondly, this application provides an intelligent navigation system for IPTV educational content based on speech recognition, the intelligent navigation system for IPTV educational content based on speech recognition comprising: The parsing module is used to perform hierarchical structure parsing of IPTV educational content through an educational knowledge graph neural network, and to construct a three-level semantic tree of knowledge points, courses, and subjects. The modeling module is used to model the cognitive level of users' learning behavior on the IPTV platform and generate user cognitive profiles. The generation module is used to perform semantic-level perceptual analysis on the voice data collected by the IPTV remote control, construct the educational conversation state by combining the historical learning trajectory, and generate query representation vectors through the educational knowledge graph neural network. The matching module is used to perform semantic matching calculations between the query representation vector and the three-level semantic tree of knowledge points-courses-subjects, and combine it with the user's cognitive profile to adapt to the cognitive level and filter out candidate educational content. The evaluation module is used to evaluate the learning context coherence of candidate educational content across different rounds based on the educational session status, and generate intelligent navigation results for IPTV educational content.

[0007] Thirdly, a speech recognition-based intelligent navigation device for IPTV educational content is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the speech recognition-based intelligent navigation device for IPTV educational content to execute the aforementioned speech recognition-based intelligent navigation method for IPTV educational content.

[0008] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the above-described intelligent navigation method for IPTV educational content based on speech recognition.

[0009] The technical solution provided in this application constructs a three-level semantic tree of knowledge points, courses, and subjects using an educational knowledge graph neural network. This fundamentally solves the problem that existing technologies cannot understand the hierarchical structure of educational content, enabling IPTV platforms to accurately grasp the logical relationships and dependencies between different educational content, avoiding the fragmented content recommendation problem caused by traditional keyword matching. The generation of user cognitive profiles, through in-depth analysis of users' learning behavior patterns on the IPTV platform, establishes a precise cognitive level assessment system. Compared to the crude user profiles based solely on viewing history in existing technologies, this system can more accurately identify users' true learning abilities and cognitive development stages. The innovative technology of constructing educational conversation states by combining semantic level perception and parsing with historical learning trajectories enables the system to understand the deep educational intent behind user voice queries and organically integrate current needs with historical learning backgrounds, significantly improving the intelligence level of voice interaction. The semantic matching calculation of query representation vectors and the three-level semantic tree, combined with the cognitive level adaptation mechanism of user cognitive profiles, achieves dual-precision positioning for content recommendation, ensuring both semantic relevance and cognitive suitability.

[0010] The message passing mechanism of graph neural networks enables the structured representation and relational modeling of educational knowledge, allowing the system to simulate the knowledge organization thinking of human teachers and providing a scientific basis for personalized learning path planning. The application value of the cross-round learning context coherence evaluation algorithm lies in its ability to maintain the continuity of knowledge and cognitive consistency among users across multiple learning sessions, avoiding the problems of learning content skipping and cognitive overload common in traditional recommendation systems. It is particularly suitable for educational scenarios requiring systematic knowledge construction. The unique contribution of the semantic hierarchical perception parsing algorithm in educational voice interaction lies in its ability to recognize language expression patterns and conceptual levels specific to the educational field. Compared to general speech recognition technology, it has a significant accuracy advantage in handling complex educational queries, effectively reducing user interaction costs and improving the fluency of the learning experience. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of an embodiment of the intelligent navigation method for IPTV educational content based on speech recognition in this application. Figure 2 This is a schematic diagram illustrating the changes in the duration of a user's viewing of educational content on an IPTV platform within one week, as described in this application embodiment. Figure 3 This is a schematic diagram of one embodiment of the intelligent navigation system for IPTV educational content based on speech recognition in this application. Figure 4 This is a schematic block diagram of the structure of an IPTV educational content intelligent navigation device based on voice recognition in an embodiment of the present invention. Detailed Implementation

[0013] This application provides a method and system for intelligent navigation of IPTV educational content based on speech recognition. The terms first, second, third, fourth, etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms include or have, and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0014] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the intelligent navigation method for IPTV educational content based on speech recognition in this application includes: Step S1: Use an educational knowledge graph neural network to perform hierarchical structure analysis on IPTV educational content and construct a three-level semantic tree of knowledge points, courses, and subjects; Step S2: Model the cognitive level of users' learning behavior on the IPTV platform and generate user cognitive profiles; Step S3: Perform semantic hierarchical perception analysis on the voice data collected by the IPTV remote control, construct the educational conversation state by combining the historical learning trajectory, and generate query representation vectors through the educational knowledge graph neural network; Step S4: Perform semantic matching calculations between the query representation vector and the three-level semantic tree of knowledge points-courses-subjects, combine it with the user's cognitive profile to adapt to the cognitive level, and filter out candidate educational content; Step S5: Based on the educational session status, evaluate the cross-round learning context coherence of candidate educational content and generate intelligent navigation results for IPTV educational content.

[0015] It is understood that the executing entity of this application can be an IPTV educational content intelligent navigation system based on voice recognition, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiment uses a server as an example for illustration.

[0016] Specifically, this paper constructs a hierarchical semantic structure for IPTV educational content using an educational knowledge graph neural network, addressing the problem that existing voice navigation systems cannot understand the knowledge dependencies within educational content. The educational knowledge graph neural network is a graph neural network architecture specifically designed to handle knowledge relationships in the educational domain, comprising three core components: a graph embedding layer, a graph attention mechanism, and a hierarchical aggregation function. The graph embedding layer converts the text descriptions and video tags of the IPTV educational content into numerical vectors. The graph attention mechanism calculates the semantic association strength between different knowledge points, and the hierarchical aggregation function organizes knowledge points into course and subject levels according to educational principles. Specifically, text preprocessing first extracts educational keywords such as functions, derivatives, and calculus from the educational videos. Then, the graph embedding layer encodes these keywords into 256-dimensional vector representations. The graph attention mechanism identifies the pre-learning relationship between functions and derivatives by calculating the attention weights between vectors. The hierarchical aggregation function aggregates related knowledge points such as functions, derivatives, and limits into basic calculus course nodes. Multiple related courses are further aggregated into mathematics subject nodes, forming a three-level semantic tree structure of knowledge point-course-subject.

[0017] User cognitive profiles are generated by analyzing learning behavior data on the IPTV platform to build personalized cognitive models. The raw learning behavior data includes four dimensions: viewing duration, pause frequency, replay count, and skipping behavior. Time-series analysis statistically analyzes changes in user behavior patterns within specific time windows. Attention level is calculated as the ratio of viewing duration to total content duration, and comprehension speed is determined based on the inverse relationship between replay count and pause frequency. Cognitive level assessment correlates learning behavior feature vectors with difficulty labels in a three-level semantic tree, calculating user scores for mastery in different knowledge domains such as mathematics, physics, and chemistry, forming a cognitive ability assessment matrix. The user cognitive profile, as a data structure, stores information on the user's knowledge base level, learning preference type, and cognitive development stage. The knowledge base level records the set of knowledge points the user has mastered, the learning preference type identifies the content format the user prefers, and the cognitive development stage reflects the user's current learning ability level.

[0018] Semantic hierarchical perceptual analysis processes voice data collected from IPTV remote controls to extract semantic information specific to the education field. The voice data first undergoes noise reduction and endpoint detection preprocessing, followed by semantic analysis to identify educational keywords and learning intentions. Educational keywords include subject names, knowledge point terms, and learning action vocabulary. Learning intentions are categorized into types such as course search, knowledge query, and difficulty adjustment. Historical learning trajectory data records the user's past learning topic sequences and timestamp information. Time-series learning behavior sequence analysis identifies the user's learning patterns and progress evolution trends. Current voice query features are combined with historical learning trajectory vectors through a feature fusion algorithm to generate an educational conversation state that includes immediate needs and historical context. The message passing mechanism of the educational knowledge graph neural network propagates semantic information across the graph structure. Node features are updated through weighted aggregation of neighboring nodes, outputting a query representation vector that integrates knowledge relevance.

[0019] Semantic matching calculation uses the cosine similarity algorithm to measure the relevance between the query representation vector and each node of the three-level semantic tree. Cosine similarity assesses semantic similarity by calculating the cosine of the angle between two vectors, with values ​​ranging from -1 to +1; the closer to +1, the stronger the semantic similarity. The semantic matching score matrix records the similarity values ​​between the query vector and all educational content nodes. A preset threshold filters out low-relevance content, forming an initial list of educational content. Cognitive level adaptation calculates the matching degree between the content's difficulty level and the cognitive ability assessment matrix in the user's cognitive profile. This assesses the degree to which the content's cognitive requirements match the user's cognitive level. The cognitive adaptation score reflects the content's suitability for the user; content exceeding the cognitive level threshold is retained as candidate educational content.

[0020] The cross-round learning context coherence assessment is based on the rationality of the learning sequence of candidate content by analyzing the state of the educational session. The learning coherence score is obtained by calculating the semantic relevance between candidate content and historical learning topics; content with strong relevance receives a higher coherence score. The knowledge structure coherence evaluation judges the logical order between content based on the knowledge dependencies in the three-level semantic tree; content with sufficient prior knowledge receives a better structure evaluation. The cross-round learning pattern vector records the user's knowledge acquisition path and progress changes in multiple learning sessions; the learning progress facilitation is quantified by analyzing the contribution of candidate content to the user's long-term learning goals. The context consistency scoring mechanism comprehensively considers the position and role of content in the overall learning context; the evaluation results guide content sorting and selection, generating personalized IPTV educational content intelligent navigation results that conform to the user's cognitive development patterns.

[0021] In one specific embodiment, step S1 includes: The text descriptions and video tags of IPTV educational content are preprocessed to extract keywords and knowledge concept identifiers in the education field, thus obtaining the original feature data of the educational content. The original feature data of educational content is input into the graph embedding layer of the educational knowledge graph neural network for vector encoding to obtain the embedding vector of educational content nodes; Based on the embedding vectors of educational content nodes, the semantic association weights between knowledge points are calculated through the graph attention mechanism of the educational knowledge graph neural network, a knowledge point layer relationship graph is constructed, and a knowledge point layer node set is obtained. Based on the knowledge point layer node set, the course level clustering is performed through the hierarchical aggregation module of the educational knowledge graph neural network. Related knowledge points are organized into course nodes, resulting in a course layer node set and a subject layer node set, which are combined to form a three-level semantic tree of knowledge point-course-subject.

[0022] Specifically, IPTV educational content preprocessing extracts structured information from the metadata of educational videos, addressing the problem of existing technologies being unable to understand the semantic level of educational content. Preprocessing first cleanses the titles, descriptions, and tags of the educational videos, removing punctuation, stop words, and irrelevant characters. Then, it identifies key concepts through an educational domain dictionary. Educational domain keywords include subject-specific terms such as functions, derivatives, and chemical bonds, while knowledge concept identifiers include cognitive level markers such as basic concepts, application examples, and comprehensive exercises. Text segmentation divides the descriptive content into independent lexical units, part-of-speech tagging identifies the grammatical attributes of nouns, verbs, and adjectives, and named entity recognition extracts names of people, places, and technical terms. The original feature data of the educational content is stored in a structured form, containing numerical representations across three dimensions: word sequences, word frequency statistics, and semantic tags.

[0023] The graph embedding layer transforms the raw feature data of educational content into a high-dimensional vector representation, and uses the Word2Vec algorithm to map words to a continuous numerical space. The Word2Vec algorithm learns word vectors based on the contextual relationships of words in the text, and captures semantic similarity between words through neural network training. The graph embedding layer receives a preprocessed word sequence, with each word corresponding to a fixed-dimensional vector representation. Multiple word vectors are merged into a content-level embedding vector through average pooling or attention-weighted summation. The embedding vector of an educational content node has a fixed dimensional length, and each numerical component in the vector represents the feature strength of that content in a specific semantic dimension. The numerical distribution of the embedding vectors reflects the semantic features of the educational content; semantically similar content is closer in the vector space, while content with significant semantic differences is farther apart.

[0024] The graph attention mechanism calculates semantic association weights between knowledge points based on the embedding vectors of educational content nodes, and uses a self-attention algorithm to analyze the interdependencies between vectors. The self-attention algorithm calculates attention scores through linear transformations of query vectors, key vectors, and value vectors, and these scores are then normalized using softmax to obtain the weight distribution. The graph attention mechanism uses the embedding vector of each educational content as both a query vector and a key vector, calculating the attention weight between any two content items. The attention weight values ​​reflect the strength of the association between knowledge points; knowledge point pairs with higher weight values ​​have stronger semantic relevance or learning dependencies. The knowledge point layer relationship graph stores the attention weights in the form of an adjacency matrix, where each element represents the association strength of the corresponding knowledge point pair. The knowledge point layer node set contains all educational content nodes involved in relationship construction and their inter-node weight information.

[0025] Hierarchical aggregation organizes the knowledge point layer node set into course and subject levels through an educational knowledge graph neural network, and uses a spectral clustering algorithm to identify semantically similar knowledge point groups. The spectral clustering algorithm performs eigenvalue decomposition based on the Laplacian matrix of the graph, projecting the high-dimensional graph structure into a low-dimensional space to facilitate cluster analysis. The Laplacian matrix is ​​calculated by the difference between the degree matrix and the adjacency matrix, and the eigenvectors correspond to the main connected components in the graph structure. The clustering algorithm groups semantically related knowledge points into course nodes, each representing a set of knowledge points with inherent logical connections. The course layer node set contains all identified course nodes and a list of their contained knowledge points, while the subject layer node set is formed through further aggregation of course nodes. The three-level semantic tree of knowledge point-course-subject organizes the hierarchical relationship of educational content in a tree structure, with leaf nodes corresponding to specific knowledge points, intermediate nodes corresponding to course categories, and the root node corresponding to subject domains.

[0026] In one specific embodiment, step S2 includes: Data on users' viewing time, pause frequency, replay frequency, and skipping behavior on the IPTV platform are collected to obtain raw data on user learning behavior. By performing time-series analysis on raw user learning behavior data, we can calculate user focus and comprehension speed indicators for different educational content and obtain learning behavior feature vectors. Cognitive level assessment is performed based on learning behavior feature vectors combined with difficulty tags in a three-level semantic tree of knowledge points, courses, and subjects. The score of the user's mastery in each knowledge domain is calculated to obtain a cognitive ability assessment matrix. Based on the cognitive ability assessment matrix, a personal profile data structure is constructed that includes the user's knowledge base, learning preferences, and cognitive development stage, thus obtaining the user's cognitive profile.

[0027] Specifically, user learning behavior data collection records user interaction with educational content through the real-time monitoring function of the IPTV set-top box, solving the problem of inaccurate assessment of user cognitive levels in existing technologies. Viewing duration data records the continuous time period from the start to the end of playback; the IPTV set-top box records a timestamp every second, accumulating the total viewing time. Pause frequency data tracks the number of times the user actively pauses the video during viewing; each pause triggers a counter increment, and the pause frequency is calculated by dividing the total number of pauses by the total viewing time. Replay frequency data records the user's repeated viewing of the same educational content segment; when the user rewinds and replays the viewed content using the remote control, the replay counter automatically increments. Skip behavior data monitors user actions such as fast-forwarding or chapter skipping, recording the start time, end time, and duration of each skip. The raw user learning behavior data is stored in time-series format, with each record containing five fields: user identifier, content identifier, behavior type, timestamp, and numerical parameters.

[0028] like Figure 2 As shown, Figure 2 This chart shows the changes in user viewing time for educational content on the IPTV platform over a week. The horizontal axis represents dates (Monday to Sunday), and the vertical axis represents viewing time (in minutes). The line graph reflects the temporal changes in viewing time within the raw user learning behavior data, with convex dots representing times of high user learning activity and concave dots representing times of relatively low activity. This data provides the foundation for subsequent time-series analysis, used to calculate user focus and comprehension speed indicators for different educational content, thereby generating a learning behavior feature vector. The data fluctuation patterns shown in the chart reflect users' individualized learning habits and time management preferences, providing crucial behavioral characteristic data support for constructing user cognitive profiles.

[0029] Time-series analysis processes raw user learning behavior data, calculating focus and comprehension speed metrics using a sliding window algorithm. The sliding window algorithm sets a fixed time window length, sliding the window across the timeline to analyze changes in behavioral patterns over different time periods. The focus metric is calculated as the ratio of continuous viewing time to the total content viewing time. Continuous viewing time refers to the longest viewing period without pausing or skipping; a higher ratio indicates stronger user focus. The comprehension speed metric is determined based on the inverse relationship between replay frequency and pause frequency. Content with many replays or frequent pauses is considered more difficult to understand, and users understand it more slowly. The time-series analysis algorithm standardizes the data across four dimensions: viewing time, pause frequency, replay frequency, and skipping behavior, eliminating the influence of different units on the analysis results. Standardization uses the Z-score method, subtracting the mean from the raw values ​​and dividing by the standard deviation to obtain a standard normally distributed value. The learning behavior feature vector stores the standardized behavioral metrics as a multi-dimensional array, with each dimension of the vector corresponding to a quantitative representation of a learning behavior.

[0030] The cognitive level assessment correlates learning behavior feature vectors with difficulty labels in a three-level semantic tree of knowledge points, courses, and subjects, and uses a weighted average algorithm to calculate the user's mastery level in each knowledge domain. Difficulty labels are divided into three levels—beginner, intermediate, and advanced—based on the cognitive complexity of the educational content, each with a different numerical weight. The weighted average algorithm multiplies the user's learning performance at a specific difficulty level by the corresponding weight and then sums them to obtain a comprehensive mastery score. The calculation process for the mastery score includes multiplying the focus index by the difficulty weight, analyzing the match between the comprehension speed index and cognitive requirements, and assessing the correspondence between learning completion and content coverage. The cognitive ability assessment matrix is ​​organized in a two-dimensional table format, with rows corresponding to different knowledge domains and columns corresponding to different cognitive ability dimensions. Each element in the matrix stores the user's mastery score at the intersection of a specific knowledge domain and cognitive dimension. The rows of the matrix include subject categories such as mathematics, physics, chemistry, and language arts, while the columns include cognitive categories such as memory ability, comprehension ability, application ability, and analytical ability.

[0031] The user cognitive profile is constructed based on a cognitive ability assessment matrix to establish a multi-layered personal learning profile data structure. The knowledge base layer records the set of knowledge points the user has mastered. It identifies knowledge points whose mastery scores exceed a threshold by traversing the cognitive ability assessment matrix, and stores the identifiers of these knowledge points in the knowledge base set. The learning preference layer analyzes the differences in user learning performance across different content types, identifying the user's preferred learning methods and content formats. Learning preferences are determined by comparing the user's attention span and comprehension speed across video, audio, and text-based educational content; content types with higher attention span and comprehension speed are marked as the user's preferred types. The cognitive development stage layer assesses the user's current developmental level based on their comprehensive performance across different cognitive ability dimensions. Cognitive development stages include classifications from Piaget's theory of cognitive development, such as the concrete operational stage and the formal operational stage. The user cognitive profile is organized in a tree structure, with the root node storing basic user information, and child nodes corresponding to detailed data across the three dimensions: knowledge base, learning preferences, and cognitive development stage.

[0032] In one specific embodiment, step S3 includes: Semantic hierarchical perceptual analysis is performed on the voice data collected by the IPTV remote control to extract educational keywords and intent features from the voice, thereby obtaining the current voice query features; Construct a time-series learning behavior sequence based on users' historical learning trajectory data, analyze the changes in users' learning topics and progress in different time periods, and obtain historical learning trajectory vectors; The current voice query features are fused with historical learning trajectory vectors to construct session state data containing the current query intent and historical learning context, thus obtaining the educational session state. The educational session state is input into the educational knowledge graph neural network for semantic encoding. The message passing mechanism of the graph neural network generates a vector representation that integrates knowledge relevance, thus obtaining the query representation vector.

[0033] Specifically, the semantic-level perceptual parsing process handles the voice data collected by the IPTV remote control, extracting semantic information specific to the education field through multi-level semantic analysis, thus addressing the problem that existing speech recognition technologies cannot understand educational contexts. The voice data first undergoes acoustic feature extraction, converting the audio signal into a Mel-frequency cepstral coefficient sequence. Mel-frequency cepstral coefficients are numerical representations reflecting the spectral characteristics of speech, effectively capturing phoneme information. The acoustic model uses a recurrent neural network to process temporal speech features, mapping the Mel-frequency cepstral coefficient sequence to a phoneme probability distribution, which represents the likelihood of different phonemes corresponding to speech segments. The language model performs probability calculations based on an education-related vocabulary database, which includes professional terms and common expressions from disciplines such as mathematics, physics, and chemistry. The language model establishes probability distributions by statistically analyzing the frequency and combination patterns of these words in educational contexts. Educational keyword extraction uses a named entity recognition algorithm to locate subject-specific terms, knowledge point names, and learning action words from the recognized text. The named entity recognition algorithm uses a conditional random field model to label the semantic categories of words. Intent feature extraction is based on a predefined educational intent classification system, including types such as content search, difficulty adjustment, progress query, and review. The intent classifier uses a support vector machine algorithm to determine the user's learning objectives based on keyword combinations and grammatical structure. Current voice query features are stored in structured data format, containing four fields: recognized text, keyword list, intent category, and confidence score.

[0034] The historical learning trajectory vector is constructed based on the user's past learning behavior sequence, employing a temporal analysis algorithm to identify learning patterns and progress evolution trends. The user's historical learning trajectory data includes timestamps, learning content identifiers, learning duration, completion status, and other dimensional information. Timestamps record the occurrence time of learning activities, and learning content identifiers are associated with specific nodes in a three-level semantic tree of knowledge point-course-subject. The temporal learning behavior sequence arranges the user's learning activities in chronological order, with each element representing a detailed record of a learning session. Learning topic analysis uses clustering algorithms to identify the knowledge domains the user focuses on in different time periods, grouping similar learning activities into the same topic group based on the semantic similarity of the learning content. Progress change analysis assesses the changing trend of mastery by comparing the user's repeated learning behavior on the same knowledge points; a decrease in the number of repeated learning sessions indicates improved mastery, and a shortened learning duration indicates improved comprehension efficiency. The historical learning trajectory vector uses a bag-of-words model to encode the user's learning preferences and knowledge structure. The bag-of-words model uses the knowledge points the user has learned as vocabulary and the learning frequency as word frequency, generating a high-dimensional sparse vector representation. Each dimension of the vector corresponds to a node in the three-level semantic tree of knowledge point-course-subject. The dimension value reflects the user's familiarity with the knowledge point and the time invested in learning it.

[0035] Feature fusion processing combines the features of the current voice query with historical learning trajectory vectors to construct a comprehensive representation that includes immediate needs and historical context. The fusion algorithm employs an attention mechanism to calculate the relevance weights between the current query and historical learning content. This mechanism determines the strength of the association by calculating the semantic similarity between query keywords and historical learning knowledge points. Similarity calculation is based on the cosine distance of word vectors, which are obtained through a pre-trained educational domain language model. Higher similarity values ​​indicate a stronger association between the current query and historical learning content. The session state data structure comprises four components: current query information, relevant historical learning records, contextual association weights, and a time decay factor. The current query information stores the speech recognition results and intent classification. Relevant historical learning records contain past learning activities semantically related to the current query. Contextual association weights quantify the degree of association between the current query and historical learning content. The time decay factor adjusts its weight based on the temporal distance of historical learning activities, with learning activities closer to the current time receiving a higher weight. The educational session state merges multiple historical learning records into a unified contextual representation through a weighted summation. The weights in the weighted summation are determined by the product of the contextual association weights and the time decay factor.

[0036] The message passing mechanism of the educational knowledge graph neural network propagates semantic information based on the graph structure, encoding the educational session state into a vector representation that integrates knowledge relevance. The message passing mechanism is the core computational process of the graph neural network, updating node features through information exchange between nodes. Each node collects messages from its neighbors and fuses them with its own features. Message computation employs a combination of linear transformations and nonlinear activation functions. The linear transformation maps neighbor node features to the message space through matrix multiplication, while the nonlinear activation function introduces nonlinear expressive power to enhance the model's fitting ability. The aggregation function merges messages from multiple neighbor nodes into a single representation, using summation or average pooling to handle variable-length message sequences. The node update function combines the current node features with the aggregated neighbor messages to calculate new node features, using a gating mechanism to control the fusion ratio of historical and new information. After multiple rounds of message passing, the feature vector of the educational session state node in the educational knowledge graph neural network contains semantic information from relevant knowledge points, and the numerical distribution of the vector reflects the position and relevance of the current query within the entire knowledge system. The query representation vector, as the feature representation of the educational session state node, contains the semantic content of the current voice query, the user's historical learning background, and structured information about relevant knowledge points.

[0037] In one specific embodiment, step S4 includes: The semantic matching score matrix is ​​obtained by calculating the cosine similarity between the query representation vector and the vectors of each node in the three-level semantic tree of knowledge point-course-subject. Based on the semantic matching score matrix, educational content nodes with similarity higher than a preset threshold are filtered to construct a preliminary matching candidate set of content and obtain a preliminary list of educational content. The difficulty level of each item in the initial educational content list is matched with the cognitive ability assessment matrix in the user's cognitive profile to calculate the degree of fit between the content difficulty and the user's cognitive level, and a cognitive fit score is obtained. The initial list of educational content is sorted and filtered based on cognitive fit scores, and educational content with fit scores exceeding the cognitive level threshold is retained to obtain candidate educational content.

[0038] Specifically, cosine similarity calculation semantically matches the query representation vector with the node vectors in the three-level semantic tree of knowledge points, courses, and subjects, solving the problem of inaccurate matching between user voice queries and educational content in existing technologies. Cosine similarity is a mathematical method for measuring the cosine of the angle between two vectors in vector space. It is obtained by calculating the dot product of the vectors and dividing by the product of their magnitudes. The query representation vector serves as the target vector, containing the semantic features of the user's voice query and historical learning context information. Each node in the three-level semantic tree corresponds to a vector representation, and the node vector encodes the knowledge attributes and semantic features of the educational content. The similarity calculation process first calculates the dot product between the query vector and the node vector. The dot product is obtained by multiplying the corresponding dimension values ​​and then summing them. The dot product value reflects the degree of matching between the two vectors in each semantic dimension. The vector magnitude is calculated by taking the square root of the sum of the squares of the values ​​in each dimension. The magnitude represents the length of the vector in the high-dimensional space. The cosine similarity value ranges from -1 to +1. The closer the value is to +1, the higher the semantic similarity; the closer it is to -1, the greater the semantic difference; and close to zero, the semantics are unrelated. The semantic matching score matrix stores the calculation results in a two-dimensional table. The rows correspond to the query representation vector, the columns correspond to the nodes in the three-level semantic tree, and each element in the matrix stores the cosine similarity value between the corresponding query and the node.

[0039] The filtering mechanism identifies highly relevant educational content nodes based on a semantic matching score matrix and filters out low-similarity matching results using a preset threshold. The preset threshold is determined based on the content quality requirements and user experience standards of the IPTV education platform; a threshold that is too high will result in too few candidate contents, while a threshold that is too low will introduce a large amount of irrelevant content. The filtering algorithm iterates through all values ​​in the semantic matching score matrix, recording educational content nodes with similarity higher than the preset threshold in the candidate set. The initial matching candidate set contains all educational content nodes that have passed the similarity filtering; each element in the set corresponds to an educational content and its similarity score with the query. The initial selected educational content list sorts the candidate set from highest to lowest similarity score, reflecting the matching priority of different educational content with the user's query. Each entry in the list contains four fields: content identifier, content title, similarity score, and content attributes. Content attributes include metadata information such as subject classification, difficulty level, and duration.

[0040] Cognitive fit calculation matches the difficulty level of the initially selected educational content with the user's cognitive profile, assessing the degree of fit between the content's cognitive requirements and the user's cognitive abilities. Content difficulty levels are categorized into three levels: beginner, intermediate, and advanced, based on the cognitive complexity of the educational content. Each level corresponds to different cognitive ability requirements and prior knowledge conditions. The cognitive ability assessment matrix in the user's cognitive profile records the user's mastery level in different knowledge domains and cognitive dimensions; the matrix values ​​reflect the user's current cognitive development level and learning ability. Fit calculation determines the degree of fit by comparing the gap between the content difficulty level and the user's cognitive level. The degree of fit is represented by a numerical score; a high score indicates a good match between the content difficulty and the user's ability, while a low score indicates an unsuitable level. The calculation process first converts the content difficulty level into a numerical representation: beginner corresponds to value one, intermediate to value two, and advanced to value three. Then, it extracts the mastery score for the corresponding knowledge domain from the user's cognitive ability assessment matrix. The fit score is calculated by the absolute value of the difference between the content difficulty value and the user's mastery score; a smaller difference indicates a higher fit, and a larger difference indicates a lower fit. The cognitive fit score standardizes the fit score to eliminate the influence of different units on the comparison results. The standardized score is used for subsequent sorting and filtering operations.

[0041] The sorting and filtering process reorders the initial list of educational content based on cognitive suitability scores, prioritizing content with high suitability. A stable sorting algorithm ensures that content with the same suitability score maintains its original similarity ranking, guaranteeing that both semantic matching and cognitive suitability evaluation results are reasonably reflected. The cognitive level threshold is set based on the user's learning goals and cognitive development stage, taking into account both the user's current ability level and a reasonable level of challenge. The filtering process iterates through the sorted content list, retaining educational content with suitability scores exceeding the cognitive level threshold and removing content with insufficient suitability. The candidate educational content, as the filtering result, includes high-quality content recommendations that have undergone dual filtering based on semantic matching and cognitive suitability. The number of content items is controlled within a reasonable range based on the IPTV interface display capabilities and user convenience.

[0042] In one specific embodiment, step S5 includes: Based on the historical learning trajectory vector in the educational session state, the learning sequence coherence analysis of candidate educational content is performed, the correlation strength between each content and the user's historical learning topics is calculated, and the learning coherence score is obtained. The candidate educational content is sorted according to the knowledge dependency relationship in the three-level semantic tree of knowledge point-course-subject, and the pre-knowledge association and learning path rationality between the content are analyzed to obtain an evaluation of the coherence of the knowledge structure. Based on the learning coherence score and knowledge structure coherence evaluation, cross-round contextual consistency is calculated to assess the role of content recommendations in promoting the user's overall learning process, thus obtaining contextual coherence evaluation results. Based on the contextual coherence assessment results, candidate educational content is sorted and filtered to generate a personalized recommendation sequence that conforms to the user's learning path and cognitive development, resulting in intelligent navigation results for IPTV educational content.

[0043] Specifically, the learning sequence coherence analysis calculates the correlation strength between candidate educational content and the user's past learning topics based on the historical learning trajectory vector in the educational session state, addressing the problem of existing technologies failing to maintain the continuity of learning across learning rounds. The historical learning trajectory vector records the user's learning activity sequence in different time periods. Each dimension in the vector corresponds to a specific node in the three-level semantic tree of knowledge point-course-subject, and the dimension value reflects the user's learning engagement and mastery level on that knowledge point. The correlation strength calculation uses a weighted cosine similarity algorithm to calculate the similarity between the vector representation of the candidate educational content and the historical learning trajectory vector. The similarity value reflects the degree of matching between the content and the user's historical learning topics. The weighting mechanism assigns different weights based on the time distance of the learning activities; learning activities closer to the current time have a higher weight, while the weight of learning activities further away gradually decreases. The time decay function adopts an exponential decay model, and the decay coefficient is determined based on the forgetting curve characteristics of the educational content, ensuring that the learning coherence analysis considers both historical context and highlights recent learning priorities. The learning coherence score is obtained through the standardization process of weighted similarity calculation. Standardization eliminates the influence of different learning trajectory lengths on the comparison results. The higher the score, the stronger the continuity between the candidate content and the user's learning context.

[0044] The knowledge structure coherence evaluation analyzes the logical order of candidate educational content based on knowledge dependency relationships in a three-level semantic tree of knowledge points, courses, and subjects. Knowledge dependency relationships describe the prerequisite learning requirements between different knowledge points; prerequisite knowledge points must be mastered before subsequent knowledge points to form a knowledge structure. The dependency graph is represented as a directed graph, where nodes correspond to knowledge points, directed edges represent dependency directions, and edge weights reflect dependency strength. A topological sorting algorithm processes the knowledge dependency graph, generating a learning order of knowledge points that satisfies dependency constraints. Topological sorting ensures that prerequisite knowledge points appear in the sequence before knowledge points that depend on them. Candidate educational content determines its learning priority based on the position of its associated knowledge points in the topological sort; content corresponding to knowledge points with earlier positions has higher learning priority. Learning path rationality analysis checks for knowledge jumps or missing prerequisite knowledge in the candidate content sequence. Jump detection judges the rationality of the learning difficulty gradient by comparing the distance between knowledge points of adjacent content. The knowledge structure coherence evaluation comprehensively considers the position of content in the knowledge dependency graph, the completeness of its relation to prerequisite knowledge, and the continuity of the learning path. The evaluation result is expressed numerically as the rationality of the recommended knowledge structure.

[0045] Cross-round contextual consistency calculation comprehensively analyzes learning coherence scores and knowledge structure coherence evaluations to assess the role of content recommendations in promoting the user's overall learning process. Contextual consistency refers to the degree to which recommended content maintains logical coherence and goal consistency across multiple learning sessions. Recommendations with high consistency can form a learning loop and build a knowledge system. The calculation process employs a multi-objective optimization method, using learning coherence and knowledge structure coherence as two optimization objectives, and seeking a balance point through the concept of Pareto optimality. The weighting mechanism dynamically adjusts the importance of the two objectives based on the user's learning stage and cognitive characteristics. Beginners prioritize the completeness of the knowledge structure, while advanced learners focus more on continuity with previous learning. The degree of promotion of learning progress is quantified by analyzing the contribution of recommended content to the completeness of the user's knowledge graph. The contribution calculation considers the knowledge gaps filled, the weak links strengthened, and the new knowledge areas expanded by the content. The contextual coherence evaluation results represent the value and importance of each candidate content in the overall learning process in the form of a comprehensive score. The score comprehensively reflects the immediate relevance and long-term learning value of the content.

[0046] The ranking and filtering process optimizes and reorganizes candidate educational content based on contextual coherence assessment results, generating personalized recommendation sequences that align with users' learning paths and cognitive development patterns. The ranking algorithm employs a multi-keyword ranking strategy: the primary keyword is contextual coherence score, the secondary keyword is cognitive fit score, and the tertiary keyword is semantic matching score. This multi-level ranking ensures optimal configuration across multiple dimensions. The filtering mechanism sets a limit on the number of recommendations, controlling the quantity based on the IPTV interface's display capabilities and user convenience. This limit avoids the negative impact of selection overload on user decision-making. The personalized recommendation sequence organizes the content list according to the ranking results. Each item in the list includes a title, description, difficulty level, estimated learning time, and other information to help users make appropriate learning choices. The intelligent navigation results for IPTV educational content are output in a structured data format, including a list of recommended content, explanations of recommendations, and suggested learning paths. These navigation results are presented to users through the IPTV user interface, supporting voice control and remote control operation.

[0047] In one specific embodiment, the process of performing cross-round contextual consistency calculation based on learning coherence score and knowledge structure coherence evaluation can specifically include the following steps: The learning coherence score and the knowledge structure coherence evaluation are weighted and integrated, and coherence weight coefficient and structure weight coefficient are set to obtain a comprehensive coherence index. A temporal context window is constructed based on multi-round learning history data in the educational session state. The knowledge acquisition pattern and learning progress evolution of users in continuous learning sessions are analyzed to obtain cross-round learning pattern vectors. By performing correlation analysis between the comprehensive coherence index and the cross-round learning pattern vector, the contribution of candidate educational content to users' long-term learning goals and the support for the integrity of the knowledge system are calculated, and a quantitative value of the learning process promotion is obtained. Based on the learning process, a contextual consistency scoring mechanism is established using quantifiable values ​​to assess the rationality and necessity of each candidate educational content within the user's overall learning context, thereby obtaining contextual coherence assessment results.

[0048] Specifically, the weighted fusion calculation combines the learning coherence score and the knowledge structure coherence evaluation numerically. By setting coherence weight coefficients and structure weight coefficients, it controls the importance ratio of the two evaluation dimensions, addressing the problem in existing technologies of balancing immediate learning needs with long-term knowledge structure construction. The coherence weight coefficient reflects the degree of influence of the user's historical learning trajectory on the current recommendation; a higher coefficient value indicates greater emphasis on continuity with historical learning, while a lower coefficient value indicates greater focus on the systematic nature of the knowledge system. The structure weight coefficient measures the strength of the role of knowledge dependencies in recommendation decisions; a high coefficient value prioritizes content with complete prior knowledge, while a low coefficient value allows for a certain degree of knowledge leap. The weight coefficients are dynamically adjusted according to the user's learning stage and cognitive characteristics. In the initial learning stage, the structure weight coefficient is higher to ensure a solid knowledge foundation, while in the advanced stage, the coherence weight coefficient is higher to maintain learning continuity. The weighted fusion uses a linear combination method, multiplying the learning coherence score by the coherence weight coefficient and the knowledge structure coherence evaluation by the structure weight coefficient. The two products are added together to obtain a comprehensive coherence index, which comprehensively reflects the candidate content's overall performance across the two evaluation dimensions.

[0049] The temporal context window is constructed based on multi-round learning history data within the educational session state, and analyzes changes in user behavior patterns during continuous learning sessions using sliding window technology. The temporal context window is a data analysis technique that divides a user's learning history into fixed-length time segments in chronological order, with each window containing records of learning activities within a specific time range. The window length is determined based on the learning cycle of the educational content and the user's learning frequency; a window that is too short cannot capture learning patterns, while a window that is too long will contain outdated learning information. Knowledge acquisition pattern analysis identifies user learning preferences and habits by statistically analyzing indicators such as the distribution of knowledge points learned by the user within the window, the allocation of learning time, and changes in learning difficulty. Learning progress evolution analysis compares the user's learning outcomes and ability improvements within different time windows, tracking the user's development trajectory in terms of knowledge mastery, learning efficiency, and cognitive level. Cross-round learning pattern vectors encode the results of window analysis into numerical vectors. Each dimension of the vector corresponds to a learning pattern feature, such as learning frequency, difficulty preference, topic concentration, and progress speed; the vector value reflects the intensity of the user's performance on that feature.

[0050] Correlation analysis correlates comprehensive coherence indicators with cross-cycle learning pattern vectors to assess the matching degree between candidate educational content and user learning patterns. The correlation calculation uses the Pearson correlation coefficient method to measure the strength of the linear relationship between two variables. The correlation coefficient ranges from -1 to +1, with positive values ​​indicating positive correlation and negative values ​​indicating negative correlation; the larger the absolute value, the stronger the correlation. The contribution to long-term learning goals is calculated by analyzing the promoting effect of candidate content on various dimensions of the user learning pattern vector. This contribution assesses the content's effect on increasing user learning frequency, enhancing difficulty adaptation ability, and improving the knowledge system. The support for knowledge system integrity analyzes the position and role of candidate content in the user's overall knowledge structure, assessing the contribution of the content to knowledge gaps, strengthening weak links, and building knowledge connections, thus contributing to the integrity of the knowledge system. The quantified value of learning progress promotion combines the contribution to long-term learning goals and the support for knowledge system integrity into a single value using a weighted average method. This quantified value reflects the degree to which candidate content promotes the user's overall learning development.

[0051] The contextual consistency scoring mechanism establishes a comprehensive evaluation system for candidate content based on quantifiable values ​​that promote learning progress, assessing the rationality and necessity of each candidate educational content within the user's overall learning context. The scoring mechanism employs a multi-level evaluation structure: the first level evaluates the matching degree between the content and current learning needs; the second level evaluates the continuity of the content with historical learning; and the third level evaluates the preparatory role of the content for future learning. The rationality assessment analyzes the suitability of candidate content at the current learning stage, including a comprehensive consideration of factors such as difficulty matching, time allocation, and cognitive load. The necessity assessment determines the importance of candidate content in the user's knowledge system construction; highly necessary content plays a crucial role in the integrity of the knowledge structure, while less necessary content is considered supplementary or extensional learning. The contextual coherence assessment combines the results of the rationality and necessity assessments to form a score for each candidate content. Content with higher scores is ranked higher in the recommendation sequence, while content with lower scores is ranked lower or filtered out.

[0052] The above describes the intelligent navigation method for IPTV educational content based on speech recognition in the embodiments of this application. The following describes the intelligent navigation system for IPTV educational content based on speech recognition in the embodiments of this application. Please refer to [link / reference]. Figure 3 One embodiment of the intelligent navigation system for IPTV educational content based on speech recognition in this application includes: The parsing module is used to perform hierarchical structure parsing of IPTV educational content through an educational knowledge graph neural network, and to construct a three-level semantic tree of knowledge points, courses, and subjects. The modeling module is used to model the cognitive level of users' learning behavior on the IPTV platform and generate user cognitive profiles. The generation module is used to perform semantic-level perceptual analysis on the voice data collected by the IPTV remote control, construct the educational conversation state by combining the historical learning trajectory, and generate query representation vectors through the educational knowledge graph neural network. The matching module is used to perform semantic matching calculations between the query representation vector and the three-level semantic tree of knowledge points-courses-subjects, and combine it with the user's cognitive profile to adapt to the cognitive level and filter out candidate educational content. The evaluation module is used to evaluate the learning context coherence of candidate educational content across different rounds based on the educational session status, and generate intelligent navigation results for IPTV educational content.

[0053] above Figure 3 The intelligent navigation system for IPTV educational content based on speech recognition in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The intelligent navigation device for IPTV educational content based on speech recognition in this embodiment of the invention will be described in detail from the perspective of hardware processing.

[0054] Reference Figure 4 This invention also provides a voice recognition-based intelligent navigation device for IPTV educational content. This voice recognition-based intelligent navigation device for IPTV educational content can be a server, and its internal structure can be as follows: Figure 4 As shown, the voice recognition-based IPTV educational content intelligent navigation device includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor, designed as a computer, provides computing and control capabilities. The memory of the voice recognition-based IPTV educational content intelligent navigation device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the voice recognition-based IPTV educational content intelligent navigation device stores the data corresponding to this embodiment. The network interface of the voice recognition-based IPTV educational content intelligent navigation device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.

[0055] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the intelligent navigation device for IPTV educational content based on voice recognition applied thereto.

[0056] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the speech recognition-based intelligent navigation method for IPTV educational content.

[0057] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0058] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a voice recognition-based IPTV educational content intelligent navigation device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0059] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for intelligent navigation of IPTV educational content based on speech recognition, characterized in that, The method includes: Step S1: Use an educational knowledge graph neural network to perform hierarchical structure analysis on IPTV educational content and construct a three-level semantic tree of knowledge points, courses, and subjects; Step S2: Model the cognitive level of users' learning behavior on the IPTV platform and generate user cognitive profiles; Step S3: Perform semantic hierarchical perceptual analysis on the voice data collected by the IPTV remote control, construct the educational conversation state by combining the historical learning trajectory, and generate a query representation vector through the educational knowledge graph neural network; Step S4: Perform semantic matching calculation between the query representation vector and the three-level semantic tree of knowledge point-course-subject, combine it with the user cognitive profile to adapt the cognitive level, and filter out candidate educational content; Step S5: Based on the educational session state, evaluate the cross-round learning context coherence of the candidate educational content and generate intelligent navigation results for IPTV educational content.

2. The intelligent navigation method for IPTV educational content based on speech recognition according to claim 1, characterized in that, Step S1 includes: The text descriptions and video tags of IPTV educational content are preprocessed to extract keywords and knowledge concept identifiers in the education field, thus obtaining the original feature data of the educational content. The original feature data of the educational content is input into the graph embedding layer of the educational knowledge graph neural network for vector encoding to obtain the educational content node embedding vector; Based on the embedded vectors of the educational content nodes, the semantic association weights between knowledge points are calculated through the graph attention mechanism of the educational knowledge graph neural network, a knowledge point layer relationship graph is constructed, and a knowledge point layer node set is obtained. Based on the knowledge point layer node set, the course level clustering is performed through the hierarchical aggregation module of the educational knowledge graph neural network. The relevant knowledge points are organized into course nodes, resulting in a course layer node set and a subject layer node set, which are combined to form a three-level semantic tree of knowledge point-course-subject.

3. The intelligent navigation method for IPTV educational content based on speech recognition according to claim 1, characterized in that, Step S2 includes: Data on users' viewing time, pause frequency, replay frequency, and skipping behavior on the IPTV platform are collected to obtain raw data on user learning behavior. The raw data of user learning behavior is subjected to time-series analysis to calculate the user's focus index and comprehension speed index on different educational content, and a learning behavior feature vector is obtained. Based on the learning behavior feature vector and the difficulty tags in the three-level semantic tree of knowledge point-course-subject, the cognitive level is assessed, the user's mastery score in each knowledge domain is calculated, and the cognitive ability assessment matrix is ​​obtained. Based on the cognitive ability assessment matrix, a personal profile data structure containing the user's knowledge base, learning preferences, and cognitive development stage is constructed to obtain the user's cognitive profile.

4. The intelligent navigation method for IPTV educational content based on speech recognition according to claim 1, characterized in that, Step S3 includes: Semantic hierarchical perceptual analysis is performed on the voice data collected by the IPTV remote control to extract educational keywords and intent features from the voice, thereby obtaining the current voice query features; Construct a time-series learning behavior sequence based on users' historical learning trajectory data, analyze the changes in users' learning topics and progress in different time periods, and obtain historical learning trajectory vectors; The current voice query features are fused with the historical learning trajectory vector to construct session state data containing the current query intent and historical learning context, thus obtaining the educational session state. The educational session state is input into the educational knowledge graph neural network for semantic encoding. A vector representation that integrates knowledge relevance is generated through the message passing mechanism of the graph neural network to obtain the query representation vector.

5. The intelligent navigation method for IPTV educational content based on speech recognition according to claim 1, characterized in that, Step S4 includes: The semantic matching score matrix is ​​obtained by calculating the cosine similarity between the query representation vector and the node vectors in the knowledge point-course-subject three-level semantic tree. Based on the semantic matching score matrix, educational content nodes with similarity higher than a preset threshold are filtered to construct a preliminary matching content candidate set and obtain a preliminary list of educational content. The difficulty level of each item in the initial educational content list is matched with the cognitive ability assessment matrix in the user's cognitive profile to calculate the degree of fit between the content difficulty and the user's cognitive level, and a cognitive fit score is obtained. The initial list of educational content is sorted and filtered based on the cognitive fit score, and educational content with a fit score exceeding the cognitive level threshold is retained to obtain candidate educational content.

6. The intelligent navigation method for IPTV educational content based on speech recognition according to claim 1, characterized in that, Step S5 includes: Based on the historical learning trajectory vector in the educational session state, the candidate educational content is analyzed for learning sequence coherence. The correlation strength between each content and the user's historical learning topics is calculated to obtain a learning coherence score. The candidate educational content is sorted according to the knowledge dependency relationship in the three-level semantic tree of knowledge point-course-subject, and the pre-knowledge association and learning path rationality between the content are analyzed to obtain an evaluation of the coherence of the knowledge structure. Based on the learning coherence score and the knowledge structure coherence evaluation, cross-round contextual consistency calculation is performed to evaluate the role of content recommendations in promoting the user's overall learning process, and the contextual coherence evaluation result is obtained. Based on the contextual coherence assessment results, the candidate educational content is sorted and filtered to generate a personalized recommendation sequence that conforms to the user's learning path and cognitive development, thus obtaining intelligent navigation results for IPTV educational content.

7. The intelligent navigation method for IPTV educational content based on speech recognition according to claim 6, characterized in that, The step of calculating cross-round contextual consistency based on the learning coherence score and the knowledge structure coherence evaluation, assessing the role of content recommendations in promoting the user's overall learning process, and obtaining contextual coherence evaluation results includes: The learning coherence score and the knowledge structure coherence evaluation are weighted and integrated, and coherence weight coefficient and structure weight coefficient are set to obtain a comprehensive coherence index. Based on the multi-round learning history data in the educational session state, a temporal context window is constructed to analyze the user's knowledge acquisition pattern and learning progress evolution in the continuous learning session, and to obtain a cross-round learning pattern vector. The correlation analysis is performed between the comprehensive coherence index and the cross-round learning pattern vector to calculate the contribution of candidate educational content to the user's long-term learning goals and the support for the integrity of the knowledge system, thereby obtaining a quantitative value for the promotion of learning progress. Based on the learning process, a contextual consistency scoring mechanism is established using quantifiable values ​​to evaluate the rationality and necessity of each candidate educational content within the user's overall learning context, thereby obtaining contextual coherence evaluation results.

8. An intelligent navigation system for IPTV educational content based on speech recognition, characterized in that, For implementing the intelligent navigation method for IPTV educational content based on speech recognition as described in any one of claims 1-7, the intelligent navigation system for IPTV educational content based on speech recognition comprises: The parsing module is used to perform hierarchical structure parsing of IPTV educational content through an educational knowledge graph neural network, and to construct a three-level semantic tree of knowledge points, courses, and subjects. The modeling module is used to model the cognitive level of users' learning behavior on the IPTV platform and generate user cognitive profiles. The generation module is used to perform semantic-level perceptual analysis on the voice data collected by the IPTV remote control, construct the educational conversation state by combining the historical learning trajectory, and generate query representation vectors through the educational knowledge graph neural network. The matching module is used to perform semantic matching calculations between the query representation vector and the three-level semantic tree of knowledge points-courses-subjects, and combine it with the user's cognitive profile to adapt to the cognitive level and filter out candidate educational content. The evaluation module is used to evaluate the learning context coherence of candidate educational content across different rounds based on the educational session status, and generate intelligent navigation results for IPTV educational content.

9. A smart navigation device for IPTV educational content based on voice recognition, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the intelligent navigation method for IPTV educational content based on voice recognition as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it causes the processor to execute the intelligent navigation method for IPTV educational content based on speech recognition as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Online classroom intelligent recommendation method for education robot

    CN120067442A

  • Online teaching interaction method based on multi-modal knowledge graph, medium and equipment

    CN120339011A

  • Patient education auxiliary system based on deep learning

    CN120356688A

  • Educational resource intelligent recommendation method and system based on big data driving

    CN120653852A

  • Children cognition-based content recommendation method and system

    CN120780747A

Cited By

  • Growth file dynamic recording and tracking system for children with development disorder

    CN121528431A