Iptv education content intelligent navigation method and system based on voice recognition

By constructing an educational knowledge graph neural network and user cognitive profiles, the problem of IPTV education platforms being unable to understand the knowledge level of educational content and user cognition has been solved, enabling personalized educational content recommendations and learning path planning, and improving the coherence and accuracy of the learning experience.

CN120935409BActive Publication Date: 2025-12-23CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511476994.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-12-23
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing IPTV education platforms are unable to understand the knowledge levels and logical relationships of educational content, resulting in a lack of systematic and personalized content recommendations. Furthermore, they cannot accurately assess users' cognitive levels and learning abilities, leading to a mismatch between content difficulty and user capabilities.

Method used

By constructing a three-level semantic tree of knowledge points, courses, and subjects through an educational knowledge graph neural network, and combining user cognitive profiles and semantic level perception parsing, query representation vectors are generated for semantic matching and cognitive level adaptation. This assesses the contextual coherence of learning across rounds and enables intelligent navigation of educational content.

Benefits of technology

It improves the accuracy of educational content recommendations and the coherence of learning paths, ensuring that recommended content matches the user's cognitive level, avoiding content fragmentation and cognitive overload, and enhancing the smoothness of the learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935409B_ABST
    Figure CN120935409B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, and discloses an IPTV education content intelligent navigation method and system based on voice recognition. The method comprises the following steps: constructing a knowledge point-course-subject three-level semantic tree through an education knowledge graph neural network; cognitively modeling user learning behaviors to generate a user cognitive profile; performing semantic analysis on voice data to generate a query vector in combination with a historical trajectory; matching the query vector with the semantic tree and screening candidate content in combination with the cognitive profile; and performing coherence evaluation based on a conversation state to generate a navigation result. The application solves the problem that an existing IPTV education platform cannot understand the knowledge hierarchy of education content and the user cognitive development law, leading to the problem that recommended content lacks systematicness and individualization. The accuracy of education content recommendation and the coherence of a learning path are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to an IPTV education content intelligent navigation method and system based on voice recognition. BACKGROUND

[0002] The existing IPTV education platform mainly adopts content recommendation technology based on keyword matching and collaborative filtering. Users search for education content through remote control input or simple voice instructions, and the system recommends content according to content tags and user historical viewing records. These technologies have certain effects in processing the basic search needs of users, can provide relevant education content recommendations according to the explicit behavior data of users, and have been widely used in commercial online education platforms.

[0003] However, the existing technology has significant deficiencies: first, the traditional keyword matching cannot understand the knowledge level and logical relationship of the education content, resulting in a lack of systematicness and coherence of the recommended content; second, the existing user portrait construction is mainly based on viewing history and click behavior, which cannot accurately evaluate the cognitive level and learning ability of users, and is prone to the problem of mismatch between content difficulty and user ability; third, the traditional voice recognition technology lacks deep understanding of the semantics of the education field, and cannot handle complex education query intentions and context associations. SUMMARY

[0004] The present application provides an IPTV education content intelligent navigation method and system based on voice recognition, which is used to solve the problem that the existing IPTV education platform cannot understand the knowledge level structure of education content and the user cognitive development rule, resulting in a lack of systematicness and individualization of the recommended content. The accuracy of education content recommendation and the coherence of learning path are improved.

[0005] In a first aspect, the present application provides an IPTV education content intelligent navigation method based on voice recognition, which comprises:

[0006] Step S1: performing hierarchical structure analysis on IPTV education content through an education knowledge graph neural network, and constructing a knowledge point-course-discipline three-level semantic tree;

[0007] Step S2: modeling the cognitive level of the user's learning behavior on the IPTV platform, and generating a user cognitive profile;

[0008] Step S3: performing semantic level perception analysis on the voice data collected by the IPTV remote control, constructing an education conversation state in combination with the historical learning trajectory, and generating a query representation vector through the education knowledge graph neural network;

[0009] Step S4: performing semantic matching calculation on the query representation vector and the knowledge point-course-subject three-level semantic tree, combining the user cognitive profile to perform cognitive level adaptation, and screening to obtain candidate educational content;

[0010] Step S5: performing cross-turn learning context coherence evaluation on the candidate educational content based on the education session state, and generating an IPTV educational content intelligent navigation result.

[0011] In a second aspect, the present application provides an IPTV educational content intelligent navigation system based on voice recognition, comprising:

[0012] The parsing module is configured to perform hierarchical structure parsing on the IPTV educational content through an educational knowledge graph neural network, and construct a knowledge point-course-subject three-level semantic tree.

[0013] The modeling module is configured to model the cognitive level of the user's learning behavior on the IPTV platform, and generate a user cognitive profile.

[0014] The generation module is configured to perform semantic hierarchical perception analysis on voice data collected by an IPTV remote controller, construct an education session state in combination with a historical learning trajectory, and generate a query representation vector through an educational knowledge graph neural network.

[0015] The matching module is configured to perform semantic matching calculation on the query representation vector and the knowledge point-course-subject three-level semantic tree, and perform cognitive level adaptation in combination with the user cognitive profile, to screen to obtain candidate educational content.

[0016] The evaluation module is configured to perform cross-turn learning context coherence evaluation on the candidate educational content based on the education session state, and generate an IPTV educational content intelligent navigation result.

[0017] In a third aspect, an IPTV educational content intelligent navigation device based on voice recognition is provided, comprising a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory, so that the IPTV educational content intelligent navigation device based on voice recognition performs the above-mentioned IPTV educational content intelligent navigation method based on voice recognition.

[0018] In a fourth aspect, a computer readable storage medium is provided, the computer readable storage medium storing instructions, when running on a computer, causing the computer to perform the above-mentioned IPTV educational content intelligent navigation method based on voice recognition.

[0019] The technical scheme provided in the application fundamentally solves the problem that the prior art cannot understand the knowledge hierarchy of education content, enables the IPTV platform to accurately grasp the logical relationship and dependency relationship between different education content, and avoids the content recommendation fragmentation problem caused by traditional keyword matching. The generation of the user cognitive profile establishes a precise cognitive level evaluation system by deeply analyzing the learning behavior pattern of the user on the IPTV platform, and can more accurately identify the real learning ability and cognitive development stage of the user compared with the coarse user portrait of the prior art which only relies on the viewing history. The technical innovation of constructing the education conversation state by combining the semantic level perception analysis and the historical learning track enables the system to understand the deep education intention of the user voice query, and organically integrates the current demand with the historical learning background, thereby significantly improving the intelligent level of voice interaction. The semantic matching calculation of the query representation vector and the three-level semantic tree in combination with the cognitive level adaptation mechanism of the user cognitive profile realizes the double precise positioning of content recommendation, which guarantees semantic relevance and ensures cognitive suitability.

[0020] The message passing mechanism of the graph neural network realizes the structured representation and association modeling of education knowledge, so that the system can simulate the knowledge organization thinking of human teachers and provide a scientific basis for personalized learning path planning. The application value of the cross-turn learning context continuity evaluation algorithm lies in its ability to maintain the knowledge continuity and cognitive consistency of the user in multiple learning sessions, avoiding the learning content jumping and cognitive overload problems commonly seen in traditional recommendation systems, and being particularly suitable for education scenarios that require systematic knowledge construction. The unique contribution of the semantic level perception analysis algorithm in education voice interaction lies in its ability to identify language expression patterns and concept levels specific to the education field, which has a significant accuracy advantage over general voice recognition technology in processing complex education queries, effectively reducing the user's interaction cost and improving the smoothness of the learning experience. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative labor.

[0022] Figure 1 An embodiment schematic diagram of the IPTV education content intelligent navigation method based on voice recognition in the embodiments of the present application;

[0023] Figure 2 A schematic diagram of the change of the length of time of the user watching education content in a week on the IPTV platform in the embodiments of the present application;

[0024] Figure 3 Figure 1 is a schematic diagram of an embodiment of the IPTV education content intelligent navigation system based on voice recognition in the present application;

[0025] Figure 4 Figure 2 is a schematic diagram of the structure of an embodiment of the IPTV education content intelligent navigation device based on voice recognition in the present application. DETAILED DESCRIPTION

[0026] The present application provides an IPTV education content intelligent navigation method and system based on voice recognition. The terms first, second, third, fourth, etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term includes or has and any variation thereof is intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] For the sake of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 An embodiment of the IPTV education content intelligent navigation method based on voice recognition in the present application includes:

[0028] Step S1: performing hierarchical structure analysis on the IPTV education content through an education knowledge graph neural network, and constructing a knowledge point-course-discipline three-level semantic tree;

[0029] Step S2: modeling the cognitive level of the user's learning behavior on the IPTV platform, and generating a user cognitive profile;

[0030] Step S3: performing semantic hierarchical perception analysis on the voice data collected by the IPTV remote controller, constructing an education conversation state in combination with the historical learning trajectory, and generating a query representation vector through the education knowledge graph neural network;

[0031] Step S4: performing semantic matching calculation on the query representation vector and the knowledge point-course-discipline three-level semantic tree, performing cognitive level adaptation in combination with the user cognitive profile, and screening to obtain candidate education content;

[0032] Step S5: performing cross-round learning context coherence evaluation on the candidate education content based on the education conversation state, and generating an IPTV education content intelligent navigation result.

[0033] It can be understood that the subject performing the present application can be an IPTV education content intelligent navigation system based on voice recognition, and can also be a terminal or a server, which is not limited here. The server is taken as an example for performing the present application.

[0034] Specifically, the hierarchical semantic structure of the IPTV education content is constructed through the education knowledge graph neural network, so as to solve the problem that the voice navigation system in the prior art cannot understand the knowledge dependency relationship of the education content. The education knowledge graph neural network is a graph neural network architecture specially processing the knowledge relationship in the education field, which includes three core components of a graph embedding layer, a graph attention mechanism and a hierarchical aggregation function. The graph embedding layer converts the text description and video label of the IPTV education content into a numerical vector, the graph attention mechanism calculates the semantic correlation strength between different knowledge points, and the hierarchical aggregation function organizes the knowledge points into course and subject levels according to the educational principles. Specifically, the text preprocessing first extracts the education keywords such as function, derivative and calculus from the education video, and then the graph embedding layer encodes these keywords into a 256-dimensional vector representation. The graph attention mechanism identifies the pre-learning relationship between the function and the derivative by calculating the attention weight between the vectors, and the hierarchical aggregation function aggregates the related knowledge points such as function, derivative and limit into a calculus basic course node, and further aggregates multiple related courses into a mathematics subject node, forming a three-level semantic tree structure of knowledge point-course-subject.

[0035] The user cognitive profile is generated by analyzing the learning behavior data on the IPTV platform. The learning behavior raw data includes four dimensions of watching time, pause frequency, replay times and jumping behavior. The time series analysis statistically analyzes the behavior pattern change of the user in a specific time window, the concentration index is calculated by the ratio of the watching time to the total content time, and the understanding speed index is determined according to the inverse relationship between the replay times and the pause frequency. The cognitive level evaluation associates and analyzes the learning behavior feature vector with the difficulty label in the three-level semantic tree, calculates the mastery degree score of the user in different knowledge fields such as mathematics, physics and chemistry, and forms a cognitive ability evaluation matrix. The user cognitive profile stores the knowledge base level, learning preference type and cognitive development stage information of the user as a data structure, wherein the knowledge base level records the set of knowledge points mastered by the user, the learning preference type identifies the content form preferred by the user, and the cognitive development stage reflects the current learning ability level of the user.

[0036] The semantic level perception analysis process extracts semantic information specific to the education field by processing voice data collected by an IPTV remote controller. The voice data is first preprocessed by noise reduction and endpoint detection, and then semantic analysis is performed to identify education keywords and learning intentions in the voice. The education keywords include subject names, knowledge point terms, and learning action vocabulary, and the learning intentions include course search, knowledge query, and difficulty adjustment. Historical learning trajectory data records the user's past learning theme sequence and timestamp information, and the time sequence learning behavior sequence analysis identifies the user's learning mode and progress evolution trend. The current voice query features and historical learning trajectory vectors are combined through a feature fusion algorithm to generate an education conversation state that includes immediate needs and historical background. The message passing mechanism of the education knowledge graph neural network propagates semantic information on the graph structure, and the node features are updated through weighted aggregation of neighbor nodes, outputting a query representation vector that integrates knowledge relevance.

[0037] The semantic matching calculation uses a cosine similarity algorithm to measure the relevance of the query representation vector to each node of the three-level semantic tree. The cosine similarity evaluates semantic similarity by calculating the cosine of the angle between two vectors, with a value range of -1 to 1. The closer to 1, the more similar the semantics. The semantic matching score matrix records the similarity values of the query vector to all education content nodes, and a preset threshold is used to filter low-relevance content to form a preliminary education content list. Cognitive level adaptation calculates the matching degree between content difficulty levels and the cognitive ability evaluation matrix in the user's cognitive profile, evaluates the fit between content cognitive requirements and user cognitive levels, and the cognitive adaptation score reflects the suitability of the content for the user. Content above the cognitive level threshold is retained as candidate education content.

[0038] Cross-session learning context coherence evaluation analyzes the reasonableness of the learning sequence of candidate content based on the education conversation state. The learning coherence score is obtained by calculating the semantic relevance of candidate content to historical learning themes, and content with high relevance obtains a higher coherence score. Knowledge structure coherence evaluation determines the logical order between content based on the knowledge dependency relationship in the three-level semantic tree, and content with sufficient mastery of prerequisite knowledge obtains a better structure evaluation. The cross-session learning mode vector records the user's knowledge acquisition path and progress change law in multiple learning sessions, and the learning progress promotion degree is quantified by analyzing the contribution of candidate content to the user's long-term learning goals. The context consistency score mechanism considers the position and role of content in the overall learning context, and the evaluation results guide content sorting and selection to generate personalized IPTV education content intelligent navigation results that conform to the user's cognitive development law.

[0039] In a specific embodiment, step S1 includes:

[0040] The text description and video tags of the IPTV education content are preprocessed, education field keywords and knowledge concept identifiers are extracted, and original feature data of the education content is obtained;

[0041] The original feature data of the education content is input into a graph embedding layer of an education knowledge graph neural network for vectorization coding, and an education content node embedding vector is obtained;

[0042] Based on the education content node embedding vector, a semantic association weight between knowledge points is calculated through a graph attention mechanism of the education knowledge graph neural network, a knowledge point layer relationship graph is constructed, and a knowledge point layer node set is obtained;

[0043] According to the knowledge point layer node set, a course level clustering is performed through a hierarchical aggregation module of the education knowledge graph neural network, related knowledge points are organized into course nodes, a course layer node set and a discipline layer node set are obtained, and a knowledge point-course-discipline three-level semantic tree is formed by combination.

[0044] Specifically, the IPTV education content preprocessing extracts structured information from the metadata of the education video, solving the problem that the existing technology cannot understand the semantic hierarchy of the education content. The preprocessing first performs text cleaning on the title, description text and tag information of the education video, removes punctuation marks, stop words and irrelevant characters, and then identifies key concepts through education field dictionary matching. The education field keywords include discipline terms such as functions, derivatives and chemical bonds, and the knowledge concept identifiers include cognitive level markers such as basic concepts, application examples and comprehensive exercises. The description content is divided into independent lexical units by text segmentation, the grammatical properties of nouns, verbs and adjectives are identified by part-of-speech tagging, and the names of persons, places and professional terms are extracted by named entity recognition. The original feature data of the education content is stored in a structured form, including numerical representation of three dimensions of lexical sequence, word frequency statistics and semantic tags.

[0045] The graph embedding layer converts the original feature data of the education content into a high-dimensional vector representation, and uses the Word2Vec algorithm to map the words to a continuous numerical space. The Word2Vec algorithm learns word vectors based on the context relationship of words in the text, and captures the semantic similarity between words through neural network training. The graph embedding layer receives the preprocessed word sequence, each word corresponds to a fixed-dimensional vector representation, and multiple word vectors are combined into a content-level embedding vector through average pooling or attention weighted summation. The education content node embedding vector has a fixed dimension length, and each numerical component in the vector represents the feature intensity of the content in a specific semantic dimension. The numerical distribution of the embedding vector reflects the semantic features of the education content, and the contents with similar semantics are closer in the vector space, and the contents with larger semantic difference are farther in the vector space.

[0046] The graph attention mechanism calculates the semantic correlation weight between knowledge points based on the education content node embedding vector, and analyzes the mutual dependence relationship between vectors by using a self-attention algorithm. The self-attention algorithm calculates an attention score through linear transformation of a query vector, a key vector and a value vector, and the attention score is normalized by using a softmax to obtain a weight distribution. The graph attention mechanism takes the embedding vector of each education content as the query vector and the key vector at the same time, and calculates the attention weight between any two contents. The attention weight value reflects the correlation strength between knowledge points, and the knowledge points with high weight values have stronger semantic correlation or learning dependence relationship. The knowledge point layer relationship graph stores the attention weight in the form of an adjacency matrix, and each element in the matrix represents the correlation strength of the corresponding knowledge point pair. The knowledge point layer node set includes all education content nodes participating in the relationship construction and the weight information related to each other.

[0047] The hierarchical aggregation organizes the knowledge point layer node set into course and discipline levels by using the education knowledge graph neural network, and identifies the knowledge point groups with similar semantics by using a spectral clustering algorithm. The spectral clustering algorithm performs eigenvalue decomposition based on the Laplacian matrix of the graph, and projects the high-dimensional graph structure into a low-dimensional space for clustering analysis. The Laplacian matrix is calculated by the difference between the degree matrix and the adjacency matrix, and the eigenvectors correspond to the main connected components in the graph structure. The clustering algorithm merges the semantically related knowledge points into course nodes, and the course node represents a set of knowledge points with inherent logical correlation. The course layer node set includes all identified course nodes and the list of knowledge points contained therein, and the discipline layer node set is formed by further aggregation of the course nodes. The knowledge point-course-discipline three-level semantic tree organizes the hierarchical relationship of education content in a tree structure, the leaf nodes of the tree correspond to specific knowledge points, the intermediate nodes correspond to course categories, and the root node corresponds to the discipline field.

[0048] In a specific embodiment, step S2 comprises:

[0049] Data collection is performed on the user's viewing duration, pause frequency, replay times and skipping behavior on the IPTV platform to obtain original data of user learning behavior;

[0050] Time series analysis is performed on the original data of user learning behavior to calculate the concentration index and understanding speed index of the user on different education contents, and a learning behavior feature vector is obtained;

[0051] Based on the learning behavior feature vector and the difficulty labels in the knowledge point-course-discipline three-level semantic tree, a cognitive level evaluation is performed, a mastery degree score of the user in each knowledge field is calculated, and a cognitive ability evaluation matrix is obtained;

[0052] According to the cognitive ability evaluation matrix, a personal profile data structure including the user's knowledge base, learning preference and cognitive development stage is constructed, and a user cognitive profile is obtained.

[0053] Specifically, the user learning behavior data collection records the user's interaction behavior with the educational content through the real-time monitoring function of the IPTV set-top box, solving the problem of inaccurate assessment of user cognitive level in the prior art. The viewing duration data records the continuous time period from the start of playback to the stop of viewing, and the IPTV set-top box records a timestamp every second, and accumulates the total viewing duration. The pause frequency data counts the number of times the user actively pauses the video during viewing, and each pause operation triggers the counter to increment, and the pause frequency is calculated by dividing the total number of pauses by the total viewing duration. The number of replays data records the user's repeated viewing behavior of the same educational content segment, and when the user rewinds and replays the viewed content through the remote control, the replay counter is automatically incremented. The skipping behavior data monitors the user's operation of skipping content through fast forward or chapter skipping, and records the start time point, end time point and skipping duration. The user learning behavior raw data is stored in the form of time series, and each record contains five fields of user identification, content identification, behavior type, timestamp and numerical parameter.

[0054] As Figure 2 shown, Figure 2 shows the change of the user's viewing duration of educational content on the IPTV platform within a week. The horizontal axis represents the date (Monday to Sunday), and the vertical axis represents the viewing duration (unit: minutes). The broken line in the figure reflects the time sequence change characteristics of the viewing duration dimension in the user learning behavior raw data, where the convex points represent time nodes with higher user learning activity, and the concave points represent time nodes with relatively lower learning activity. This data provides basic input for subsequent time series analysis, which is used to calculate the user's concentration index and understanding speed index on different educational content, and then generate a learning behavior feature vector. The data fluctuation pattern shown in the figure reflects the user's individualized learning habits and time arrangement preferences, providing important behavior feature data support for building user cognitive profiles.

[0055] The time series analysis processes the user learning behavior raw data and calculates the concentration index and understanding speed index through the sliding window algorithm. The sliding window algorithm sets a fixed time window length and slides the window on the time axis to analyze the behavior pattern changes in different time periods. The concentration index is calculated by the ratio of continuous watching time length to total content time length. The continuous watching time length refers to the longest watching time period of the user without pausing or jumping. The higher the ratio, the stronger the user's concentration. The understanding speed index is determined based on the inverse relationship between the number of replays and the pause frequency. Contents with more replays or higher pause frequency are considered to be more difficult to understand, and the user's understanding speed is slower. The time series analysis algorithm standardizes the data in four dimensions of watching time, pause frequency, replay number and jumping behavior, eliminating the influence of different dimensions on the analysis results. The standardization process uses the Z-score method to subtract the mean value and divide by the standard deviation to obtain a standard normal distribution value. The learning behavior feature vector stores the standardized behavior indicators in the form of a multi-dimensional array. Each dimension of the vector corresponds to a quantitative representation of a learning behavior.

[0056] The cognitive level evaluation associates the learning behavior feature vector with the difficulty label in the knowledge point-course-subject three-level semantic tree and calculates the user's mastery degree in each knowledge field using the weighted average algorithm. The difficulty label is divided into three levels of primary, intermediate and advanced according to the cognitive complexity of educational content, and each level corresponds to a different numerical weight. The weighted average algorithm multiplies the user's learning performance on content of a specific difficulty level by the corresponding weight, and then sums to obtain the comprehensive mastery score. The calculation process of the mastery score includes the product of the concentration index and the difficulty weight, the matching degree analysis of the understanding speed index and the cognitive requirement, and the corresponding relationship evaluation of the learning completion degree and the content coverage range. The cognitive ability evaluation matrix is organized in a two-dimensional table form, with rows corresponding to different knowledge fields and columns corresponding to different cognitive ability dimensions. Each element in the matrix stores the user's mastery score at the intersection of a specific knowledge field and a cognitive dimension. The rows of the matrix include subject categories such as mathematics, physics, chemistry and Chinese, and the columns include cognitive categories such as memory ability, understanding ability, application ability and analysis ability.

[0057] The user cognitive profile is constructed based on a cognitive ability evaluation matrix to establish a multi-level personal learning portrait data structure. The knowledge base layer records the set of knowledge points mastered by the user. By traversing the cognitive ability evaluation matrix, knowledge points with a mastery level score exceeding a threshold are identified, and the identifiers of these knowledge points are stored in the knowledge base set. The learning preference layer analyzes the user's learning performance differences in different content types, identifies the user's preferred learning methods and content forms. Learning preferences are determined by comparing the user's focus and understanding speed differences in video, audio, and graphic content. The content type with higher focus and understanding speed is marked as the user's preferred type. The cognitive development stage layer assesses the current development level based on the user's comprehensive performance in different cognitive ability dimensions. The cognitive development stage includes specific operation stage, formal operation stage, and other classifications in Piaget's cognitive development theory. The user cognitive profile is organized in a tree structure, with the root node storing the user's basic information, and the child nodes corresponding to the detailed data of the three dimensions of knowledge base, learning preference, and cognitive development stage.

[0058] In a specific embodiment, step S3 comprises:

[0059] The semantic level perception analysis is performed on the voice data collected by the IPTV remote controller to extract the education keywords and intent features in the voice, and obtain the current voice query features;

[0060] Based on the user historical learning trajectory data, a time sequence learning behavior sequence is constructed to analyze the user's learning theme and progress changes in different time periods, and obtain a historical learning trajectory vector;

[0061] The current voice query features and the historical learning trajectory vector are fused and processed to construct a conversation state data containing the current query intent and historical learning context, and obtain an education conversation state;

[0062] The education conversation state is input into the education knowledge graph neural network for semantic encoding, and a vector representation of the fusion knowledge association is generated through the message passing mechanism of the graph neural network, and a query representation vector is obtained.

[0063] Specifically, the semantic level perception parsing process IPTV remote control to collect voice data, through multi-level semantic analysis to extract the semantic information specific to the field of education, solve the problem that the voice recognition can not understand the educational context in the prior art. The voice data is first subjected to acoustic feature extraction, converting the audio signal into a sequence of mel-frequency cepstral coefficients, which are numerical representations reflecting the spectral characteristics of speech and can effectively capture the phoneme information in the speech. The acoustic model uses a recurrent neural network to process the time-series speech features, mapping the sequence of mel-frequency cepstral coefficients to a phoneme probability distribution, which represents the likelihood of different phonemes corresponding to the speech segment. The language model performs probability calculations based on an education domain vocabulary, which includes professional terms and common expressions in subjects such as mathematics, physics, and chemistry. The language model establishes a probability distribution by statistically analyzing the frequency of occurrence and combination patterns of these words in the educational context. The education keyword extraction locates subject terminology, knowledge point names, and learning action vocabulary from the recognized text using a named entity recognition algorithm, which uses a conditional random field model to annotate the semantic categories of the words. The intent feature extraction is based on a pre-defined education intent classification system, including content search, difficulty adjustment, progress query, review, etc. The intent classifier uses a support vector machine algorithm to determine the user's learning goal based on keyword combinations and grammatical structures. The current voice query features are stored in a structured data format, including four fields: recognized text, keyword list, intent category, and confidence score.

[0064] The historical learning trajectory vector is constructed based on the user's past learning behavior sequence, and a time series analysis algorithm is used to identify learning patterns and progress evolution trends. The user's historical learning trajectory data includes timestamp, learning content identifier, learning duration, completion status, and other dimension information. The timestamp records the time of the learning activity, and the learning content identifier is associated with a specific node in the knowledge point-course-subject three-level semantic tree. The time series learning behavior sequence arranges the user's learning activities in chronological order, and each element in the sequence represents a detailed record of a learning session. The learning theme analysis identifies the knowledge areas that the user focuses on at different times through a clustering algorithm. The clustering algorithm groups similar learning activities into the same theme group based on the semantic similarity of the learning content. The progress change analysis evaluates the trend of mastery by comparing the user's repeated learning behavior on the same knowledge point. A decrease in the number of repeated learning indicates an improvement in mastery, and a reduction in learning duration indicates an improvement in understanding efficiency. The historical learning trajectory vector encodes the user's learning preferences and knowledge structure using a bag-of-words model. The bag-of-words model treats the knowledge points that the user has learned as words and the learning frequency as word frequency, generating a high-dimensional sparse vector representation. Each dimension of the vector corresponds to a node in the knowledge point-course-subject three-level semantic tree, and the dimension value reflects the user's familiarity with the knowledge point and the time spent on learning.

[0065] The feature fusion process combines the current voice query feature with the historical learning trajectory vector to construct a comprehensive representation containing both immediate needs and historical background. The fusion algorithm uses an attention mechanism to calculate the relevance weight of the current query and historical learning content. The attention mechanism determines the correlation strength by calculating the semantic similarity between query keywords and historical learning knowledge points. The similarity calculation is based on the cosine distance of word vectors, which are obtained through a pre-trained language model in the education field. The higher the similarity value, the stronger the relevance between the current query and the historical learning content. The conversation state data structure includes four components: current query information, relevant historical learning records, context association weight, and time decay factor. The current query information stores the speech recognition results and intent classification, and the relevant historical learning records include past learning activities related to the semantic of the current query. The context association weight quantifies the degree of association between the current query and the historical learning content, and the time decay factor adjusts the weight size according to the time distance of the historical learning activities, with the weight of the learning activities closer to the current time being larger. The education conversation state is fused into a unified context representation by weighted summation, and the weight in the weighted summation is determined by the product of the context association weight and the time decay factor.

[0066] The message passing mechanism of the educational knowledge graph neural network propagates semantic information based on the graph structure, encoding the education conversation state into a vector representation that integrates knowledge association. The message passing mechanism is the core computation process of graph neural networks, updating node features through information exchange between nodes. Each node collects messages from neighboring nodes and fuses them with its own features. Message calculation uses a combination of linear transformation and nonlinear activation function. Linear transformation maps neighbor node features to message space through matrix multiplication, and nonlinear activation function introduces nonlinear representation ability to enhance the fitting ability of the model. The aggregation function combines messages from multiple neighbor nodes into a single representation, using summation or average pooling to handle variable-length message sequences. The node update function combines the current node feature and the aggregated neighbor message to calculate the new node feature. The update function controls the fusion ratio of historical information and new information through a gating mechanism. After multiple rounds of message passing, the feature vector of the education conversation state node contains semantic information from related knowledge points, and the numerical distribution of the vector reflects the position and association of the current query in the entire knowledge system. The query representation vector, as the feature representation of the education conversation state node, contains the semantic content of the current voice query, the user's historical learning background, and the structured information of related knowledge points.

[0067] In a specific embodiment, step S4 includes:

[0068] The query representation vector is calculated with the cosine similarity of each node vector in the knowledge point-course-subject three-level semantic tree, and a semantic matching score matrix is obtained.

[0069] Filtering the education content nodes with a similarity higher than a preset threshold based on the semantic matching score matrix, constructing a preliminary matched content candidate set, and obtaining a preliminary education content list;

[0070] Matching the difficulty level of each content in the preliminary education content list with the cognitive ability evaluation matrix in the user cognitive profile to calculate the matching degree, evaluate the adaptation degree of content difficulty and user cognitive level, and obtain a cognitive adaptation score;

[0071] According to the cognitive adaptation score, the preliminary education content list is sorted and filtered, and the education content with an adaptation score exceeding the cognitive level threshold is retained to obtain the candidate education content.

[0072] Specifically, the cosine similarity calculation performs semantic matching between the query representation vector and the node vector in the knowledge point-course-discipline three-level semantic tree, solving the problem that the prior art cannot accurately match the user voice query and the education content. Cosine similarity is a mathematical method for measuring the cosine value of the included angle between two vectors in a vector space. The similarity value is obtained by calculating the inner product of the vector divided by the product of the vector length. The query representation vector is the target vector, which contains the semantic features of the user voice query and the historical learning context information. Each node in the three-level semantic tree corresponds to a vector representation, and the node vector encodes the knowledge attributes and semantic features of the education content. The similarity calculation process first calculates the inner product of the query vector and the node vector. The inner product is obtained by multiplying the corresponding dimension values and summing them up. The inner product value reflects the matching degree of the two vectors in each semantic dimension. The vector length is calculated by the square root of the sum of the squares of each dimension value. The length represents the length of the vector in the high-dimensional space. The cosine similarity value ranges from negative one to positive one. The closer the value is to positive one, the higher the semantic similarity. The closer the value is to negative one, the greater the semantic difference. The closer the value is to zero, the less the semantic relevance. The semantic matching score matrix stores the calculation results in the form of a two-dimensional table. The rows correspond to the query representation vector, and the columns correspond to each node in the three-level semantic tree. Each element in the matrix stores the cosine similarity value of the corresponding query and node.

[0073] The screening mechanism identifies the highly relevant education content nodes based on the semantic matching score matrix, and filters the matching results with low similarity through a preset threshold. The preset threshold is determined according to the content quality requirements and user experience standards of the IPTV education platform. If the threshold is set too high, the number of candidate contents will be too small, and if the threshold is set too low, a large number of irrelevant contents will be introduced. The screening algorithm iterates through all the values in the semantic matching score matrix, and records the education content nodes with a similarity higher than the preset threshold in the candidate set. The preliminary matching content candidate set contains all the education content nodes that pass the similarity screening. Each element in the set corresponds to an education content and its similarity score with the query. The preliminary education content list sorts the candidate set in descending order of similarity scores, and the sorting result reflects the matching priority of different education contents with the user query. Each entry in the list contains four fields: content identifier, content title, similarity score, and content attributes, including subject classification, difficulty level, duration, and other metadata information.

[0074] The cognitive adaptation degree calculation matches the content difficulty in the preliminary education content list with the user cognitive profile for analysis, and evaluates the matching degree of content cognitive requirements and user cognitive ability. The content difficulty level is divided into three levels: primary, intermediate, and advanced, according to the cognitive complexity of the education content. Each level corresponds to different cognitive ability requirements and knowledge prerequisites. The cognitive ability evaluation matrix in the user cognitive profile records the user's mastery of different knowledge domains and cognitive dimensions, and the matrix values reflect the user's current cognitive development level and learning ability. The matching degree calculation determines the adaptation degree by comparing the difference between the content difficulty level and the user's cognitive level. The adaptation degree is represented by a numerical score, with a high score indicating a good match between content difficulty and user ability, and a low score indicating unsuitable difficulty. The calculation process first converts the content difficulty level to a numerical representation, with primary corresponding to a value of one, intermediate corresponding to a value of two, and advanced corresponding to a value of three. Then, the mastery score of the corresponding knowledge domain is extracted from the user's cognitive ability evaluation matrix. The adaptation degree score is calculated by the absolute value of the difference between the content difficulty value and the user's mastery score. The smaller the difference, the higher the adaptation degree, and the larger the difference, the lower the adaptation degree. The cognitive adaptation degree score is standardized to eliminate the influence of different dimensions on the comparison result. The standardized score is used for subsequent sorting and screening operations.

[0075] The sorting and screening reorders the preliminary selected education content list according to the cognitive adaptation score, and gives priority to the education content with high adaptation score. The sorting algorithm adopts stable sorting to ensure that the contents with the same adaptation score maintain the original similarity sorting. The stable sorting ensures that the evaluation results of the semantic matching and cognitive adaptation dimensions are reasonably reflected. The cognitive level threshold is set according to the user's learning goal and cognitive development stage. The threshold setting comprehensively considers the user's current ability level and reasonable challenge level. The screening process traverses the sorted content list, retains the education content with an adaptation score exceeding the cognitive level threshold, and eliminates the content items with insufficient adaptation. The candidate education content is the screening result, which contains high-quality content recommendation filtered by semantic matching and cognitive adaptation. The number of contents is controlled within a reasonable range according to the IPTV interface display capability and user selection convenience.

[0076] In a specific embodiment, step S5 comprises:

[0077] Based on the historical learning trajectory vector in the education session state, the learning sequence coherence of the candidate education content is analyzed, the association strength of each content with the user's historical learning theme is calculated, and a learning coherence score is obtained.

[0078] The candidate education content is sorted according to the knowledge dependency relationship in the three-level semantic tree of knowledge points-course-discipline, the pre-knowledge association and learning path rationality between the contents are analyzed, and a knowledge structure coherence evaluation is obtained.

[0079] According to the learning coherence score and the knowledge structure coherence evaluation, the cross-round context consistency is calculated, the promotion effect of content recommendation on the user's overall learning process is evaluated, and a context coherence evaluation result is obtained.

[0080] Based on the context coherence evaluation result, the candidate education content is sorted and screened, a personalized recommendation sequence conforming to the user's learning context and cognitive development is generated, and an IPTV education content intelligent navigation result is obtained.

[0081] Specifically, the learning sequence continuity analysis calculates the relevance strength of the candidate educational content to the user's past learning topics based on the historical learning trajectory vector in the educational conversation state, solving the problem of the inability to maintain cross-session learning continuity in the prior art. The historical learning trajectory vector records the user's learning activity sequence in different time periods, and each dimension in the vector corresponds to a specific node in the knowledge point-course-subject three-level semantic tree, and the dimension value reflects the user's learning input and mastery level at that knowledge point. The relevance strength calculation uses a weighted cosine similarity algorithm to calculate the similarity between the vector representation of the candidate educational content and the historical learning trajectory vector, and the similarity value reflects the matching degree of the content to the user's historical learning topics. The weighting mechanism assigns different weights according to the time distance of the learning activities, and the learning activities closer to the current time have larger weights, and the learning activities farther away have gradually decaying weights. The time decay function uses an exponential decay model, and the decay coefficient is determined according to the forgetting curve characteristics of the educational content, ensuring that the learning continuity analysis considers both historical background and recent learning emphasis. The learning continuity score is obtained by standardizing the weighted similarity calculation, and the standardization eliminates the influence of different learning trajectory lengths on the comparison results, and a higher score indicates that the candidate content has a stronger continuity with the user's learning context.

[0082] The knowledge structure continuity evaluation analyzes the logical order between candidate educational contents based on the knowledge dependency relationship in the knowledge point-course-subject three-level semantic tree. The knowledge dependency relationship describes the prerequisite learning requirements between different knowledge points, and the prerequisite knowledge points must be mastered before the subsequent knowledge points to form a knowledge structure. The dependency relationship graph is represented in the form of a directed graph, where the nodes correspond to knowledge points, the directed edges represent the dependency direction, and the weight of the edge reflects the dependency strength. The topological sorting algorithm processes the knowledge dependency relationship graph to generate a knowledge point learning order that satisfies the dependency constraints, and the topological sorting ensures that the prerequisite knowledge points appear before the knowledge points that depend on them in the sequence. The candidate educational content determines the learning priority according to the position of its associated knowledge points in the topological sorting, and the content corresponding to the knowledge points with a higher position has a higher learning priority. The learning path rationality analysis checks whether there are knowledge jumps or missing prerequisite knowledge in the candidate content sequence, and the jump detection judges the rationality of the learning difficulty gradient by comparing the knowledge point distance of adjacent contents. The knowledge structure continuity evaluation comprehensively considers the position of the content in the knowledge dependency graph, the completeness of the prerequisite knowledge, and the continuity of the learning path, and the evaluation result is represented in the form of a numerical value indicating the rationality degree of the knowledge structure of the content recommendation.

[0083] The cross-session context consistency calculation synthesizes the learning continuity score and the knowledge structure continuity evaluation to analyze comprehensively, and evaluates the promotion of the content recommendation to the overall learning process of the user. The context consistency refers to the degree of maintaining logical continuity and goal consistency of the recommended content in the multi-session learning conversation of the user, and the recommendation with high consistency can form a learning closed loop and knowledge system construction. The calculation process adopts a multi-objective optimization method, takes the learning continuity and the knowledge structure continuity as two optimization objectives, and finds a balance point through the concept of Pareto optimal solution. The weight distribution mechanism dynamically adjusts the importance of the two objectives according to the learning stage and cognitive characteristics of the user, and beginners pay more attention to the integrity of the knowledge structure, and advanced learners pay more attention to the continuity with the historical learning. The learning process promotion degree is quantified by analyzing the contribution degree of the recommended content to the integrity of the knowledge graph of the user, and the contribution degree calculation considers the knowledge gap filled by the content, the weak link strengthened, and the new knowledge field expanded. The context continuity evaluation result is expressed in the form of a comprehensive score to represent the value and importance of each candidate content in the overall learning process, and the score comprehensively reflects the instant relevance and long-term learning value of the content.

[0084] The sorting and screening optimizes and reorganizes the candidate educational content based on the context continuity evaluation result, and generates a personalized recommendation sequence conforming to the learning context and cognitive development law of the user. The sorting algorithm adopts a multi-key sorting strategy, the primary key is the context continuity score, the secondary key is the cognitive adaptation degree score, and the third key is the semantic matching score, and the multi-level sorting ensures that the recommended result is optimally configured in multiple dimensions. The screening mechanism sets an upper limit of the number of recommendations, controls the number of recommended contents according to the display capability of the IPTV interface and the convenience of user selection, and the number limit avoids the negative influence of selection overload on user decision. The personalized recommendation sequence organizes a content list according to the sorting result, and each content item in the list contains title, introduction, difficulty identification, estimated learning time and other information, and the information display helps the user to make appropriate learning selection. The IPTV educational content intelligent navigation result is output in a structured data format, and contains navigation information such as the recommended content list, the recommended reason explanation, and the learning path suggestion. The navigation result is presented to the user through the IPTV user interface, and supports the interactive mode of voice control and remote control operation.

[0085] In a specific embodiment, the performing step of performing a cross-session context consistency calculation according to the learning continuity score and the knowledge structure continuity evaluation can specifically include the following steps:

[0086] The learning continuity score and the knowledge structure continuity evaluation are weighted and fused to calculate, a continuity weight coefficient and a structure weight coefficient are set, and a comprehensive continuity index is obtained;

[0087] Based on the multi-round learning history data in the education conversation state, a timing context window is constructed, the knowledge acquisition mode and learning progress evolution of the user in the continuous learning conversation are analyzed, and a cross-round learning mode vector is obtained;

[0088] The comprehensive continuity index and the cross-round learning mode vector are analyzed for correlation, the contribution of the candidate education content to the long-term learning goal of the user and the support degree of the knowledge system integrity are calculated, and a learning progress promotion quantitative value is obtained;

[0089] According to the learning progress promotion quantitative value, a context consistency scoring mechanism is established to evaluate the rationality and necessity of each candidate education content in the overall learning context of the user, and a context continuity evaluation result is obtained.

[0090] Specifically, the weighted fusion calculation combines the learning continuity score and the knowledge structure continuity evaluation into a numerical value, and by setting the continuity weight coefficient and the structure weight coefficient, the importance proportion of the two evaluation dimensions is controlled, solving the problem that the existing technology cannot balance the immediate learning needs and long-term knowledge structure construction. The continuity weight coefficient reflects the influence degree of the user's historical learning trajectory on the current recommendation, and the higher the coefficient value is, the more attention is paid to the continuity of historical learning, and the lower the coefficient value is, the more attention is paid to the systematicness of the knowledge system. The structure weight coefficient measures the role intensity of knowledge dependency relationship in the recommendation decision, and a high coefficient value preferentially recommends content with complete prerequisite knowledge, and a low coefficient value allows a certain degree of knowledge jump. The setting of the weight coefficient is dynamically adjusted according to the learning stage and cognitive characteristics of the user, and the structure weight coefficient is higher in the initial learning stage to ensure a solid knowledge foundation, and the continuity weight coefficient is higher in the advanced stage to maintain learning continuity. The weighted fusion adopts a linear combination method, multiplies the learning continuity score by the continuity weight coefficient, multiplies the knowledge structure continuity evaluation by the structure weight coefficient, and adds the two products to obtain the comprehensive continuity index, which comprehensively reflects the comprehensive performance of the candidate content in the two evaluation dimensions.

[0091] The time series context window is constructed based on the multi-round learning history data in the education session state, and the behavior pattern change of the user in the continuous learning session is analyzed through a sliding window technique. The time series context window is a data analysis technique that divides the learning history of the user in chronological order into fixed-length time segments, and each window contains learning activity records within a specific time range. The window length is determined according to the learning period of the education content and the learning frequency of the user. If the window length is too short, the learning pattern cannot be captured, and if the window length is too long, outdated learning information will be included. The knowledge acquisition pattern analysis identifies the learning preferences and habits of the user by statistical indicators such as the distribution of knowledge points learned by the user within the window, the allocation of learning time, and the change of learning difficulty. The learning progress evolution analysis compares the learning achievements and ability improvement of the user in different time windows, and tracks the development trajectory of the user in terms of knowledge mastery, learning efficiency, and cognitive level. The cross-round learning mode vector encodes the results of window analysis into a numerical vector, and each dimension of the vector corresponds to a learning mode feature, such as learning frequency, difficulty preference, theme concentration, and progress speed. The vector value reflects the performance intensity of the user in that feature.

[0092] The correlation analysis associates the coherence indicators and the cross-round learning mode vector to evaluate the matching degree of the candidate education content and the user learning mode. The correlation calculation uses the Pearson correlation coefficient method to measure the linear relationship strength between two variables. The correlation coefficient value ranges from negative one to positive one, positive value indicates positive correlation, negative value indicates negative correlation, and the absolute value is larger, indicating stronger correlation. The long-term learning goal contribution degree is calculated by analyzing the promoting effect of the candidate content on each dimension of the user learning mode vector. The contribution degree evaluates the promoting effect of the content on the user's learning frequency improvement, difficulty adaptation ability enhancement, and knowledge system perfection. The knowledge system integrity support degree analyzes the position and role of the candidate content in the user's overall knowledge structure, and evaluates the contribution of the content to the knowledge system integrity, such as filling the knowledge gap, strengthening the weak link, and constructing the knowledge connection. The learning progress promotion quantization value combines the long-term learning goal contribution degree and the knowledge system integrity support degree into a single numerical value through weighted average method, and the quantization value reflects the promotion degree of the candidate content to the overall learning development of the user.

[0093] The context consistency scoring mechanism is based on the learning progress facilitation metric to establish a comprehensive evaluation system of candidate content, and to evaluate the rationality and necessity of each candidate educational content in the overall learning context of the user. The scoring mechanism adopts a multi-level evaluation structure, the first level evaluates the matching degree of the content with the current learning needs, the second level evaluates the continuity of the content to the historical learning, and the third level evaluates the preparatory role of the content to the future learning. The rationality evaluation analyzes the appropriateness of the candidate content in the current learning stage, including the comprehensive consideration of factors such as difficulty matching, time arrangement, and cognitive load. The necessity evaluation judges the importance of the candidate content in the construction of the user's knowledge system, and the content with high necessity plays a key role in the integrity of the knowledge structure, and the content with low necessity belongs to supplementary or expanded learning. The context coherence evaluation combines the results of rationality and necessity evaluation to form the score of each candidate content, and the content with high score is ranked in the front in the recommendation sequence, and the content with low score is ranked in the back or filtered.

[0094] The IPTV educational content intelligent navigation method based on voice recognition in the embodiments of the present application is described above, and the IPTV educational content intelligent navigation system based on voice recognition in the embodiments of the present application is described below. Please refer to Figure 3 An embodiment of the IPTV educational content intelligent navigation system based on voice recognition in the embodiments of the present application includes:

[0095] The parsing module is configured to perform hierarchical structure parsing on the IPTV educational content through the educational knowledge graph neural network, and construct a knowledge point-course-discipline three-level semantic tree.

[0096] The modeling module is configured to model the learning behavior of the user on the IPTV platform to generate a user cognitive profile.

[0097] The generation module is configured to perform semantic hierarchical perception analysis on the voice data collected by the IPTV remote controller, construct an educational conversation state in combination with the historical learning track, and generate a query representation vector through the educational knowledge graph neural network.

[0098] The matching module is configured to perform semantic matching calculation on the query representation vector and the knowledge point-course-discipline three-level semantic tree, and perform cognitive level adaptation in combination with the user cognitive profile to filter and obtain candidate educational content.

[0099] The evaluation module is configured to perform cross-turn learning context coherence evaluation on the candidate educational content based on the educational conversation state to generate an IPTV educational content intelligent navigation result.

[0100] The above Figure 3The voice recognition based IPTV education content intelligent navigation system in the embodiment of the present application is described in detail from the perspective of the modular functional entity, and the voice recognition based IPTV education content intelligent navigation device in the embodiment of the present application is described in detail from the perspective of hardware processing.

[0101] With reference to Figure 4 The voice recognition based IPTV education content intelligent navigation device in the embodiment of the present application can be a server, and the internal structure thereof can be as shown in Figure 4 The voice recognition based IPTV education content intelligent navigation device comprises a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. The processor of the computer is used to provide computing and control capabilities. The memory of the voice recognition based IPTV education content intelligent navigation device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the voice recognition based IPTV education content intelligent navigation device is used to store the corresponding data in the embodiment. The network interface of the voice recognition based IPTV education content intelligent navigation device is used to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the above method.

[0102] Those skilled in the art can understand that Figure 4 The structure shown in the embodiment is only a block diagram of part of the structure related to the present application scheme, and does not constitute a limitation on the voice recognition based IPTV education content intelligent navigation device to which the present application scheme is applied.

[0103] The present application further provides a computer readable storage medium, which can be a non-volatile computer readable storage medium or a volatile computer readable storage medium. The computer readable storage medium stores instructions, and when the instructions are run on a computer, the computer executes the steps of the voice recognition based IPTV education content intelligent navigation method.

[0104] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, system and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described herein.

[0105] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a voice recognition based IPTV education content intelligent navigation device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0106] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A voice recognition based intelligent navigation method for IPTV educational contents, characterized in that, The method comprises: Step S1: hierarchical structure analysis of IPTV education content is performed by an education knowledge graph neural network to construct a three-level semantic tree of knowledge points-courses-disciplines; Step S2: a cognitive level model of the learning behavior of a user on an IPTV platform is established to generate a user cognitive profile; Step S3: semantic hierarchical perception analysis is performed on voice data collected by an IPTV remote controller, and an education conversation state is constructed in combination with historical learning tracks to generate a query representation vector by the education knowledge graph neural network, comprising: semantic hierarchical perception analysis is performed on voice data collected by an IPTV remote controller to extract education keywords and intent features in the voice to obtain current voice query features; a time sequence learning behavior sequence is constructed based on historical learning track data of a user to analyze learning themes and progress changes of the user at different time periods to obtain a historical learning track vector; the current voice query features and the historical learning track vector are fused to construct conversation state data containing current query intent and historical learning context to obtain an education conversation state; wherein the conversation state data contains current query information, relevant historical learning records, context association weights and time decay factors, and the education conversation state fuses multiple historical learning records into a unified context representation by weighted summation, and the weights in the weighted summation are determined by the product of the context association weights and the time decay factors; the education conversation state is input into the education knowledge graph neural network for semantic coding to generate a vector representation fused with knowledge association by a message passing mechanism of the graph neural network to obtain the query representation vector; Step S4: semantic matching calculation is performed on the query representation vector and the three-level semantic tree of knowledge points-courses-disciplines, and cognitive level adaptation is performed in combination with the user cognitive profile to screen candidate education content; Step S5: cross-episode learning context coherence evaluation is performed on the candidate education content based on the education conversation state to generate an IPTV education content intelligent navigation result.

2. The voice recognition based intelligent navigation method for IPTV educational contents according to claim 1, wherein, The step S1 comprises: Text descriptions and video tags of IPTV education content are preprocessed to extract education field keywords and knowledge concept identifiers to obtain education content original feature data; The education content original feature data is input into a graph embedding layer of the education knowledge graph neural network for vectorization coding to obtain education content node embedding vectors; Based on the education content node embedding vectors, a graph attention mechanism of the education knowledge graph neural network is used to calculate semantic association weights between knowledge points to construct a knowledge point layer relationship graph to obtain a knowledge point layer node set; Based on the knowledge point layer node set, a hierarchical aggregation module of the education knowledge graph neural network is used for course level clustering to organize related knowledge points into course nodes to obtain a course layer node set and a discipline layer node set, which are combined to form a three-level semantic tree of knowledge points-courses-disciplines.

3. The voice recognition based intelligent navigation method for IPTV educational contents according to claim 1, wherein, The step S2 comprises: Data collection is performed on the viewing duration, pause frequency, replay times and skipping behavior of a user on an IPTV platform to obtain user learning behavior original data; The user learning behavior original data is subjected to time series analysis, and a concentration index and an understanding speed index of the user on different education contents are calculated to obtain a learning behavior feature vector; Based on the learning behavior feature vector and the difficulty labels in the knowledge point-course-subject three-level semantic tree, a cognitive level is evaluated, a mastery degree score of the user in each knowledge field is calculated, and a cognitive ability evaluation matrix is obtained; According to the cognitive ability evaluation matrix, a personal archive data structure including a user knowledge base, a learning preference and a cognitive development stage is constructed, and a user cognitive archive is obtained.

4. The voice recognition based intelligent navigation method for IPTV educational contents according to claim 1, wherein, The step S4 comprises: The query representation vector is subjected to cosine similarity calculation with each node vector in the knowledge point-course-subject three-level semantic tree to obtain a semantic matching score matrix; Based on the semantic matching score matrix, an education content node with a similarity higher than a preset threshold is screened, a preliminary matched content candidate set is constructed, and an initial education content list is obtained; The difficulty level of each content in the initial education content list is matched with the cognitive ability evaluation matrix in the user cognitive archive, the adaptation degree of content difficulty and user cognitive level is evaluated, and a cognitive adaptation degree score is obtained; According to the cognitive adaptation degree score, the initial education content list is sorted and screened, and an education content with an adaptation degree score exceeding a cognitive level threshold is retained to obtain a candidate education content.

5. The voice recognition based intelligent navigation method for IPTV educational contents according to claim 1, wherein, The step S5 comprises: Based on the historical learning trajectory vector in the education session state, a learning sequence coherence analysis is performed on the candidate education content, an association strength of each content with a user historical learning theme is calculated, and a learning coherence score is obtained; The candidate education content is sorted according to the knowledge dependency relationship in the knowledge point-course-subject three-level semantic tree, and the pre-knowledge association and learning path rationality between the contents are analyzed to obtain a knowledge structure coherence evaluation; According to the learning coherence score and the knowledge structure coherence evaluation, a cross-round context consistency calculation is performed, the promoting effect of content recommendation on the overall learning process of the user is evaluated, and a context coherence evaluation result is obtained; Based on the context coherence evaluation result, the candidate education content is sorted and screened to generate a personalized recommendation sequence conforming to the user learning context and cognitive development, and an IPTV education content intelligent navigation result is obtained.

6. The voice recognition based intelligent navigation method for IPTV educational contents according to claim 5, wherein, The cross-round context consistency calculation based on the learning coherence score and the knowledge structure coherence evaluation, the promoting effect of content recommendation on the overall learning process of the user is evaluated, and a context coherence evaluation result is obtained, comprising: The learning coherence score and the knowledge structure coherence evaluation are subjected to weighted fusion calculation, a coherence weight coefficient and a structure weight coefficient are set, and a comprehensive coherence index is obtained; Based on the multi-round learning historical data in the education session state, a time series context window is constructed, the knowledge acquisition mode and learning progress evolution of the user in a continuous learning session are analyzed, and a cross-round learning mode vector is obtained; Correlation analysis is performed on the comprehensive coherence index and the cross-episode learning mode vector to calculate the contribution of the candidate educational content to the long-term learning goal of the user and the support degree of the knowledge system integrity, and a learning process promotion quantitative value is obtained; A context consistency scoring mechanism is established according to the learning process promotion quantitative value to evaluate the rationality and necessity of each candidate educational content in the overall learning context of the user, and a context coherence evaluation result is obtained.

7. A voice recognition based intelligent navigation system for IPTV educational contents, characterized in that, The voice recognition-based IPTV educational content intelligent navigation system for implementing the voice recognition-based IPTV educational content intelligent navigation method according to any one of claims 1 to 6 comprises: The parsing module is configured to perform hierarchical structure parsing on the IPTV educational content through an educational knowledge graph neural network to construct a knowledge point-course-discipline three-level semantic tree. The modeling module is configured to model the learning behavior of the user on the IPTV platform to generate a user cognitive profile. The generation module is configured to perform semantic hierarchical perception analysis on the voice data collected by the IPTV remote controller, construct an educational conversation state in combination with the historical learning trajectory, generate a query representation vector through the educational knowledge graph neural network, and comprises: performing semantic hierarchical perception analysis on the voice data collected by the IPTV remote controller, extracting educational keywords and intent features in the voice to obtain current voice query features; constructing a time-series learning behavior sequence based on the historical learning trajectory data of the user, analyzing the learning theme and progress change of the user at different time periods to obtain a historical learning trajectory vector; performing fusion processing on the current voice query features and the historical learning trajectory vector to construct conversation state data containing the current query intent and the historical learning context, and obtaining an educational conversation state; wherein the conversation state data comprises current query information, relevant historical learning records, context association weights, and time decay factors, the educational conversation state fuses multiple historical learning records into a unified context representation through weighted summation, and the weight in the weighted summation is determined by the product of the context association weight and the time decay factor; inputting the educational conversation state into the educational knowledge graph neural network for semantic encoding, generating a vector representation fused with knowledge association through the message passing mechanism of the graph neural network, and obtaining a query representation vector; The matching module is configured to perform semantic matching calculation on the query representation vector and the knowledge point-course-discipline three-level semantic tree, perform cognitive level adaptation in combination with the user cognitive profile, and filter to obtain candidate educational content. The evaluation module is configured to perform cross-episode learning context coherence evaluation on the candidate educational content based on the educational conversation state to generate an IPTV educational content intelligent navigation result.

8. A voice recognition based IPTV educational content intelligent navigation apparatus, characterized by, The computer program causes the processor to perform the voice recognition-based IPTV educational content intelligent navigation method according to any one of claims 1 to 6 when the computer program is run on the processor.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program causes the processor to perform the voice recognition-based IPTV educational content intelligent navigation method according to any one of claims 1 to 6 when the computer program is run on the processor.

Citation Information

Patent Citations

  • Online classroom intelligent recommendation method for education robot

    CN120067442A

  • Online teaching interaction method based on multi-modal knowledge graph, medium and equipment

    CN120339011A