A Human-Computer Interaction Method Based on the Fusion of Generative Artificial Intelligence and Multimodality

By integrating sparse adaptive networks and graph convolutional networks combined with reinforcement learning methods, the integration of multimodal education data and personalized learning path generation problems are solved, the accuracy and flexibility of educational content are improved, and the teacher-student interaction and resource allocation in higher education are enhanced.

CN119961524BActive Publication Date: 2025-07-22CIVIL AVIATION UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429561.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively integrate and process multimodal education data, the accuracy of the generated content and the delay in responding to personalized learning needs, resulting in insufficient teacher-student interaction and uneven allocation of educational resources in higher education.

Method used

The integrated sparse adaptive network is used to extract and fusion multimodal data, combine graph convolution networks and reinforcement learning, dynamically adjust the recommended paths, generate personalized learning paths, and optimize educational content in real time.

Benefits of technology

It improves the accuracy and flexibility of generative artificial intelligence to generate educational content, enhances the processing ability of multimodal data, dynamically adapts to personalized learning needs, and improves the adaptability and specificity of teaching content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961524B_ABST
    Figure CN119961524B_ABST
Patent Text Reader

Abstract

The present invention discloses a human-computer interaction method based on generative artificial intelligence and multimodal fusion, belonging to the field of artificial intelligence technology. The method includes acquiring multimodal data and preprocessing the data; using an integrated sparse adaptive network to extract and fuse features of the preprocessed multimodal data; based on the fused multimodal data, using generative artificial intelligence to generate initial learning materials; constructing a knowledge graph, and the generative artificial intelligence generates a personalized initial path based on the multi-dimensional knowledge graph according to the personalized needs of learners; using a graph convolutional network and reinforcement learning to dynamically adjust the path; collecting dynamic feedback information of learners, adjusting the path weights according to the dynamic feedback information of learners, and optimizing the recommended path in real time. By adopting the above method, the present invention enhances the ability to cope with the complexity of multimodal data and improves the accuracy and flexibility of generative artificial intelligence in generating educational content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a human-computer interaction method based on generative artificial intelligence and multimodal fusion. Background Art

[0002] With the rapid development of information technology, the integration of artificial intelligence (AI) and education has deepened, and AI is becoming the core driving force for the modernization and intelligent transformation of the education system. In the field of higher education, traditional teaching methods are usually difficult to promote sufficient teacher-student interaction, resulting in low student participation and unmet needs for personalized learning experience. The problem of uneven distribution of educational resources also limits the access to high-quality educational opportunities, further widening the gap in educational equity. Although online education has brought flexibility to the learning environment, it has also exposed deficiencies in teaching interaction and learning outcome evaluation. To solve the above problems, using artificial intelligence (AI) to enhance human-computer interaction, optimize the allocation of educational resources, and achieve personalized teaching has become one of the key strategies for reform. The progress of generative AI technology has shown revolutionary potential in fields such as natural language processing, image synthesis, and audio generation. These technologies can create high-quality and adaptable content and provide a certain degree of personalized services based on user feedback. Despite this, there are still many major challenges when applying generative AI to educational scenarios. For example, how to effectively integrate and process multimodal educational data, ensure the accuracy and relevance of generated content, and quickly respond to personalized learning needs are all important problems that need to be solved.

[0003] Multimodal fusion technology provides innovative solutions to challenges in higher education by integrating multiple data types such as text, audio, and vision. It enables artificial intelligence systems to understand educational contexts more comprehensively, thereby significantly improving the diversity and accuracy of content generation. However, despite these advantages, the inherent complexity of multimodal data, the delay in system responsiveness, and the continuous changes in personalized learning needs have become major obstacles to the development of current models. Summary of the invention

[0004] The purpose of the present invention is to provide a human-computer interaction method based on generative artificial intelligence and multimodal fusion, which uses an integrated sparse adaptive network to perform feature fusion on multimodal data, and uses a graph convolutional network and reinforcement learning to dynamically adjust the recommended path. While dynamically adapting to personalized learning needs, it enhances the ability to cope with the complexity of multimodal data and improves the accuracy and flexibility of generative artificial intelligence in generating educational content.

[0005] To achieve the above object, the present invention provides a human-computer interaction method based on generative artificial intelligence and multimodal fusion, the steps comprising:

[0006] S1. Obtain text, audio, and visual multimodal data, and preprocess the multimodal data;

[0007] S2. Use an integrated sparse adaptive network to extract and fuse features from the preprocessed multimodal data;

[0008] S3. Based on the fused multimodal data in step S2, use generative artificial intelligence to generate initial learning materials and form a dynamic resource library;

[0009] S4. Construct a multi-dimensional knowledge graph including course relationships, knowledge point associations, and learner individual data. The generative artificial intelligence generates a personalized initial path based on the multi-dimensional knowledge graph according to the personalized needs of the learner;

[0010] S5. Based on the initial path in step S4, use a graph convolutional network to embed the structured content of the multi-dimensional knowledge graph, extract feature vectors, and use reinforcement learning to iteratively adjust the feature vectors to generate an optimal learning path;

[0011] S6. Collect dynamic feedback information of learners, adjust the path weights according to the dynamic feedback information of learners, and optimize the recommended path in real time.

[0012] Preferably, step S2 specifically includes: The integrated sparse adaptive network screens key feature vectors in the multimodal data through an adaptive threshold mechanism. For text features, BERT is used to extract semantic embeddings, and high-frequency keywords are retained through sparsification processing. For audio features, MFCC is used for feature extraction, and key segments are screened through an attention mechanism. For visual features, CNN is used to extract image features, and significant regions are retained through sparse coding. Then, the sparsified feature vectors are concatenated into a joint representation and input into a fully connected layer for deep modeling.

[0013] Preferably, the initial learning materials in step S3 include text explanations, charts, and audio.

[0014] Preferably, constructing the multi-dimensional knowledge graph including course relationships, knowledge point associations, and learner individual data in step S4 includes:

[0015] The knowledge graph includes courses V, teachers Teacher, providers Provider, knowledge points Kpoint, and exercise entities Exercise, and its expression is:

[0016] ;

[0017] Three entities are interconnected through the relationship Relation, and the expression is:

[0018] ;

[0019] Among them, the relation set Relation includes the prerequisite relation (Is_pre_to) between courses, the inclusion relation (v_contain_p) between courses and knowledge points, and the enrollment relation (enrolled_in) between learners and courses;

[0020] The relationship between each entity is represented as a triple, and the expression is:

[0021] ;

[0022] Among them, r represents the relationship, represents the entity set, and represent two entities;

[0023] The knowledge graph includes the knowledge graph at the course level and the knowledge graph at the knowledge point level. The knowledge graph at the course level and the knowledge graph at the knowledge point level are associated to form a structured knowledge graph.

[0024] Preferably, in step S4, the generative artificial intelligence generates a personalized initial path based on the multi-dimensional knowledge graph according to the personalized needs of the learner, including:

[0025] Input the learning goal and current knowledge level of the learner, and let represent the target course of the learner, represent the current knowledge base of the learner. Assume that the learner has mastered all the prerequisite courses and courses included in the knowledge base. The generative artificial intelligence filters according to the prerequisite relationship between courses, and the remaining course set constitutes the starting point of the learning path;

[0026] Introduce the concept of an adaptive learning path, define the relationship constraint R, build the unique connection of adjacent course nodes in the learning path according to the constraint relationship, and recommend a series of related courses according to the specific target category of the learner;

[0027] Score the importance of each course, evaluate it through three dimensions: course score, course difficulty, and course node centrality. Introduce the k-means clustering algorithm to divide the learners into three levels, and the generative artificial intelligence generates a personalized initial path according to the learner level.

[0028] Preferably, in step S5, use the graph convolutional network to embed the structured content of the multi-dimensional knowledge graph and extract the feature vector, including: Initialize a feature vector for each node according to the input stage of the GCN encoder, generate text embeddings for courses and knowledge points using a pre-trained language model, generate image embeddings for learning resources using a convolutional neural network, and use multi-layer graph convolutional operations to aggregate the data of adjacent nodes and gradually update the node features. The expression of the graph convolutional operation for each layer is:

[0029] ;

[0030] In the formula, represents the node feature matrix of the l-th layer; represents the adjacency matrix with self-loops added; represents the degree matrix; represents the weight matrix of the l-th layer; represents the activation function.

[0031] Preferably, reinforcement learning is used to iteratively adjust the feature vectors to generate an optimal learning path, including: the reinforcement learning state space S, the action set A, the reward mechanism R, and the policy function π. Use the algorithm to iteratively optimize the value table to approximate the optimal policy, The value represents the expected cumulative reward obtained by executing a specific action starting from a given state. The iterative optimization update formula is expressed as:

[0032] ;

[0033] In the formula, represents the current state; represents the current action; represents the immediate reward; is the learning rate; is the discount factor; represents the subsequent state; represents the action taken in the next state;

[0034] Use the Markov decision process to build a basic model for the interaction between the agent and the environment to optimize the long-term reward accumulation. By optimizing the agent, an optimal learning path is generated.

[0035] Therefore, the present invention adopts the above-mentioned human-computer interaction method based on generative artificial intelligence and multimodal fusion, and has the following beneficial effects:

[0036] (1) Using an integrated sparse adaptive network to fuse multimodal data, efficiently integrating and deeply modeling diverse educational data, not only enhancing the ability of generative artificial intelligence to process and generate various data modalities in an educational context, but also enhancing the adaptability and specificity of teaching content creation;

[0037] (2) Using a graph convolutional network to extract intricate relationship features from complex educational graphs, generating a personalized recommendation list based on multi-dimensional data, and reinforcement learning iteratively improves these recommendations by incorporating feedback on the learner's progress, ensuring that the learning plan dynamically adapts to changing needs in real time.

[0038] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0039] Figure 1 It is a method framework diagram of an embodiment of the present invention;

[0040] Figure 2 It is a knowledge graph structure diagram of an embodiment of the present invention;

[0041] Figure 3 It is a personalized initial path generation structure diagram of an embodiment of the present invention;

[0042] Figure 4 It is a comparison chart of AUC values between different algorithms of an embodiment of the present invention;

[0043] Figure 5 It is a comparison chart of F1 values between different algorithms of an embodiment of the present invention;

[0044] Figure 6 It is a comparison chart of Top-N values of different algorithms of an embodiment of the present invention on the EdTechX dataset;

[0045] Figure 7 It is a comparison chart of Top-N values of different algorithms of an embodiment of the present invention on the Learning Equality dataset;

[0046] Figure 8 It is a comparison chart of Top-N values of different algorithms of an embodiment of the present invention on the X5gon dataset. Detailed Embodiments

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0048] Embodiment

[0049] The present invention provides a human-computer interaction method based on generative artificial intelligence and multimodal fusion. The method framework diagram is as Figure 1 shown, and the steps include:

[0050] S1. Obtain multi-modal heterogeneous data such as text (course descriptions, student feedback), audio (lecture recordings, voice interactions), and vision (charts, videos), and preprocess the multi-modal data, including normalization, denoising, and aligning timestamps (for audio-video and text synchronization).

[0051] S2. Use an integrated sparse adaptive network to extract and fuse features from the preprocessed multi-modal data, specifically including:

[0052] The integrated sparse adaptive network screens key feature vectors in the multi-modal data through an adaptive threshold mechanism. For text features, BERT is used to extract semantic embeddings, and sparsification processing retains high-frequency keywords. For audio features, MFCC is used for feature extraction, and key segments are screened through an attention mechanism. For vision features, CNN is used to extract image features, and sparse coding retains significant regions. Then, the sparsified feature vectors are concatenated into a joint representation and input into a fully connected layer for deep modeling.

[0053] S3. Based on the fused multi-modal data in step S2, use generative artificial intelligence to generate initial learning materials, including text explanations, charts, and audio, such as well-structured explanatory texts, charts, and audio, etc., and form a dynamic resource library.

[0054] S4. Construct a multi-dimensional knowledge graph including course relationships, knowledge point associations, and learner individual data. The generative artificial intelligence generates a personalized initial path based on the multi-dimensional knowledge graph according to the personalized needs of the learner.

[0055] The knowledge graph is a basic framework representing entities such as courses, knowledge points, learners, learning tasks, etc., and the relationships between them. For example, the "contains" relationship may connect a course with a specific knowledge point, while the "takes" relationship connects a student with a course. Through the above structured representation, the contextualization and organization necessary for effective learning path recommendation are achieved. In a teaching context, the knowledge graph operates as a dynamic and evolving system rather than a static knowledge base, and these graphs are continuously updated to incorporate new academic viewpoints, learning resources, and behavioral data generated through learner interactions. Formally, the knowledge graph is represented as a set of triples:

[0056] ;

[0057] where: represents the head entity, the initiator of the relationship; represents the relationship connecting the two entities; represents the tail entity, the recipient of the relationship.

[0058] The knowledge graph includes course V, teacher Teacher, provider Provider, knowledge point Kpoint, and exercise entity Exercise, and its expression is:

[0059] 。

[0060] Three entities are interconnected through the relation Relation, and the expression is:

[0061] ;

[0062] Among them, the relation set Relation includes the prerequisite relation (Is_pre_to) between courses, the inclusion relation (v_contain_p) between courses and knowledge points, and the enrollment relation (enrolled_in) between learners and courses.

[0063] The relationship between each entity is represented as a triple, and the expression is:

[0064] ;

[0065] Among them, represents the entity set, and represent two entities.

[0066] In the learning context based on courses, the knowledge graph includes the knowledge graph at the course level (MCCKG) and the knowledge graph at the knowledge point level (KPKG). As Figure 2 shown, the knowledge graph at the course level and the knowledge graph at the knowledge point level are associated to form a structured knowledge graph.

[0067] In a fragmented and unorganized online course environment, learners face considerable difficulties in system navigation and achieving learning goals. The generative artificial intelligence, based on the learners' learning goals and current knowledge levels, utilizes the relationships between courses in the constructed knowledge graph, and comprehensively considers factors such as interests, learning preferences, and course characteristics to generate a set of comprehensive potential learning paths and customize the most suitable path for each learner. The structure is as Figure 3 shown, specifically:

[0068] Input the learners' learning goals and current knowledge levels, and let represent the learners' target courses, represent the learners' current knowledge base. Assume that the learners have mastered all the prerequisite courses and courses included in the knowledge base. Based on the above assumptions, if (such as "Is_pre_to" or "Is_used_to_Al") there is a prerequisite relationship between the courses, then these prerequisite courses are excluded from In addition, similarly, if there is an inclusion relationship, the included courses will also be excluded. After screening, the remaining course set constitutes the starting point of the learning path, and any course within this set may become the starting point of the personalized learning journey. Among them, the field "Is_pre_to" is used to represent the prerequisite relationship between courses, mainly to clarify the sequential dependence relationship between courses to assist in teaching arrangements and students' course selection and other tasks. For example, if the value of the "Is_pre_to" field of Course B is Course A, it means that Course A is a prerequisite for Course B, and students need to complete the study of Course A first before they can study Course B. "Is_used_to_AI" is used to indicate whether a certain course is used in the teaching related to artificial intelligence. For example, if the value of the "Is_used_to_Al" field of a certain course is "True", it means that this course is used in the teaching of the artificial intelligence direction and may involve relevant contents such as artificial intelligence algorithms and models; if the value is "False", it means that this course has nothing to do with the use of artificial intelligence.

[0069] Introduce the concept of an adaptive learning path, define the relationship constraint R, build the unique connection of adjacent course nodes in the learning path according to the constraint relationship, and recommend a series of relevant courses according to the specific target category of the learner.

[0070] Evaluate the importance of each course, evaluate it through three dimensions: course score (ratLo), course difficulty (diffLoi), and the centrality of the course node (centLoi). Introduce the k-means clustering algorithm to divide learners into three levels: beginners, intermediate learners, and advanced learners. The generative artificial intelligence generates a personalized initial path according to the learner level. The classification shows that beginners tend to prefer simpler courses, while advanced learners are more likely to engage in challenging courses. Therefore, sort the courses according to the learner type by difficulty, and the corresponding difficulty scores are shown in Table 1. For beginners, the easy-to-learn courses get the lowest difficulty scores, the normal courses are assigned to the next level, and the difficult-to-learn courses get the highest scores.

[0071] Table 1 Course Difficulty Levels

[0072] ;

[0073] S5. Based on the initial path in step S4, use the graph convolutional network to embed the structured content of the multi-dimensional knowledge graph, extract the feature vectors, and use reinforcement learning to iteratively adjust the feature vectors to generate the optimal learning path. Specifically:

[0074] The graph convolutional network initializes a feature vector for each node according to the input stage of the GCN encoder, uses a pre-trained language model to generate text embeddings for courses and knowledge points, uses a convolutional neural network to generate image embeddings for learning resources, and adopts multi-layer graph convolutional operations to aggregate data of adjacent nodes and gradually update node features. The expression of the graph convolutional operation for each layer is as follows:

[0075] ;

[0076] In the formula, represents the node feature matrix of the l-th layer; represents the adjacency matrix with self-loops added; represents the degree matrix; represents the weight matrix of the l-th layer; represents the activation function.

[0077] For the state space S, action set A, reward mechanism R, and policy function π of reinforcement learning, use algorithm to iteratively optimize value table to approximate the optimal policy. The value represents the expected cumulative reward obtained by executing a specific action starting from a given state. The iterative optimization update formula is expressed as:

[0078] ;

[0079] In the formula, represents the current state; represents the current action; represents the immediate reward; is the learning rate; is the discount factor; represents the subsequent state; represents the action taken in the next state.

[0080] Use the Markov decision process to build a basic model for the interaction between the agent and the environment to optimize long-term reward accumulation. By optimizing the agent, an optimal learning path is generated. In the Markov decision process, the task of the "agent" is to derive a policy (π(s)), which maps the "state" to an action. The policy (π(s)) defines the action that the agent should take when in the state (s). Its goal is to find an optimal policy , maximizing the total "reward" accumulated over time. In the field of learning path recommendation, the learner plays the role of the "agent", the courses or knowledge points represent the "states" in the environment, and the actions correspond to the choices made by the learner, such as choosing a certain knowledge point to learn. The "reward" reflects the learner's progress or satisfaction, guiding the system to continuously adjust the recommendation strategy. By leveraging reinforcement learning, the recommendation content can be iteratively adjusted to ensure that the personalized learning path is consistent with the learner's changing needs and feedback.

[0081] In a recommendation system based on graph convolutional networks, the nodes represent entities such as learners, courses, and knowledge points, while the edges describe the relationships between these entities, such as "learner A has completed course B" or "course B contains knowledge point C". Information from neighboring nodes is aggregated through iterative multi-layer graph convolutional operations to refine the feature representation of each node. This process can generate comprehensive node embeddings, facilitating the recommendation of highly relevant courses or knowledge points to learners. In a teaching context, graph convolutional networks go beyond mapping student-course interactions and reveal deeper interdependencies between courses. For example, the prerequisites of a course, represented by the "prerequisite" relationship, outline the basic knowledge required before progressing to advanced topics. By leveraging the propagation mechanism of GCNs, the recommendation system can automatically derive successive course requirements or propose optimal learning paths, thereby improving educational efficiency and outcomes. The combination of graph convolutional networks and reinforcement learning further enhances the adaptability and intelligence of the recommendation system by combining the advantages of relational data extraction and dynamic policy optimization.

[0082] S6. Collect dynamic feedback information of the learner, adjust the path weights according to the dynamic feedback information of the learner, and optimize the recommended path in real time.

[0083] In addition to creating personalized learning materials that are consistent with the course objectives and tailored to the individual needs of students, generative artificial intelligence can also generate various teaching materials that can assist teachers according to the course syllabus and teaching objectives, such as experimental reports, case analyses, and discussion prompts, enriching the teaching experience and providing targeted learning resources for students. For visual teaching, generative AI can automatically generate images, animated charts, and videos that are consistent with the course content.

[0084] The present invention also provides a human-computer interaction system, applying the above human-computer interaction method, including account login, new account registration, resource search within computer science, personalized course and knowledge point recommendation, and user evaluation participation.

[0085] To evaluate the effectiveness of the graph convolutional network and reinforcement learning algorithm (GCN+RL) in terms of personalized recommendation paths for learners, a comparative experiment was conducted, comparing this method with baseline methods such as GCN, Knowledge Graph Convolutional Network (KGCN), and Knowledge Graph Attention Network (KGAT). To facilitate model training and performance evaluation, publicly available educational datasets were used: the EdTechX educational dataset, the Kaggle educational equity course dataset, and the X5gon open-source educational platform dataset, and a sparse adaptive network was used for efficient feature fusion to help generative artificial intelligence manage complex multimodal data and ensure accurate feature extraction.

[0086] The study selected 220 students from different disciplines as experimental subjects to participate in the experiment. These students who had completed at least one course were evenly divided into two groups: the experimental group and the control group. The experimental group used an optimized generative artificial intelligence model (RL+GCN) for personalized learning path recommendation, and the control group used a traditional content-based algorithm for learning path recommendation. The accuracy of personalized recommendation was evaluated using two metrics: click-through rate (CTR) and retention rate (RR), and the calculation formulas are as follows:

[0087] 。

[0088] 。

[0089] The AUC and F1 values of different learning path recommendation algorithms on 3 datasets are as Figure 4 and Figure 5 shown. Through Figure 4 and Figure 5 it can be seen that the RL+GCN algorithm is superior to other algorithms and obtained the highest AUC and F1 values on the EdTechX dataset, which are 0.915 and 0.838 respectively.

[0090] In practical performance evaluation, the average CTR of the experimental group was 28.5% (±3.2%), higher than 20.7% (±3.8%) of the control group. Similarly, the retention rate of the experimental group reached 72.1% (±4.5%), while that of the control group was 61.3% (±5.2%). Generally speaking, the accuracy of personalized recommendations in the experimental group increased by 37.9% compared with the control group. The degree of coincidence between the personalized learning paths generated by the RL+GCN model and the optimal learning paths increased to 68.5%, an increase of 46.5% compared with the baseline. In the simulated dynamic learning demand scenario, the average response time of the model decreased to 0.3 seconds, and the response speed increased by 44.1%. The efficient feature fusion mechanism of the sparse adaptive network helps to more accurately identify the interests and needs of learners, thus generating more relevant recommendations. Compared with traditional knowledge graph-based recommendation algorithms, the sparse adaptive network can better manage the complex relationships in multi-modal data and improve the accuracy and relevance of recommendations.

[0091] The Top-N recommendation results of different algorithms are as Figure 6 , Figure 7 and Figure 8 shown. As the value of K increases, the recall rate rises steadily. It can be seen from the figure that the RL+GCN model is superior to other baseline methods in terms of recall rate, demonstrating its excellent comprehensive performance. The model simplifies the non-linear activation and feature transformation steps in the graph convolution process, significantly reducing the training complexity.

[0092] Therefore, the present invention adopts the above-mentioned human-computer interaction method based on generative artificial intelligence and multi-modal fusion, uses an integrated sparse adaptive network to perform feature fusion on multi-modal data, and uses a graph convolution network and reinforcement learning to dynamically adjust the recommendation path. While dynamically adapting to personalized learning needs, it enhances the ability to handle the complexity of multi-modal data and improves the accuracy and flexibility of the educational content generated by generative artificial intelligence.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A human-computer interaction method based on the integration of generative artificial intelligence and multi-modal, characterized in that the steps Including: S1. Obtain text, audio, and visual multimodal data, and preprocess the multimodal data; S2. Use an integrated sparse adaptive network to extract and fuse features from the preprocessed multimodal data, including: The integrated sparse adaptive network screens key feature vectors in the multimodal data through an adaptive threshold mechanism. For text features, BERT is used to extract semantic embeddings, and sparsification processing retains high-frequency keywords. For audio features, MFCC is used for feature extraction, and key segments are screened through an attention mechanism. For visual features, CNN is used to extract image features, and sparse coding retains significant regions. Then, the sparsified feature vectors are concatenated into a joint representation and input into a fully connected layer for deep modeling; S3. Based on the fused multimodal data in step S2, use generative artificial intelligence to generate initial learning materials and form a dynamic resource library; S4. Construct a multi-dimensional knowledge graph including course relationships, knowledge point associations, and learner individual data. The generative artificial intelligence generates a personalized initial path based on the multi-dimensional knowledge graph according to the personalized needs of the learner; S5. Based on the initial path in step S4, use a graph convolutional network to embed the structured content of the multi-dimensional knowledge graph, extract feature vectors, and use reinforcement learning to iteratively adjust the feature vectors to generate an optimal learning path, where the feature vectors represent courses or knowledge points; Using reinforcement learning to iteratively adjust the feature vector to generate an optimal learning path includes: setting the reinforcement learning state space S, action set A, reward mechanism R, and policy function , using the Q-Learning algorithm to iteratively optimize the Q-value table to approximate the optimal policy. The Q-value represents the expected cumulative reward obtained by executing a specific action starting from a given state. The iterative optimization update formula is expressed as: ; In the formula, represents the current state; represents the current action; represents the immediate reward; is the learning rate; is the discount factor; represents the subsequent state; represents the action taken in the next state; Use the Markov decision process to construct a basic model for optimizing long-term reward accumulation through the interaction between the agent and the environment. The task of the agent is to derive a policy that maps states to actions. The policy defines the actions that the agent should take in state s. The goal is to find an optimal policy to maximize the total reward accumulated over time. Here, the learner represents the agent, the course or knowledge point represents the state, the action represents the choice made by the learner, and the reward represents the progress or satisfaction of the learner. The learner's goal is optimized through reinforcement learning to obtain the optimal path; S6. Collect dynamic feedback information of the learner, adjust the path weights according to the dynamic feedback information of the learner, and optimize the recommended path in real time.

2. The human-computer interaction method based on the fusion of generative artificial intelligence and multi-modal according to claim 1, characterized in that: The initial learning materials in step S3 include text explanations, charts, and audio.

3. A human-computer interaction method based on the integration of generative artificial intelligence and multi-modalities according to claim 1, characterized in that, The construction of the multi-dimensional knowledge graph including course relationships, knowledge point associations, and learner individual data in step S4 includes: The knowledge graph includes courses V, teachers Teacher, providers Provider, knowledge points Kpoint, and exercise entities Exercise, and its expression is: ; Three entities are connected to each other through the relationship Relation, and the expression is: ; Among them, the relationship set Relation includes the prerequisite relationship between courses, the inclusion relationship between courses and knowledge points, and the study relationship between learners and courses; The relationship between each entity is represented as a triple, and the expression is: ; Among them, represents a relationship, represents a set of entities, and represents two entities; The knowledge graph includes a knowledge graph at the course level and a knowledge graph at the knowledge point level. The knowledge graph at the course level and the knowledge graph at the knowledge point level are associated to form a structured knowledge graph.

4. A human-computer interaction method based on the integration of generative artificial intelligence and multi-modal, as claimed in claim 3, wherein, The generative artificial intelligence generates a personalized initial path based on the multi-dimensional knowledge graph according to the personalized needs of the learner in step S4 includes: Input the learning goals and current knowledge level of the learner, and let represent the target courses of the learner, represent the learner's current knowledge base. Assume that the learner has mastered all the prerequisite courses and courses included in the knowledge base. The generative artificial intelligence filters according to the prerequisite relationships between the courses, and the remaining course set forms the starting point of the learning path; Introduce the concept of an adaptive learning path, define the relationship constraint R, build the unique connection of adjacent course nodes in the learning path according to the constraint relationship, and recommend a series of related courses according to the specific target category of the learner; Score the importance of each course, evaluate it through three dimensions: course score, course difficulty, and course node centrality, introduce the k-means clustering algorithm, divide the learners into three levels, and the generative artificial intelligence generates a personalized initial path according to the learner level.

5. The human-computer interaction method based on the fusion of generative artificial intelligence and multi-modalities according to claim 4, characterized in that, In step S5, a graph convolutional network is used to embed the structured content of the multi-dimensional knowledge graph. The extracted feature vectors include: initializing a feature vector for each node according to the input stage of the GCN encoder, generating text embeddings for courses and knowledge points using a pre-trained language model, generating image embeddings for learning resources using a convolutional neural network, aggregating data of adjacent nodes using multi-layer graph convolutional operations, and gradually updating the node features. The expression of the graph convolutional operation for each layer is: ; In the formula, represents the node feature matrix of the l-th layer; represents the adjacency matrix with self-loops added; represents the degree matrix; represents the weight matrix of the l-th layer; represents the activation function.

Citation Information

Patent Citations

  • Knowledge graph personalized learning path recommendation method based on RankNet-transformer

    CN113239209A

  • A teaching optimization method based on big data informationization

    CN119741175A