Multi-mode interactive learning and consulting integrated platform

Through a multi-layered architecture consisting of an edge interaction layer, an edge node collaboration layer, and a cloud-enabled layer, the problem of transmission lag, data loss, and privacy leakage in multimodal learning platforms under weak network conditions and on low-end devices has been solved, enabling efficient, accurate, and privacy-protected learning and consulting services in weak network environments.

CN122069276APending Publication Date: 2026-05-19SHANDONG TRANSPORT VOCATIONAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610140148.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-02
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing multimodal interactive learning platforms suffer from problems such as transmission lag, data loss, privacy leaks, and rapid power consumption on weak networks and low-end devices, making them difficult to adapt to the needs of complex learning scenarios.

Method used

It adopts a three-layer architecture consisting of an edge interaction layer, an edge node collaboration layer, and a cloud-enabled layer. Through multimodal data acquisition, lightweight processing, dynamic decision scheduling, lightweight federated learning, and adaptive energy consumption optimization, it achieves localized data processing and privacy protection.

Benefits of technology

Ensuring the smoothness and accuracy of multimodal learning consultation in weak network environments reduces reliance on cloud computing power and network bandwidth, improves data extraction accuracy and privacy protection, and extends device battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069276A_ABST
    Figure CN122069276A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode interactive learning and consulting integrated platform which comprises an edge end interaction layer, an edge node collaboration layer and a cloud end enabling layer. The edge end interaction layer is composed of a multi-modal data acquisition module, an intelligent state sensing module, a core data processing module, a local learning caching module, an energy consumption self-adaption module and a dynamic decision scheduling module, and is used for acquisition, state sensing, precise processing, local caching, energy consumption optimization and transmission priority scheduling of multi-modal interaction data; the edge node collaboration layer is composed of a collaboration transmission relay module, a regional resource cache pool and a lightweight federated node module, and is used for realizing data temporary storage, incremental relay transmission, localized learning resource support and regional model aggregation; and the cloud enabling layer is provided with a lightweight learning knowledge base, a federated learning coordinator and a global strategy optimization module, and is used for providing learning resources of different learning sections and different subjects, planning the whole process of federated learning as a whole and optimizing a global operation strategy of the platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer application technology, and in particular to a multimodal interactive learning and consulting integrated platform. Background Technology

[0002] Against the backdrop of the deepening digital transformation of education, multimodal interactive learning and consultation platforms have become a core carrier for breaking through the time and space limitations of traditional education and realizing personalized learning. Integrating multimodal interaction methods such as text, voice, and handwriting, combined with intelligent Q&A and resource linkage functions, they are reshaping the ecological landscape of teaching and learning. Currently, mainstream platforms are introducing technologies such as federated learning and edge computing to attempt to resolve the contradiction between data privacy protection and resource sharing efficiency, and to promote the upgrading of educational services towards precision and convenience.

[0003] However, existing platforms still face multiple technical bottlenecks in practical implementation, making it difficult to adapt to the needs of complex learning scenarios:

[0004] Rural campuses and corners of campuses often suffer from bandwidth of less than 1Mbps and frequent signal fluctuations. Multimodal data (voice, handwritten formulas, etc.) transmission is prone to stuttering and interruption. Existing solutions either rely on strong cloud network support or simply compress data, resulting in semantic loss and failing to guarantee the continuity of consultation.

[0005] Most platforms simply pile up modal functions, lacking lightweight collaborative processing strategies. This leads to a contradiction between redundant data transmission and insufficient accuracy in core semantic extraction. Furthermore, cross-modal interactions are prone to conflicts, making it difficult to adapt to the precise needs of learning and consulting.

[0006] Traditional centralized model training requires the collection of sensitive data such as students' handwritten notes and voice consultations, which poses a high risk of privacy leakage. Although some platforms have introduced federated learning, the models are large and rely on cloud computing power, making it impossible to run efficiently on low-end learning devices.

[0007] Students often use low-end tablets and smartphones with limited battery life. Existing platforms lack dynamic power consumption adjustment mechanisms, and running them at full power under weak network conditions can easily lead to rapid power consumption, affecting the continuity of learning.

[0008] Therefore, a multimodal interactive learning and consulting integrated platform is proposed. Summary of the Invention

[0009] This application aims to at least partially solve one of the technical problems in the aforementioned technologies.

[0010] To achieve the above objectives, the first aspect of this application proposes a multimodal interactive learning and consulting integrated platform, comprising an edge interaction layer, an edge node collaboration layer, and a cloud-enabled layer, wherein:

[0011] The edge interaction layer consists of a multimodal data acquisition module, an intelligent state perception module, a core data processing module, a local learning cache module, an energy consumption adaptive module, and a dynamic decision scheduling module. It is used for the acquisition, state perception, accurate processing, local caching, energy consumption optimization, and transmission priority scheduling of multimodal interaction data.

[0012] The edge node collaboration layer consists of a collaborative transmission relay module, a regional resource cache pool, and a lightweight federated node module, which are used to realize data temporary storage, incremental relay transmission, local learning resource support, and regional model aggregation.

[0013] The cloud-enabled layer includes a lightweight learning knowledge base, a federated learning coordinator, and a global strategy optimization module, which are used to provide learning resources by grade level and subject, coordinate the entire federated learning process, and optimize the platform's global operation strategy.

[0014] In addition, the multimodal interactive learning and consulting integrated platform proposed in this application may also have the following additional technical features:

[0015] As a further description of the above technical solution:

[0016] The multimodal data acquisition module supports the synchronous acquisition and collaborative interaction of three core modal data types: text, voice, and handwriting. The intelligent state perception module collects multi-dimensional state information locally at the edge through a lightweight sensor fusion algorithm with an acquisition latency of ≤50ms. The multi-dimensional state information specifically includes network status, device status, and learning consultation status, providing a basis for subsequent data processing and strategy scheduling.

[0017] As a further description of the above technical solution:

[0018] The core data processing module adopts a modal collaborative processing approach to adapt to the needs of multimodal learning and consultation. It performs TF-Lite keyword extraction, education domain vocabulary filtering, and state weighting optimization on text data, endpoint detection, acoustic model recognition of educational terms, and adaptive speech rate segmentation on speech data, and lightweight CNN recognition, stroke priority sorting, and test point association extraction on handwritten data. The overall core data extraction accuracy rate is ≥96%.

[0019] As a further description of the above technical solution:

[0020] The dynamic decision-making and scheduling module adjusts the priority of multimodal data transmission based on multi-dimensional state weight factors, with the weight allocation being 0.4 for network quality, 0.3 for the urgency of the learning scenario, and 0.3 for user interaction preferences.

[0021] The dynamic decision-making and scheduling module works in conjunction with the transmission relay module to achieve multi-protocol fusion transmission.

[0022] As a further description of the above technical solution:

[0023] The regional resource cache pool stores high-frequency learning and consulting resources by grade level and subject, including formula templates, Q&A scripts for knowledge points, case studies of incorrect questions, and scenario rule templates, through a two-way linkage mechanism of local caching and edge caching.

[0024] As a further description of the above technical solution:

[0025] The lightweight federated node module and the cloud-based federated learning coordinator work together to build a lightweight federated learning framework. Participants include edge devices and edge nodes. The cloud does not participate in the transmission of original learning data to ensure privacy. Each edge device optimizes the semantic completion model based on local learning consultation interaction data.

[0026] As a further description of the above technical solution:

[0027] The semantic keep-alive completion module adopts a three-level completion strategy. The basic layer achieves fast matching based on the scenario rule templates shared by local and edge nodes. The enhancement layer achieves global semantic adaptation through a word vector model optimized by federated learning. The bottom layer combines user historical learning consultation data with collaborative recommendations from edge nodes to achieve context-related completion.

[0028] As a further description of the above technical solution:

[0029] The local learning cache module adopts a multi-dimensional weighted caching strategy based on the LRU algorithm, adding energy consumption weight and learning priority weight. When the device is in a low power state, low priority cached content is automatically cleared, and high-frequency core consultation data and resources related to users' weak knowledge points are retained first.

[0030] As a further description of the above technical solution:

[0031] The energy consumption adaptive module dynamically adjusts the platform's operating mode based on the device's power consumption and network status. In low power and weak network scenarios, it automatically shuts down non-core modality recognition functions and retains only core consulting service capabilities. Once the network and device status are restored, it automatically restarts the full-function mode.

[0032] Advantages of this invention:

[0033] The multimodal interactive learning and consulting platform of this application adopts a three-layer architecture of edge-end interactive processing, edge node collaboration, and cloud empowerment. The core logic is pushed down to the edge end and edge nodes, which greatly reduces the dependence on cloud computing power and network bandwidth, perfectly adapts to weak network and low-end learning device scenarios, and avoids the transmission lag problem of traditional pure cloud architecture.

[0034] Multimodal processing and transmission strategies are precisely adapted to educational scenarios. The combination of modal collaborative algorithms enables accurate extraction and lightweight compression of core data. Combined with dynamic priority scheduling and multi-protocol fusion transmission, the extraction accuracy is guaranteed to be ≥96% and the end-to-end latency is <450ms, balancing information integrity and transmission efficiency, provided that the data volume in a single round is <8KB.

[0035] Balancing privacy protection and model optimization, the lightweight federated learning framework only transmits incremental model parameters. Combined with differential privacy technology, it eliminates the leakage of students' sensitive learning data at the source. The global model size is less than 80MB, enabling efficient operation on low-end devices and resolving the contradiction between high privacy risks and difficult model deployment in traditional solutions.

[0036] It possesses strong practicality and scalability, with a consultation success rate of ≥99% in weak network environments. It is adaptable to scenarios across all educational stages and can be horizontally integrated with more subject resources and interactive modes, providing technical support for the balanced coverage of educational resources and forming a differentiated competitive barrier.

[0037] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0038] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0039] Figure 1 This is a schematic diagram of the module connections of a multimodal interactive learning and consulting integrated platform according to an embodiment of this application. Detailed Implementation

[0040] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0041] The multimodal interactive learning and consulting integrated platform of this application embodiment will be described below with reference to the accompanying drawings.

[0042] like Figure 1As shown in Embodiment 1 of this application, the multimodal interactive learning and consulting integrated platform may include an edge interaction layer, an edge node collaboration layer, and a cloud empowerment layer. The edge interaction layer consists of a multimodal data acquisition module, an intelligent state perception module, a core data processing module, a local learning cache module, an energy consumption adaptive module, and a dynamic decision scheduling module. It is used for the acquisition, state perception, precise processing, local caching, energy consumption optimization, and transmission priority scheduling of multimodal interactive data. The edge node collaboration layer consists of a collaborative transmission relay module, a regional resource cache pool, and a lightweight federated node module. It is used to realize data temporary storage, data augmentation, and other functions. The solution features relay transmission, localized learning resource support, and regional model aggregation. The cloud-enabled layer includes a lightweight learning knowledge base, a federated learning coordinator, and a global strategy optimization module. These modules provide learning resources by grade level and subject, coordinate the entire federated learning process, and optimize the platform's global operation strategy. In this solution, the edge devices handle interactive data collection, core processing, and energy consumption optimization. Edge nodes act as relay hubs, enabling data storage, regional resource sharing, and model aggregation. The cloud is only responsible for global resource coordination and strategy optimization and does not participate in front-end core computation. This achieves lightweight adaptation for multimodal learning consultation in weak network environments.

[0043] like Figure 1 As shown:

[0044] The multimodal data acquisition module supports the synchronous acquisition and collaborative interaction of three core modalities: text, voice, and handwriting. The intelligent state perception module collects multi-dimensional state information locally at the edge using a lightweight sensor fusion algorithm, with an acquisition latency of ≤50ms. The multi-dimensional state information specifically includes network status, device status, and learning / consultation status, providing a basis for subsequent data processing and strategy scheduling. This solution clearly states that the multimodal data acquisition module supports synchronous interaction of three core modalities: text, voice, and handwriting, and is compatible with multiple terminal devices. The intelligent state perception module collects network, device, and learning / consultation status data locally at the edge using a lightweight sensor fusion algorithm, with the acquisition latency controlled within 50ms, providing real-time and accurate decision-making basis for subsequent data processing and strategy scheduling.

[0045] like Figure 1 As shown:

[0046] The core data processing module employs a modal collaborative processing approach to adapt to the needs of multimodal learning and consultation. For text data, it performs TF-Lite keyword extraction, educational terminology lexicon filtering, and state-weighted optimization. For speech data, it performs endpoint detection, educational terminology acoustic model recognition, and adaptive speech rate segmentation. For handwritten data, it performs lightweight CNN recognition, stroke priority sorting, and test point association extraction. The overall core data extraction accuracy is ≥96%. This solution utilizes a modal collaborative algorithm in the core data processing module to customize processing strategies for the characteristics of text, speech, and handwritten data. Text focuses on keyword and educational terminology extraction; speech filters redundant information while retaining core semantics; and handwritten data recognizes core symbols and test point association logic. Simultaneously, through algorithm optimization, it achieves an extraction accuracy of ≥96% and a lightweight compression effect of <8KB per round of data.

[0047] like Figure 1 As shown:

[0048] The dynamic decision-making and scheduling module adjusts the priority of multimodal data transmission based on multi-dimensional state weight factors. The weight allocation is 0.4 for network quality, 0.3 for learning scenario urgency, and 0.3 for user interaction preference. The dynamic decision-making and scheduling module works in conjunction with the collaborative transmission relay module to achieve multi-protocol fusion transmission. This solution dynamically adjusts the transmission priority of multimodal data based on the weight allocation of 0.4 for network quality, 0.3 for scenario urgency, and 0.3 for user preference. At the same time, it works in conjunction with the collaborative transmission relay module of the edge node and adopts a UDP and TCP multi-protocol fusion transmission strategy. Core data uses UDP to ensure low latency, while incremental data uses lightweight TCP to ensure reliability, ultimately achieving an end-to-end response latency of <450ms.

[0049] like Figure 1 As shown:

[0050] The regional resource cache pool stores high-frequency learning and consulting resources by grade level and subject, including formula templates, Q&A scripts for knowledge points, error analysis cases, and scenario rule templates. It uses a two-way linkage mechanism of local caching and edge caching to store high-frequency learning and consulting resources by grade level and subject. Through the two-way linkage mechanism of local caching and edge caching, the solution enables local retrieval of resources and improves the cache hit rate to ≥85%, supporting users to view consultation records offline and quickly initiate similar questions.

[0051] like Figure 1 As shown:

[0052] The lightweight federated node module and the cloud-based federated learning coordinator collaborate to build a lightweight federated learning framework. Participants include edge devices and edge nodes. The cloud does not participate in the transmission of original learning data to ensure privacy. Each edge device optimizes the semantic completion model based on local learning consultation interaction data. This solution constructs a lightweight federated learning framework, limiting participants to edge devices and edge nodes. The cloud acts only as a coordinator and does not access the original learning data. Each edge device optimizes the model based on local data, uploading only parameter increments of <2KB / round. Edge nodes aggregate to generate regional models, which are then synchronized to the cloud to form a global model of <80MB. Differential privacy technology is used to add noise to the parameter increments to protect user privacy.

[0053] like Figure 1 As shown:

[0054] The semantic keep-alive completion module adopts a three-level completion strategy. The basic layer achieves fast matching based on scene rule templates shared by local and edge nodes. The enhancement layer achieves global semantic adaptation through a word vector model optimized by federated learning. The bottom layer combines user historical learning consultation data with edge node collaborative recommendations to achieve context-related completion. Through the collaboration of the three-level strategy, the solution achieves a completion accuracy of ≥92% for incomplete data and a completion latency of <180ms.

[0055] like Figure 1 As shown:

[0056] The local learning cache module adopts a multi-dimensional weighted caching strategy based on an improved LRU algorithm, adding energy consumption weight and learning priority weight. When the device is in a low power state, low-priority cached content is automatically cleared, while high-frequency core information data and resources related to users' weak knowledge points are prioritized for retention. This solution, through the local learning cache module based on an improved LRU algorithm, adds energy consumption weight and learning priority weight. When the device is in a low power state, low-frequency, unrelated low-priority cached content is automatically cleared, while high-frequency core information data and resources related to users' weak knowledge points are prioritized for retention, balancing caching efficiency and device battery life.

[0057] like Figure 1 As shown:

[0058] The energy consumption adaptive module dynamically adjusts the platform's operating mode based on device power and network status. In low power and weak network scenarios, it automatically shuts down non-core modal recognition functions, retaining only core consultation service capabilities. Once the network and device status are restored, it automatically restarts the full-function mode. This solution's energy consumption adaptive module dynamically switches the platform's operating mode based on device power and network status. In low power and weak network scenarios, it shuts down non-core functions such as redundant handwriting stroke detection and voice emotion analysis, retaining only core consultation services. Once the network and device status are restored, it automatically restarts the full-function mode, achieving an energy consumption reduction rate of ≥30%.

[0059] Example 2: The following is a specific case for further explanation:

[0060] This embodiment focuses on the math learning scenario of eighth graders in rural middle schools, adapting to low-end Android tablets used by students (MediaTek G85 CPU, 4GB RAM). It supports dual-modal learning consultation, including handwritten formulas and voice follow-up questions, even in a weak campus network environment (network speed fluctuations of 300Kbps-800Kbps, average packet loss rate of 3%-5%). The core objective is to achieve accurate Q&A for quadratic function-related questions, while ensuring smooth interaction, accurate answers, and device battery life throughout. The platform adopts a three-tier deployment architecture: edge, edge nodes, and cloud. The edge is deployed on student tablets, integrating a lightweight model package (including a multimodal processing algorithm, a local caching module, and an adaptive power consumption module), running on the TF-Lite 2.10 framework. The edge nodes are small campus edge servers based on Ubuntu. System 22.04 provides collaborative transmission relay, regional resource caching pool, and federated node aggregation services, synchronously caching mathematics resources for different educational stages; the cloud is a regional education cloud server, deploying a lightweight learning knowledge base and a federated learning coordinator, which does not participate in front-end real-time calculations, but is only responsible for global parameter synchronization and resource updates. The implementation process of each core module is as follows:

[0061] Implementation of the intelligent state perception module:

[0062] Relying on the tablet's built-in sensors and lightweight detection software, three types of state data are captured in real time with a 100ms acquisition cycle, and the latency is strictly controlled within 45ms throughout the process. The specific acquisition and calculation methods are as follows:

[0063] Network status probing uses the ICMP lightweight message mechanism, sending one 64B ICMP message every 100ms. After collecting five sets of data, the average value is taken for calculation. The network speed calculation formula is: Network speed (Kbps) = (Message size × 8 × Number of successful receptions) / (Total probing time (ms)) × 1000; The packet loss rate calculation formula is: Packet loss rate (%) = (Number of transmissions - Number of successful receptions) / Number of transmissions × 100.

[0064] Device status is collected via Android system API, including average CPU utilization, remaining battery power (percentage), and storage utilization within 500ms. Remaining computing power is quantified by a coefficient: computing power coefficient = (1 - CPU utilization) × 0.6 + (1 - storage utilization) × 0.4, with a value range of 0-1. A higher value indicates more sufficient computing power for the device.

[0065] The learning consultation status recognition combines text keywords, voice semantics, and handwritten content to match preset scene tags (such as quadratic function, axis of symmetry, Q&A), and simultaneously marks the learning stage (second year of junior high school) and consultation intent (such as seeking analysis), providing scene support for subsequent data processing.

[0066] Implementation of the multimodal core data processing module:

[0067] For students to write quadratic function formulas by hand The platform processes the dual-modal input, including the voice inquiry about how the axis of symmetry is calculated, separately for each modality, and finally completes data fusion, as detailed below:

[0068] The text data (generated from the core semantics of speech-to-text transcription) was processed using the TF-Lite keyword extraction model. Based on an educational vocabulary containing over 2000 mathematical terms, TF-IDF values ​​were calculated to select core vocabulary. The calculation formula is: That is (document) Chinese terminology Number of occurrences / documents (Total word count) × log(Total document count / (Including terms)) (Number of documents + 1), finally extract The top 3 weighted keywords (quadratic function, axis of symmetry, calculation) are used to filter out redundant words such as "how" and "calculate".

[0069] The speech data is first filtered for silence segments using the endpoint detection algorithm (VAD), and then the 13-dimensional Mel frequency cepstral coefficients (MFCC) are extracted as features and input into the educational terminology acoustic model for processing. The recognition accuracy reaches 94.2%, and the amount of speech keyframe data after compression is only 1.2KB.

[0070] Handwritten data is processed using a lightweight CNN model. The input is a 28×28 grayscale image. The model contains one convolutional layer with a 3×3 kernel (outputting 32 channels), one 2×2 pooling layer, and one fully connected layer, which can accurately recognize core symbols. Simultaneously, the cosine similarity is calculated using a test point association algorithm, with the formula: ,in For handwritten feature vectors, (The vector is a feature library of test points). A similarity of ≥0.8 indicates a successful association. Handwritten redundant correction marks are automatically removed. The core data size after processing is 2.8KB, and the overall processing latency is 78ms.

[0071] After fusion processing, the total size of the core data for a single round of interaction is 6KB, and the core information extraction accuracy rate reaches 96.5%, meeting the preset technical indicators.

[0072] Implementation of dynamic decision scheduling and transmission module:

[0073] The module calculates interaction priority based on a weighted formula, allocates transmission resources reasonably, and ensures transmission stability in weak network environments through multi-protocol optimization, as detailed below:

[0074] Priority calculation uses the formula (in Priority is determined by a value between 0 and 1. For network quality coefficients, To determine the urgency of the scenario, (For user preference coefficients), providing a basis for data transmission sorting;

[0075] The multi-protocol transmission strategy was optimized. Core data (handwritten symbols, keywords) was encapsulated using UDP (port 5000, packet size 1KB) and combined with LZ4 compression algorithm (compression ratio 4:1). The dynamic bitrate was adjusted to 200Kbps. Incremental data (such as supplementary Q&A steps) used TCP breakpoint resume transmission (fragment size 512B, timeout retransmission time 1s). By simplifying the three-way handshake to two, the transmission latency was reduced. At the same time, relying on campus edge node relay (the distance between the tablet and the node is 50 meters), the node temporarily stores the transmitted data. When the network fluctuates, only the unacknowledged fragments are retransmitted, and finally the end-to-end latency of 420ms is achieved.

[0076] Implementation of lightweight federated learning and semantic completion module:

[0077] The module optimizes model performance through lightweight federated learning, improves question-answering accuracy by combining a three-level semantic completion mechanism, and balances privacy protection and user experience, as detailed below:

[0078] The federated learning framework uses the Word2Vec word vector model, with a vector dimension of 100, a window size of 5, and 128 hidden layer nodes. The global model size is 75MB. To reduce transmission costs and protect privacy, each tablet iterates locally for 10 rounds before uploading only the parameter differences. The incremental data size per round is only 1.8KB, and privacy protection is achieved by adding Laplace noise, as shown in the formula. Privacy budget =0.1, noise range controlled within [-0.05, 0.05], region aggregation is performed by edge nodes, using a weighted average algorithm. (Assign device weights based on data volume), and then aggregate and synchronize them to the cloud to update global parameters;

[0079] The three-level semantic completion mechanism is progressive, with the basic layer matching the template for solving the quadratic function symmetry axis of the edge node cache. The enhancement layer, through a federated optimized word vector model, associates semantic information from the symmetry axis calculation steps and uses the underlying layer to access the user's historical consultation records (such as previous consultation records of similar question types) to supplement the context. Ultimately, semantic completion is achieved, clarifying the consultation requirement as finding a quadratic function. The axis of symmetry requires analytical steps, with a completion accuracy of 93% and a processing delay of 165ms.

[0080] Implementation of adaptive caching and energy optimization modules:

[0081] The module balances resource utilization efficiency and device battery life through an improved caching strategy and adaptive power consumption adjustment, as detailed below:

[0082] Multidimensional weighted caching is based on an improved LRU algorithm, and the formula for calculating cache weight is as follows: (in For resource access frequency, As energy consumption weight, (Based on learning priority), prioritize caching high-frequency, low-energy-consumption, and high-priority mathematical resources.

[0083] The adaptive power consumption adjustment is optimized for low battery and weak network scenarios. When the tablet's remaining battery is ≤20% and the network speed is <500Kbps, it automatically shuts down non-core functions such as redundant handwriting stroke detection and voice emotion analysis, retaining only the core recognition module, which can reduce power consumption by 32%. When the battery recovers to 30%, it automatically restarts the full-function mode to ensure that the learning experience is not affected.

[0084] In summary, the multimodal interactive learning and consultation integrated platform according to Embodiment 2 of this application achieves a consultation success rate of 99.2% and an end-to-end latency of 420ms under weak network (500Kbps) and low power (20%) scenarios. It also boasts a core data extraction accuracy of 96.5%, a missing data completion accuracy of 93%, and a 32% reduction in device energy consumption. All of these meet the preset technical solution indicators. Furthermore, the process is reproducible, the modules are deployable, and it is well-suited to the learning and consultation needs of rural middle schools. Localized edge node support reduces cloud bandwidth usage (by 90% compared to traditional solutions), and federated learning ensures that students' handwritten notes, voice recordings, and other private data are not leaked, complying with educational data compliance requirements.

[0085] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0086] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0087] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A multimodal interactive learning and consulting integrated platform, characterized in that: It includes an edge interaction layer, an edge node collaboration layer, and a cloud enabling layer, among which: The edge interaction layer consists of a multimodal data acquisition module, an intelligent state perception module, a core data processing module, a local learning cache module, an energy consumption adaptive module, and a dynamic decision scheduling module. It is used for the acquisition, state perception, accurate processing, local caching, energy consumption optimization, and transmission priority scheduling of multimodal interaction data. The edge node collaboration layer consists of a collaborative transmission relay module, a regional resource cache pool, and a lightweight federated node module, which are used to realize data temporary storage, incremental relay transmission, local learning resource support, and regional model aggregation. The cloud-enabled layer includes a lightweight learning knowledge base, a federated learning coordinator, and a global strategy optimization module, which are used to provide learning resources by grade level and subject, coordinate the entire federated learning process, and optimize the platform's global operation strategy.

2. The multimodal interactive learning and consulting integrated platform according to claim 1, characterized in that, The multimodal data acquisition module supports the synchronous acquisition and collaborative interaction of three core modal data types: text, voice, and handwriting. The intelligent state perception module collects multi-dimensional state information locally at the edge through a lightweight sensor fusion algorithm with an acquisition latency of ≤50ms. The multi-dimensional state information specifically includes network status, device status, and learning consultation status, providing a basis for subsequent data processing and strategy scheduling.

3. The multimodal interactive learning and consulting integrated platform according to claim 1, characterized in that, The core data processing module adopts a modal collaborative processing approach to adapt to the needs of multimodal learning and consultation. It performs TF-Lite keyword extraction, education domain vocabulary filtering, and state weighting optimization on text data, endpoint detection, acoustic model recognition of educational terms, and adaptive speech rate segmentation on speech data, and lightweight CNN recognition, stroke priority sorting, and test point association extraction on handwritten data. The overall core data extraction accuracy rate is ≥96%.

4. The multimodal interactive learning and consulting integrated platform according to claim 1, characterized in that, The dynamic decision-making and scheduling module adjusts the priority of multimodal data transmission based on multi-dimensional state weight factors, with the weight allocation being 0.4 for network quality, 0.3 for the urgency of the learning scenario, and 0.3 for user interaction preferences. The dynamic decision-making and scheduling module works in conjunction with the transmission relay module to achieve multi-protocol fusion transmission.

5. The multimodal interactive learning and consulting integrated platform according to claim 1, characterized in that, The regional resource cache pool stores high-frequency learning and consulting resources by grade level and subject, including formula templates, Q&A scripts for knowledge points, case studies of incorrect questions, and scenario rule templates, through a two-way linkage mechanism of local caching and edge caching.

6. The multimodal interactive learning and consulting integrated platform according to claim 1, characterized in that, The lightweight federated node module and the cloud-based federated learning coordinator work together to build a lightweight federated learning framework. Participants include edge devices and edge nodes. The cloud does not participate in the transmission of original learning data to ensure privacy. Each edge device optimizes the semantic completion model based on local learning consultation interaction data.

7. The multimodal interactive learning and consulting integrated platform according to claim 1, characterized in that, The semantic keep-alive completion module adopts a three-level completion strategy. The basic layer achieves fast matching based on the scenario rule templates shared by local and edge nodes. The enhancement layer achieves global semantic adaptation through a word vector model optimized by federated learning. The bottom layer combines user historical learning consultation data with collaborative recommendations from edge nodes to achieve context-related completion.

8. The multimodal interactive learning and consulting integrated platform according to claim 1, characterized in that, The local learning cache module adopts a multi-dimensional weighted caching strategy based on the LRU algorithm, adding energy consumption weight and learning priority weight. When the device is in a low power state, low priority cached content is automatically cleared, and high-frequency core consultation data and resources related to users' weak knowledge points are retained first.

9. The multimodal interactive learning and consulting integrated platform according to claim 1, characterized in that, The energy consumption adaptive module dynamically adjusts the platform's operating mode based on the device's power consumption and network status. In low power and weak network scenarios, it automatically shuts down non-core modality recognition functions and retains only core consulting service capabilities. Once the network and device status are restored, it automatically restarts the full-function mode.