A kind of integrated wisdom desk of early childhood STEM evaluation teaching and control management method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-08-11
AI Technical Summary
本发明用于幼儿 STEM 教学中的组网管理、算力分配、教学评估及服务器协同,能够解决当前学前教育信息化基础设施薄弱、教学评估主观、资源利用率低的问题
(1)灵活性高:分布式动态组网算法支持1-8桌灵活组网,课程角色切换无需物理操作,适配不同教学场景与班级规模。
Smart Images

Figure CN122311977B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of new information technology and artificial intelligence technology, and in particular relates to a smart desk integrating STEM assessment and teaching for young children and a control and management method. Background Technology
[0002] Facial recognition can accurately analyze student identity, emotions, activity status, and behavior, enabling automated monitoring and digital assessment, effectively improving teaching quality and solving problems faced by traditional early childhood education. However, achieving automated monitoring of 10-30 students per classroom requires expensive cloud servers and cloud graphics processing units (GPUs). More importantly, this relies on high network bandwidth, but many kindergartens currently suffer from weak IT infrastructure, low broadband speeds, and low levels of intelligence, making it difficult to implement many IT-based educational applications. In addition, existing GPU-enabled real-time video analysis servers on the market cost hundreds of thousands of yuan, with a single GPU supporting 16 channels of real-time video analysis, but even these are insufficient to meet market demand.
[0003] Research and analysis have identified the following main problems in early childhood STEM education: ① Existing educational assessments rely heavily on teachers' subjective judgments, resulting in underutilization of teaching data and a lack of scientific rigor and credibility in the assessment results. ② Teachers struggle to monitor the learning status of multiple children simultaneously and for extended periods, leading to insufficient personalized guidance and uneven distribution of teaching resources. ③ Traditional assessment methods are time-consuming and disruptive to normal teaching activities, often employing periodic tests rather than process-oriented evaluations. ④ Edge computing power is expensive, and there is currently no computing power implementation plan that meets the requirements for highly interactive and instant-feedback STEM teaching experiences, resulting in a lack of hardware support for high-quality teaching solutions. ⑤ Existing smart education devices use a fixed network architecture, which cannot be flexibly adjusted according to actual teaching needs and class size, leading to wasted or insufficient hardware resources. ⑥ Different courses and teaching activities require different equipment configurations and resources, resulting in overly specialized smart classroom equipment and low utilization rates. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by providing a smart desk and control management method that integrates STEM assessment and teaching for preschool children. This invention is used for network management, computing power allocation, teaching assessment, and server collaboration in preschool STEM teaching, and can solve the problems of weak information infrastructure, subjective teaching assessment, and low resource utilization in current preschool education.
[0005] The objective of this invention is achieved through the following technical solution: The first aspect of this invention provides a smart desk integrating STEM assessment and teaching for preschoolers, which internally includes an edge computing device, a sound acquisition sensor, four high-definition cameras, and a wiring enclosure. The wiring enclosure contains a sub-router and a switch. The high-definition cameras and the sound acquisition sensor are both connected to the switch, the switch is connected to the sub-router, the sub-router is connected to both the edge computing device and the main router, and the main router is connected to the edge server.
[0006] Furthermore, the edge computing device is used to realize local image recognition and voice interaction; The sound acquisition sensor is used to acquire voice signals to enable voice interaction; The high-definition camera is used to capture children's movement sequences and automatically identify the types of teaching aids placed on the table; The edge server is configured with a container pool, which includes a GPU instance pool, a course resource repository, and a model repository.
[0007] A second aspect of this invention provides a control and management method for the above-mentioned integrated smart desk for STEM assessment and teaching in preschool children, comprising the following steps: (1) Set the network scale according to teaching needs and number of students. The smart desk analyzes the required resources based on the identified teaching aid type or selected target course role, and sends a resource request to the edge server to obtain the corresponding course role container image, so as to dynamically reorganize the teaching area and change the original teaching area to the target course teaching area. (2) The smart desk calculates the load fusion score through adaptive dynamic weight calculation of teaching scenario, and adopts a two-level triggering mechanism to select to enter the warning state or the critical state, and executes resource pre-application and full resource application respectively; the edge server dynamically adjusts the attention weight of the smart desk's resource request characteristics and the edge server's state characteristics through cross-attention mechanism, and fuses them with the resource request characteristics to generate the optimal resource scheduling strategy. (3) Collect multimodal data, extract features and fuse them across modalities through a multimodal learning fusion model, obtain the scores, binary classification values and contribution of each modal data and natural language interpretation of the quantitative indicators of educational theory, so as to generate an evaluation report and output the evaluation basis and improvement suggestions. (4) A Server-only proximity collaboration algorithm is built based on the Client / Multi-Serve architecture, in which multiple edge servers form an adjacency graph to ensure the computing resources of processed requests in each time slice. Then, the surplus resources are allocated to the overloaded server through the discrete optimal transmission algorithm with entropy regularization. The neighborhood Sinkhorn-Knopp algorithm is used to iteratively solve the optimal transmission plan and generate a smooth migration execution volume to perform the migration and realize resource scheduling of multi-server collaboration.
[0008] Furthermore, the upper limit of the network scale is one network with eight desks, that is, eight smart desks are contained in the same local area network; The edge server is configured with a course resource repository and a model repository, which store course content, evaluation system, AI model, computing resources and storage space in a containerized manner to form a course role container image; the edge server is pre-loaded with all course role container images, so that after receiving a resource request from the smart desk, the corresponding course role container image can be distributed to the smart desk through elastic distribution.
[0009] Furthermore, the formula for calculating the load fusion score is as follows:
[0010] In the formula, Indicates the load fusion score. Indicates local GPU utilization. Indicates the current delay. This indicates a delay in the target. Indicates the number of video frames to be processed. Indicates the maximum tolerable number of video frames. , , These represent the parameter weights for different teaching stages; The two-level triggering mechanism specifically includes: when the load fusion score is greater than or equal to a preset first threshold for a duration of not less than a first set duration, a warning state is entered and the smart desk performs a resource pre-request; when the load fusion score is greater than or equal to a preset second threshold for a duration of not less than a second set duration, a severe state is entered and the smart desk performs a complete resource request.
[0011] Furthermore, the generation of the optimal resource scheduling strategy specifically includes: The resource request features of the smart desk are encoded into an n-dimensional query vector Q, whose dimensions include scenario type, request level, course attributes, and resource requirements; the status features of the edge server are encoded into an m-dimensional key vector K and value vector V, whose dimensions include hardware utilization, container status, and cache information. The query vector Q, key vector K, and value vector V are fed into the cross-attention mechanism to calculate the attention weights; By fusing resource request characteristics with attention weights, an optimal resource scheduling strategy is generated, including container allocation instructions, resource reservation strategies, priority adjustment, and preloading of model lists.
[0012] Furthermore, the multimodal data includes key points of children's postures, movement sequences, facial expressions and attention levels, language and dialogue, and artwork displays; The quantitative indicators of the educational theory include quantitative indicators of creation level and quantitative indicators of development level. The quantitative indicators of creation level include stability, structure, playability, completeness and narrative. The quantitative indicators of development level include health level, living level, exploration level, perception level and learning level.
[0013] Furthermore, step (3) specifically includes: Collect multimodal data and preprocess it to obtain preprocessed multimodal data, including motion trajectories, visual images, speech text, and facial expression data; Based on the preprocessed multimodal data, key features of each modality are extracted by quantization transformation method and aggregated into manual feature vectors according to time windows; Construct a multimodal learning fusion model and collect training data samples for training. During the training process, with minimizing the total loss function as the optimization objective, adjust the parameters of the multimodal learning fusion model until the preset training rounds are reached to obtain the trained multimodal learning fusion model. The reasoning process utilizes a trained multimodal learning fusion model to obtain scores, binary classification values, and the contribution of each modality's data to the scores of each educational theory quantitative indicator, along with natural language interpretation, in order to generate an evaluation report and output evaluation criteria and improvement suggestions.
[0014] Furthermore, the multimodal learning-based fusion model includes a modal feature encoder, a temporal aggregation layer, a cross-modal attention fusion unit, and an output head. The modal feature encoder includes a motion trajectory encoder, a visual image encoder, a speech-text encoder, and an expression encoder. The motion trajectory encoder encodes preprocessed motion trajectories to obtain corresponding evidence vectors; the visual image encoder encodes preprocessed visual images to obtain corresponding evidence vectors; the speech-text encoder encodes preprocessed speech-text to obtain corresponding evidence vectors; and the expression encoder encodes preprocessed expression data to obtain corresponding evidence vectors. The temporal aggregation layer performs positional encoding and average pooling on the evidence vectors corresponding to each modality to obtain the corresponding evidence vectors for each modality. The system generates a temporally-aware token sequence for each modality. This temporally-aware token sequence, along with handcrafted feature vectors, is input into a cross-modal attention fusion unit. Connection confidence graph tokens are inserted into the input temporally-aware token sequence to encode the topological information of visually detected connection points into graph vectors, which guide attention weight allocation. The final output is an aggregate vector. The output head includes a regression head, a classification head, and an interpretation head. The aggregate vector is input into these three heads respectively. The regression head outputs the scores of the quantitative indicators of educational theory, the classification head outputs the binary classification values of the quantitative indicators of educational theory, and the interpretation head, through cross-analysis of attention weights and evidence vectors from each modality, outputs the contribution of each modality's data to the scores of the quantitative indicators of educational theory and provides a natural language interpretation. The total loss function is specifically calculated as follows: the score loss is calculated based on the predicted score of the quantitative indicator of educational theory output by the multimodal learning fusion model and its corresponding real score label; the explanatory consistency loss is calculated based on the contribution of each modal data output by the multimodal learning fusion model to the score of the quantitative indicator of educational theory and its corresponding real evidence; and the total loss function is obtained by weighted summation of the score loss and the explanatory consistency loss.
[0015] Furthermore, the Server-only neighbor collaboration algorithm specifically includes: In each time slice, based on the adjacency graph, the guaranteed capacity is obtained according to the service level agreement of the edge servers; the computing power gap and surplus computing power are determined based on the load, guaranteed capacity, and elastic pool of the edge servers; the transmission cost matrix of the servers is calculated as the adjacency cost, and risk correction is performed on it according to the risk weight parameter; the kernel matrix is calculated based on the corrected adjacency cost and the entropy regularization strength parameter; based on the kernel matrix, computing power gap, and surplus computing power, the neighborhood Sinkhorn-Knopp algorithm is used to iteratively solve the dynamic adjustment variable pairs; the optimal transmission plan is calculated based on the dynamic adjustment variable pairs and the kernel matrix; the smooth migration execution amount for this time slice is calculated based on the optimal transmission plan and the smoothing factor; the migration is executed based on the smooth migration execution amount to achieve resource scheduling for multi-server collaboration.
[0016] Compared with existing methods, the beneficial effects of the present invention are as follows: (1) High flexibility: The distributed dynamic networking algorithm supports flexible networking of 1-8 tables, and the course role switching does not require physical operation, adapting to different teaching scenarios and class sizes.
[0017] (2) High computing power efficiency: The multi-level computing power scheduling algorithm realizes resource coordination between local and edge servers, with fast request response, high resource utilization, and reduced computing power cost.
[0018] (3) Assessment Science: Multimodal data fusion algorithms combined with educational theories generate quantitative and interpretable assessment reports, solving the problems of subjective and untraceable traditional assessments.
[0019] (4) High scalability: The algorithm is compatible with multi-server architecture, supports horizontal migration to various STEM platforms, and covers different age groups vertically, with a wide range of application scenarios. Attached Figure Description
[0020] Figure 1 This is a structural diagram of a smart desk integrating STEM assessment and teaching for preschoolers; Figure 2 This is a schematic diagram of the physical network composed of smart desks; Figure 3 This is a layout diagram of the STEM activity space; Figure 4 This is a flowchart of the control and management method for smart desks; Figure 5 This is a flowchart of a two-level triggering mechanism; Figure 6 This is a flowchart of the smart desk terminal triggering algorithm; Figure 7 This is a timing diagram of the three-dimensional scheduling algorithm; Figure 8 This is a flowchart of a server-side resource scheduling algorithm based on dynamic scheduling strategy learning; Figure 9This is a sequence diagram of a server-side resource scheduling algorithm based on dynamic scheduling strategy learning; Figure 10 This is a flowchart of a multimodal data fusion education assessment algorithm; Figure 11 This is a schematic diagram of the offline interactive teaching function module. Detailed Implementation
[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not intended to limit this application.
[0022] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0023] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "in response to determination," or "includes." Moreover, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process or method. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0024] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.
[0025] This invention provides a smart desk integrating STEM (Science, Technology, Engineering, and Mathematics) assessment and teaching for young children. STEM encompasses four disciplines: Science, Technology, Engineering, and Mathematics, emphasizing interdisciplinary education to cultivate comprehensive practical abilities. This smart desk enables highly interactive STEM teaching and allows for seamless, efficient, automatic, and comprehensive analysis and assessment of the teaching process.
[0026] See Figure 1 The smart desk integrating STEM assessment and teaching for preschoolers is equipped with one edge computing device, one sound acquisition sensor, four high-definition cameras, and one wiring box. The wiring box contains a sub-router and a switch. The high-definition cameras and sound acquisition sensor are connected to the switch, the switch is connected to the sub-router, the sub-router is connected to the edge computing device and the main router, and the main router is connected to the edge server.
[0027] It should be understood that a smart classroom includes an edge server, a main router, and several smart desks for integrated STEM assessment and teaching as described in this invention. Each smart desk can be configured into a specific teaching area according to teaching requirements, such as a metal play area, a clay play area, a wooden play area, a building block area, a video game area, or a creative area. Placing these smart desks in the kindergarten's STEM activity space allows teachers to seamlessly observe children's play-based learning, understand their developmental preferences and individual differences, and analyze data to personalize toys and materials, providing individualized instruction and optimizing children's development. This smart desk enables real-time intelligent recording and analysis of each child's individual autonomous activities in a natural setting, and better utilizes GPU computing resources for highly interactive STEM teaching. This provides educators and parents with valuable analytical data, including children's social skills, domain interests, and potential strengths.
[0028] Furthermore, edge computing devices are used to achieve local image recognition and voice interaction; sound acquisition sensors are used to collect voice signals to achieve voice interaction; high-definition cameras are used to capture children's action sequences and automatically identify the types of teaching aids placed on the table. Among them, four high-definition cameras are used to capture children's hand movements, details of their works, facial expressions, and the overall activity scene. The high-definition camera set in the best position can also automatically identify the types of teaching aids placed on the table, such as iron toys, clay toys, wooden toys, building blocks, etc.
[0029] Furthermore, the edge server is configured with a container pool, which contains multiple course role resources such as a GPU instance pool, a course resource repository, and a model repository.
[0030] It is worth mentioning that this invention also provides a control and management method for the integrated smart desk for STEM assessment and teaching in the above embodiments. This method includes four parts: a distributed dynamic networking algorithm, a computing power dynamic scheduling algorithm, a multimodal data fusion educational assessment algorithm, and a server-only proximity collaboration algorithm. These four parts of the algorithm can respectively realize dynamic networking applications, computing power scheduling applications, teaching assessment applications, and multi-server collaborative applications.
[0031] See Figure 4 The control and management method specifically includes the following steps: (1) Distributed dynamic networking algorithm: The networking scale is set according to teaching needs and the number of students. The smart desk analyzes the required resources based on the identified teaching aid type or the selected target course role, and sends a resource request to the edge server to obtain the corresponding course role container image, so as to dynamically reorganize the teaching area and change the original teaching area to the target course teaching area.
[0032] Furthermore, the system supports flexible network configuration of smart desks based on teaching needs and the number of students. The networking method is standardized, theoretically unlimited, and in practical applications, it supports a maximum of eight desks per network. Therefore, preferably, the maximum network size is eight desks per network, meaning eight smart desks within the same local area network.
[0033] Furthermore, the edge server is configured with a course resource repository and a model repository, storing course content, evaluation systems, AI models, computing resources, and storage space in containers to form course role container images. The edge server has pre-set container images of all course roles (such as iron toys, clay toys, wooden toys, video games, building blocks, and creative works), which can be elastically distributed to the smart desks upon receiving resource requests. When a new course role is added, only its corresponding course role container image needs to be imported into the edge server, eliminating the need for separate operations on each smart desk and effectively reducing maintenance costs.
[0034] Specifically, after automatically identifying the type of teaching aid (such as building blocks), the smart desk automatically analyzes the resources required for that type of teaching aid (such as course roles, interactive content, AI models, computing resources, evaluation systems, etc.) and sends a resource request to the edge server. Upon receiving the resource request from the smart desk, the edge server elastically distributes resources such as course roles, evaluation systems, AI models, computing resources, and storage space, and sends the corresponding course role container image to the smart desk. Furthermore, users can select a target course role on the smart desk screen, and the smart desk analyzes the required resources based on the selected target course role. Subsequently, after receiving the course role container image, the smart desk automatically loads and updates the course, the corresponding evaluation system, and the corresponding resources. This enables dynamic switching of teaching roles (such as switching from the play area to the creative area) without physically moving the desk, teaching aids, or requiring children to move. By loading the target course role container image onto the smart desk, the smart desk can be set as the teaching area corresponding to the target course role, realizing the dynamic reorganization of the teaching area and changing the original teaching area to the target course teaching area (such as changing the original iron play area to the building block area).
[0035] (2) Dynamic computing power scheduling algorithm, including a smart desk-side triggering algorithm and a server-side resource scheduling algorithm. Specifically, the smart desk-side triggering algorithm calculates a load fusion score based on adaptive dynamic weights within the teaching scenario, and employs a two-level triggering mechanism to select between a WARNING or CRITICAL state, executing resource pre-request and full resource request respectively. The server-side resource scheduling algorithm involves the edge server dynamically adjusting the attention weights of the smart desk's resource request characteristics and the edge server's state characteristics through a cross-attention mechanism, and fusing these with the resource request characteristics to generate the optimal resource scheduling strategy.
[0036] It should be understood that smart desks support distributed connectivity. In addition to the computing power resources that the smart desks themselves have to serve teaching, edge servers can provide more additional computing power and can dynamically respond to the computing power resource requests of different smart desks, thereby realizing dynamic scheduling of computing power.
[0037] Furthermore, the formula for calculating the load fusion score is as follows:
[0038] In the formula, Indicates the load fusion score. Indicates local GPU utilization. Indicates the current delay. This indicates a delay in the target. Indicates the number of video frames to be processed. Indicates the maximum tolerable number of video frames. , , These represent different parameter weights, and satisfy the following conditions: , , , The curriculum is dynamically adjusted at different teaching stages, for example, the exploratory phase is set up... Display period settings .
[0039] Furthermore, such as Figure 5 and Figure 6 As shown, the two-level triggering mechanism specifically includes: when the load fusion score is greater than or equal to a preset first threshold for a duration of not less than a first set duration, the system enters the WARNING state, at which time the smart desk performs resource pre-request; when the load fusion score is greater than or equal to a preset second threshold for a duration of not less than a second set duration, the system enters the CRITICAL state, at which time the smart desk performs a complete resource request. Subsequently, when the load fusion score is less than a preset third threshold for a duration of not less than a third set duration, the system transitions from the CRITICAL state to the WARNING state; when the load fusion score is less than a preset fourth threshold for a duration of not less than a fourth set duration, the system transitions from the WARNING state to the NORMAL state. Among them, the first threshold is less than the second threshold, the second threshold is greater than the third threshold, the third threshold is greater than the fourth threshold, and the fourth threshold is less than the first threshold. For example, the first threshold is preset to 0.75, the second threshold is preset to 0.9, the third threshold is preset to 0.85, and the fourth threshold is preset to 0.7. The first set duration is greater than the second set duration, the second set duration is less than the third set duration, the third set duration is less than the fourth set duration, and the fourth set duration is greater than the first set duration. For example, the first set duration is 5 seconds, the second set duration is 2 seconds, the third set duration is 4 seconds, and the fourth set duration is 10 seconds.
[0040] For example, when the load fusion score is greater than or equal to the first threshold for 5 seconds, it enters the WARNING state. At this time, the smart desk performs resource pre-request, pre-requesting edge server resources and warming up the required AI model, but does not activate the computing task. When the load fusion score is greater than or equal to the second threshold for 2 seconds, it enters the CRITICAL state. At this time, the smart desk performs a full resource request, immediately requesting full computing power resources from the edge server to ensure the normal operation of high computing power demand tasks (such as multimodal generation and complex object detection). When the load fusion score is less than the third threshold for 4 seconds, it transitions from the CRITICAL state to the WARNING state. When the load fusion score is less than the fourth threshold for 10 seconds, it transitions from the WARNING state to the NORMAL state. Figure 5 As shown.
[0041] Furthermore, such as Figure 8 As shown, the edge server dynamically adjusts the attention weights of the smart desk's resource request features and the edge server's state features through a cross-attention mechanism, and then fuses them with the resource request features to generate the optimal resource scheduling strategy. Specifically, the resource request features of the smart desk are encoded into an n-dimensional query vector Q, whose key dimensions include: scenario type (security alert, skill assessment, behavior analysis, etc.), request level (urgency, time limit), course attributes (current course ID, model dependencies), resource requirements (GPU memory, CPU core count), etc. The edge server's state features are encoded into an m-dimensional key vector K and value vector V, whose key dimensions include: hardware utilization (current GPU / memory / storage usage), container status (number and type of running containers), cache information (list of loaded course models), etc. Then, the query vector Q, key vector K, and value vector V are fed into the cross-attention mechanism, and the attention weights are calculated according to the following formula:
[0042] In the formula, Indicates attention weight; This is the activation function that converts scores into a probability distribution; the superscript T indicates the transpose of a vector or matrix. The dot product of the query vector Q and the transpose of the key vector K represents the similarity. This represents the dimension of the key vector K. This represents the scaling factor to prevent the gradient from vanishing due to an excessively large dot product. Finally, the resource request features are fused with the attention weights to generate the optimal resource scheduling strategy, including container allocation instructions, resource reservation strategies, priority adjustments, and preloading of model lists.
[0043] (3) Multimodal data fusion education assessment algorithm: Collect multimodal data, extract features and fuse across modalities through a multimodal learning fusion model, obtain the scores, binary classification values and contribution of each modal data and natural language interpretation of the quantitative indicators of education theory, generate an assessment report, and output the evaluation basis and improvement suggestions.
[0044] The multimodal data includes key points of children's postures, movement sequences, facial expressions and attention spans, language and dialogue, and artwork displays. The quantitative indicators of educational theory include quantitative indicators of creation level and development level. The quantitative indicators of creation level include stability, structure, playability, completeness, and narrative, while the quantitative indicators of development level include health level, living standards, exploration level, perception level, and learning level.
[0045] In this embodiment, the multimodal data fusion education assessment algorithm specifically includes the following sub-steps: (3.1) Collect multimodal data and preprocess it to obtain preprocessed multimodal data, including motion trajectories, visual images, speech text, and facial expression data.
[0046] (3.2) Based on the preprocessed multimodal data, the key features of each modality are extracted by the quantization transformation method in Table 2, and then aggregated into a manual feature vector according to the time window.
[0047] (3.3) Construct a multimodal learning fusion model and collect training data samples for training. During the training process, with minimizing the total loss function as the optimization objective, adjust the parameters of the multimodal learning fusion model until the preset training rounds are reached, and obtain the trained multimodal learning fusion model.
[0048] Furthermore, the multimodal learning-based fusion model includes multiple modal feature encoders, a temporal aggregation layer, a cross-modal attention fusion unit, and an output head. The modal feature encoders include a motion trajectory encoder, a visual image encoder, a speech-text encoder, and an expression encoder. The motion trajectory encoder encodes the preprocessed motion trajectory to obtain the corresponding evidence vector; the visual image encoder encodes the preprocessed visual image to obtain the corresponding evidence vector; the speech-text encoder encodes the preprocessed speech-text to obtain the corresponding evidence vector; and the expression encoder encodes the preprocessed expression data to obtain the corresponding evidence vector. The temporal aggregation layer performs positional encoding and average pooling on the evidence vectors corresponding to each modality to obtain a temporally aware token sequence for each modality. The temporally aware token sequence for each modality, along with the handcrafted feature vector Ft, is input into the cross-modal attention fusion unit. Connection confidence graph tokens are inserted into the input temporally aware token sequence to encode the topological information of the visually detected connection points into graph vectors, which guide the attention weight allocation. The final output is an aggregated vector. The output head includes a regression head, a classification head, and an interpretation head. The aggregation vector is input into the regression head, classification head, and interpretation head, respectively. The regression head is used to output the scores of the quantitative indicators of educational theory, such as the quantitative scores of 1-5 levels. The classification head is used to output the binary classification values of the quantitative indicators of educational theory, such as stable or unstable (Stable / Unstable). The interpretation head is used to output the contribution of each modality's data to the scores of the quantitative indicators of educational theory through cross-analysis of attention weights and evidence vectors of each modality, as well as natural language interpretations, such as the natural language interpretation of "It is recommended to strengthen the fixation of the second connection point, and repeated loosening actions were detected," and to highlight the problematic parts in the video.
[0049] Furthermore, the trained multimodal learning fusion model is obtained through the following method: First, training data samples are collected; then, each sample is input into the multimodal learning fusion model to obtain the scores, binary classification values, and the contributions and natural language interpretations of each modality's data to the scores of the educational theory quantitative indicators; then, the score loss is calculated based on the predicted scores of the educational theory quantitative indicators output by the multimodal learning fusion model and their corresponding true score labels, and the interpretation consistency loss is calculated based on the contributions of each modality's data to the scores of the educational theory quantitative indicators and their corresponding true evidence. The total loss function is obtained by weighted summation of the score loss and the interpretation consistency loss; subsequently, during training, the parameters of the multimodal learning fusion model are adjusted with minimizing the total loss function as the optimization objective until the preset training rounds are reached, finally obtaining the trained multimodal learning fusion model.
[0050] (3.4) Reasoning process: Use the trained multimodal learning fusion model to obtain the scores, binary classification values and the contribution of each modal data to the scores of each educational theory quantitative indicator and natural language interpretation, so as to generate an evaluation report and output the evaluation basis and improvement suggestions.
[0051] (4) Server-only proximity collaboration algorithm: The Server-only proximity collaboration algorithm is built based on the Client / Multi-Serve architecture, where Client represents the smart desk and Multi-Serve represents multiple edge servers. Multiple edge servers form an adjacency graph to ensure the computing resources of processed requests in each time slice. Then, the surplus resources are allocated to the overloaded server through the discrete optimal transmission algorithm with entropy regularization. The neighborhood Sinkhorn-Knopp algorithm is used to iteratively solve the optimal transmission plan and generate a smooth migration execution volume to perform the migration and realize resource scheduling of multi-server collaboration.
[0052] Furthermore, the Server-only proximity collaboration algorithm specifically includes: in each time slice, based on the adjacency graph, obtaining guaranteed capacity according to the service level agreement of the edge servers; determining the computing power gap and surplus computing power based on the load, guaranteed capacity, and elastic pool of the edge servers; calculating the transmission cost matrix of the servers as the adjacency cost, and correcting it for risk based on the risk weight parameter; calculating the kernel matrix based on the corrected adjacency cost and the entropy regularization strength parameter; iteratively solving the dynamic adjustment variable pairs using the neighborhood Sinkhorn-Knopp algorithm based on the kernel matrix, computing power gap, and surplus computing power; calculating the optimal transmission plan based on the dynamic adjustment variable pairs and the kernel matrix; calculating the smooth migration execution amount for the time slice based on the optimal transmission plan and the smoothing factor; and executing the migration based on the smooth migration execution amount to achieve resource scheduling for multi-server collaboration.
[0053] The present invention’s integrated smart desk for STEM assessment and teaching in preschools and its control and management method are described in detail below with reference to embodiments, which will make the purpose and effects of the present invention more apparent.
[0054] Example 1: Dynamic Networking and Multimodal Assessment of Thematic STEM Teaching This embodiment targets the "Building Blocks" STEM course for kindergarten middle class, realizing dynamic networking of smart desks, interactive teaching, and seamless evaluation of children's learning data. The specific process is as follows: 1.1 System Hardware Deployment and Initial Network Setup One edge server, one main router, and four smart desks are deployed in the kindergarten's STEM activity space. The edge server is configured with a container pool, which is an abstract software resource containing multiple course role resources. Each course resource can be further divided into different sub-resources. This container pool includes a GPU (Graphics Processing Unit) instance pool, a course resource repository, and a model repository. Each smart desk is equipped with one edge computing device, one sound sensor (using a microphone as the sound sensor to collect voice signals), four high-definition cameras, and one circuit box. Figure 1 As shown, the edge computing device supports local image recognition and voice interaction. Four high-definition cameras are used to capture children's hand movements, artwork details, facial expressions, and the overall activity scene. The circuit box contains sub-routers and switches. The high-definition cameras are connected to the switch, the switch is connected to the sub-routers, the sub-routers are connected to the edge computing device and the main router, and the main router is connected to the edge server. Figure 2 As shown, the four smart desks are connected to the edge server via the main router, forming a distributed architecture of one network and four desks, as illustrated. Figure 3As shown, "one network, four desks" refers to four smart desks deployed in the same local area network. These four smart desks can be in the same teaching area, such as all four smart desks being the iron play area, or they can be in different teaching areas, such as the iron play area, clay play area, wooden play area, and building block area respectively. The specific settings can be configured according to actual needs.
[0055] According to the teaching requirements, these four smart desks are respectively set up as iron play areas, clay play areas, wooden play areas, and building block areas. Teachers can set any of the four course roles (iron play / clay play / wooden play / building block) on any smart desk. Then, the high-definition camera of the smart desk automatically identifies the teaching aids (iron play aids / clay play aids / wooden play aids / building block teaching aids) on the desktop, or teachers can manually select course options (iron play course / clay play course / wooden play course / building block course) through the smart desk's screen. The smart desk then begins to parse the required resources and sends a virtual domain self-generation request to the edge server to dynamically reorganize the teaching area, changing the original teaching area to the target course teaching area. For example, if a smart desk is selected for teaching, and this desk was initially used as a play area (either as a metal play area or as a clay play area in the previous lesson), but now has building blocks on its surface, the smart desk's high-definition camera needs to automatically identify the building blocks, or the user can manually select the building block course option on the screen. After analysis, the smart desk will determine that it needs to use the building block area for teaching. Therefore, the teaching area needs to be dynamically reorganized to change the original metal play area or clay play area to the building block area corresponding to the building block course for subsequent teaching. Subsequently, the edge server sends corresponding course role container images (including teaching videos, course evaluation systems, course AI models, computing resources, storage space, etc.) to the four smart desks. For example, if four different teaching areas are set up, the edge server will accurately send the corresponding teaching resources to each smart desk, such as sending the container image for the building block course to the building block area and the container image for the metal play area. After receiving the image, the smart desk automatically loads course resources and completes the network setup. The network status is displayed in real time on the teacher's terminal (e.g., "Desks 1-4 are connected, network type: one network for four desks, resources loaded"), without the need to physically move the desks or teaching aids.
[0056] 1.2 Interactive Teaching and Dynamic Scheduling of Computing Power In addition to automatically and seamlessly observing children's learning progress and providing teachers with information on children's learning and developmental preferences and personalized teaching suggestions, smart desks can also act as educators, offering interactive and gamified teaching methods. An example of an offline interactive teaching module is shown below. Figure 11As shown, in a construction-themed teaching scenario, a large number of plastic building blocks are placed on the smart desks in the block-building teaching area. Children can independently build their creations by following the instructional videos on the smart desk screens. During this process, the smart desks record the children's creative expressions and can interact with them via audio and video.
[0057] Taking the block-building teaching area as an example, the smart desk acts as the teacher. Children follow the guided videos on the smart desk's screen to build blocks, such as constructing a sloping roof. The smart desk's local computing device handles the following basic tasks: capturing the children's movement sequences through a high-definition camera, recognizing the block teaching aids, and enabling voice interaction through a built-in microphone. For example, if a child asks, "How do I build a sloping roof?", the smart desk immediately replies, "Your house needs to be symmetrical, just like our bodies need to be balanced on both sides to be stable." After the children complete their construction, they place their small house in front of the camera for recognition. The desk then "scans" the physical house into the interactive teaching materials, allowing them to continue with subsequent contextualized course content in the form of a digital house.
[0058] The smart desk's on-device triggering algorithm: When the smart desk's local computing resources are insufficient, it can dynamically request computing resources from the edge server. Through continuous optimization, it aims to fully utilize both the smart desk's local computing resources and the edge server's resources, while ensuring that the edge server can respond appropriately to the computing resource requests. Its working procedure is as follows: Figure 6 As shown. Raw metrics such as local GPU utilization and current latency are monitored, and after processing, parameter weights are dynamically adjusted to calculate the load fusion score. The parameter weights... , , When dynamically adjusting the curriculum at different teaching stages, during the exploration phase, when children are freely building, GPU utilization is assigned a weight of 60%, indicating a high GPU load weight and the need for extensive real-time 3D rendering. Slightly higher Slightly lower, can be set , , During the exhibition period, when children are close to completing their artwork, or when digital scanning of their work is initiated, or during the artwork display phase, reasoning is given a 50% weighting for delay. Delay sensitivity is given a high weighting, requiring rapid feedback to the children's actions. Lower, Increase, you can set , , Subsequently, a two-level triggering judgment is performed based on the load fusion score, employing a two-level triggering mechanism with hysteresis characteristics: when the load fusion score... When the duration of a condition greater than or equal to the first threshold is no less than 5 seconds, the system enters a WARNING state. At this time, the smart desk performs resource pre-request, pre-requesting edge server resources and preheating the required AI model; when the load fusion score... When the duration of the second threshold is greater than or equal to 2 seconds, the system enters the CRITICAL state. At this time, the smart desk executes a full resource request and immediately requests full computing power resources from the edge server.
[0059] Server-side resource scheduling algorithm: This is a multi-request resource contention algorithm that dynamically adjusts scheduling priorities through a cross-attention mechanism.
[0060] Among them, a three-dimensional scheduling algorithm was constructed based on the urgency of the teaching scenario, the scarcity of resources, and the relevance of courses. Its working sequence is as follows: Figure 7 As shown, when the smart desk lacks resources, it submits a resource request to the edge server, specifically through the `RequestType=CRITICAL` directive. The edge server determines the scheduling priority by calculating the request scheduling priority score (`schedule_priority(request)`), and then issues a call request to the dynamic resource pool to allocate a container (`assignContainer(request)`). The allocated container instance is then returned to the smart desk. In this three-dimensional scheduling algorithm, dimension 1 is the urgency of the teaching scenario (e.g., safety > teaching); dimension 2 is the resource scarcity (e.g., `CRITICAL` > `WARNING`); and dimension 3 is the course relevance (shared model cache within the same course group). The scheduling priority can be determined through this three-dimensional scheduling algorithm, and its expression is:
[0061] In the formula, Indicates scheduling priority. Indicates the urgency of the teaching scenario. Indicates the degree of resource scarcity. Indicates course relevance. , , These represent the weights of the three dimensions: urgency of the teaching scenario, scarcity of resources, and relevance to the course. For example, in this embodiment, the weights are set as follows: , , At this point, the scheduling priority is calculated as follows: (urgency of teaching scenario × 60% + resource scarcity × 30% + course relevance × 10%) × 100.
[0062] The aforementioned manually defined three-dimensional weight formula is rather rigid and cannot adapt to more complex or unknown scenarios. Therefore, a cross-attention mechanism is needed to dynamically adjust the weights of each dimension and learn dynamic resource scheduling strategies to replace the manual formula. By constructing a dynamic correlation model between resource request features and server status features, resource scheduling decisions are transformed into an attention allocation problem in the feature space, enabling the system to autonomously learn the optimal resource scheduling strategy. Its working procedure is as follows: Figure 8 As shown.
[0063] like Figure 8 As shown, the resource request features of the smart desk are encoded as an n-dimensional query vector Q, with key dimensions including: scenario type (security alert, skills assessment, behavior analysis, etc.), request level (urgency, time limit), course attributes (current course ID, model dependencies), and resource requirements (GPU memory, number of CPU cores). The state features of the edge server are encoded as an m-dimensional key vector K and value vector V, with key dimensions including: hardware utilization (current GPU / memory / storage usage), container status (number and type of running containers), and cache information (list of loaded course models). Subsequently, the query vector Q, key vector K, and value vector V are fed into the cross-attention mechanism to calculate attention weights. Specifically, ① Feature projection: The query vector Q corresponding to the resource request features and the key vector K and value vector V corresponding to the edge server's state features are mapped to a unified representation space through a linear layer; ② Attention weight generation: The similarity between Q and K is calculated to obtain a weight matrix, which reflects the "degree of attention" of each request to each resource unit of the server, and then normalized weights are obtained through softmax normalization; ③ State feature refinement: The normalized weights are applied to V (server state features) to generate a refined feature representation that highlights the server state information most relevant to the current request. Finally, the resource request features are fused with the attention weights output by the cross-attention mechanism to generate the optimal resource scheduling strategy. The resource scheduling strategy includes container allocation instructions (which container to allocate), resource reservation strategies (which course models to preheat), priority adjustment, and preloading of the model list. The work sequence is as follows: Figure 9 As shown, both the attention decision engine and the container pool run on edge servers.
[0064] The measured data are shown in Table 1, comparing the three-dimensional scheduling algorithm based on the three-dimensional weight formula and the server-side resource scheduling algorithm based on the cross-attention mechanism.
[0065] Table 1: Comparison of measured data between 3D weighted formula scheduling and cross-attention mechanism scheduling
[0066] As shown in Table 1, the server-side resource scheduling algorithm based on the cross-attention mechanism can automatically learn the implicit relationship between resource requests and server states without requiring manual specification of weight ratios. Compared to the limitations of the three-dimensional scheduling algorithm based on the three-dimensional weight formula, the server-side resource scheduling algorithm based on the cross-attention mechanism achieves a shift from manual empirical formulas to data-driven decision-making, from linear feature combination to nonlinear relationship mining, and from static parameter configuration to continuous adaptive optimization.
[0067] 1.3 Multimodal Data Fusion Evaluation Multimodal data is collected, including: key points of children's posture, movement sequences, facial expressions and concentration, language and dialogue, and artwork display. These five aspects are used to evaluate quantitative indicators of educational theory. Specifically, based on the technology mapping and educational quantification table shown in Table 2 and the corresponding multimodal data fusion educational evaluation algorithm, children's teaching performance results are transformed into five scoring indices: stability, structure, playability, completeness, and narrative. Based on children's operational process data, five developmental level indices are calculated: health level, living level, exploration level, perceptual level, and learning level. The specific multimodal data fusion educational evaluation process is as follows: Figure 10 As shown.
[0068] Table 2: Technology Mapping and Education Quantification Table
[0069] Taking stability assessment as an example, this paper demonstrates the technical path of designing a multimodal data fusion education assessment algorithm, namely multimodal data acquisition → feature extraction → mapping calculation → output of quantitative indicators of educational theory. A multimodal data fusion education assessment algorithm was designed, realizing the fusion of physical perception and cross-modal attention.
[0070] In the multimodal data fusion educational assessment algorithm, the inputs include: RGB frames + depth information + pose keypoint trajectories (time series) + facial expression information + speech signals + artwork display (component recognition), etc. RGB frames, depth information, pose keypoint trajectories (time series), and facial expression information are acquired through a high-definition camera; speech signals are acquired through a built-in microphone; and artwork display (component recognition) is acquired through both a high-definition camera and a microphone. The outputs include: a stability score S∈[1,5], a binary classification value, and the contribution of each modality to the score S along with a natural language explanation. Based on this, an assessment report is generated, along with evaluation criteria and improvement suggestions. Specifically, this is achieved through the following process: 1) Input and Preprocessing: ① RGB frames (25-30 FPS) + depth information frames (25-30 FPS) → Generate a timestamped RGB-D frame sequence. ② Real-time estimation of pose keypoints (hand / body) trajectories → Action sequence. ③ Speech-to-text conversion for teachers and students → Keyword detection. ④ Facial expressions (facial emotion intensity vector). ⑤ Short video showcasing artwork → Shaking / looseness detection, component category, structure, position, etc.
[0071] 2) Extract key features by modality (corresponding to the quantization transformation method in Table 2): ① "Use material connectors to firmly connect materials" → Connectivity features, including: the number of "connection actions" detected within the window (action detectors detect hand actions such as insertion, buckling, and screwing), denoted by A1; the duration of continuous pressing / holding after the connection action (detected by pressure / force sensing or hand stillness), denoted by A2; visual analysis to verify whether the connection is actually formed, denoted by A3, range [0,1]; the number of times the same connection point is repeatedly reinforced (high redundancy usually improves stability), denoted by A4. ② "Use materials for modeling and assembly" → Assembly coherence features, including: the ratio of whether key target components are in place in the final work (target template matching / semantic detection), denoted by B1; the similarity between the action sequence and the ideal assembly steps (similarity calculated using a seq2seq model), denoted by B2; the seam / fitting degree between parts (based on gap detection / color bonding judgment in the image), denoted by B3, range [0,1]. ③ "Stable and stable during display" → Stability characteristics during the display period, including: the energy of optical flow / object center of mass jitter in the display short video (higher values indicate instability), denoted by C1; the count of loosening / falling events detected during the display period, denoted by C2; and the ratio of no structural changes within 5 seconds after the artwork is lifted, denoted by C3. The above features are aggregated into a manual feature vector Ft=[A1,A2,A3,A4,B1,B2,B3,C1,C2,C3] according to time windows.
[0072] 3) Multimodal learning-based fusion model: The specific design is as follows: modal feature encoder → temporal aggregation layer → cross-modal attention fusion (CMAF) → output head.
[0073] a) Modal feature encoder: ① Motion trajectory encoder Input the preprocessed keypoint trajectory and connection action timestamps, encode them using a self-attention mechanism, and output the evidence vector corresponding to the action trajectory. ② Visual image encoder Input the preprocessed final artwork image features, and output the evidence vector corresponding to the visual image. ③ Voice-to-text encoder Keyword statistics are performed through speech recognition, and evidence vectors corresponding to the speech text are output. ④ Expression encoder Input the preprocessed facial emotion results and intensity, and output the evidence vector corresponding to the expression data. .
[0074] Each modal feature encoder outputs its corresponding evidence vector. The evidence vector corresponds to the quantitative conversion algorithm in the above technology mapping and education scale. It is an interpretable quantitative feature that is used to calculate evaluation results and to trace the source of interpretation. The evidence vector includes the number of times the connector is used, the repetition rate of reinforcement actions, the structural gap rate, the amplitude of shaking, etc.
[0075] b) Temporal aggregation layer: Displays the same modality within a window. After encoding the fragments within, position encoding and average pooling are performed to preserve the time sequence information and output the time-aware token sequence for each modality.
[0076] c) Cross-modal attention fusion unit: The time-aware token sequence for each modality, along with the handcrafted feature vector Ft, is input into a lightweight cross-modal attention fusion unit. Connection graph tokens are inserted into the input time-aware token sequence. These tokens encode the topological information of visually detected connections (such as confirmed connections of certain nodes, number of repeated reinforcement iterations, etc.) into graph vectors to guide attention weight allocation, thereby making the multimodal learning-based fusion model more sensitive to connection quality evidence. Finally, the cross-modal attention fusion unit performs a weighted fusion of the time-aware token sequence and the handcrafted feature vector Ft for each modality based on the attention weights, outputting an aggregated vector. .
[0077] d) Output Header: ① Regression Header Used to output a stability score ②Classification Head The first header is used to output the binary classification value of stability, Stable / Unstable. The second header is used to output the contribution vector G of each modality data to the stability score and the natural language interpretation, where G can be expressed as:
[0078] In the formula, This indicates the contribution of the motion trajectory mode to the stability score. This indicates the contribution of the visual image modality to the stability score. This indicates the contribution of the speech-text modality to the stability score. This indicates the contribution of facial expression data modality to the stability score.
[0079] 4) Training data and loss function: Collect training data samples to train the multimodal learning fusion model. During the training process, minimize the total loss function as the optimization objective and adjust the parameters of the multimodal learning fusion model until the preset training rounds are reached.
[0080] a) Training data samples: Each sample contains manually labeled stability labels Slabel (i.e., continuous stability score or binary classification value labels) and teacher labels (e.g., insufficient connectivity).
[0081] b) Input each sample into the multimodal learning fusion model, and after the process described in step 3), obtain the output stability score, the binary classification value of stability, the contribution vector of each modality data to the stability score, and the natural language interpretation.
[0082] c) Calculate the total loss function using the following formula:
[0083]
[0084]
[0085] In the formula, Represents the total loss function; This represents the scoring loss, used to ensure that the stability score output by the multimodal learning fusion model is close to that of human evaluation. This represents the interpretation consistency loss, used to ensure consistency between the interpretation (evidence vector selection) of the multimodal learning fusion model and the basis for teacher annotation (i.e., the artificial feature vector Ft). To determine the weighting of the scoring accuracy, To control for the weight of interpretability consistency; y represents the true stability label Slabel given by the teacher / expert, which is taken from the continuous stability score, ranging from 1 to 5; This represents the stability score of the predictions made by the modality learning fusion model. The evidence selection distribution (e.g., A1, B1, or C1) generated by the interpretation head is determined by the contribution of each modal data output by the interpretation head to the stability score. This represents the i-th piece of genuine evidence manually labeled, i.e., which features the teacher labeled to support the score.
[0086] From the total loss function Optimizing the parameters of a multimodal learning fusion model can improve its generalization and interpretability. Furthermore, more complex loss terms, such as multimodal consistency loss, can be added to further improve the model.
[0087] 5) Reasoning and Explainable Output (Feedback Available to Teachers / Parents): The output of the reasoning process includes not only a stability score. It also outputs: modal contribution vector G (e.g., visual 0.45, action 0.30, tactile 0.15, language 0.10); evidence prompts: based on the cross-analysis of A1, A2, B1, B2, etc. with attention weights, it generates natural language prompts, such as: "It is recommended to strengthen the fixation of the second connection point (repeated loosening action detected)" or "The work wobbles slightly when displayed, which may be due to loose seams"; visualization: the problem area is highlighted on the video interface and short video clips are provided as evidence.
[0088] 6) An instructional prediction engine is included. Based on the scores, binary classification values, and contributions of each modality's data to the scores of various educational theory quantitative indicators obtained through the reasoning process, along with natural language interpretation and historical trends, it generates personalized assessment reports and outputs evaluation criteria and improvement suggestions. This provides teachers and parents with scientifically based, visualized assessment reports of children's learning and growth.
[0089] Furthermore, in this embodiment, the evaluation of the educational quantification table shown in Table 2 mentions a STEM teaching evaluation method based on interpretive heads and evidence vectors. This method processes the raw data through multimodal feature extraction, extracts the displayed quantitative features, and constructs a hand-crafted feature vector as the evidence vector provided by the teacher (obtained through step 2 above). This achieves a closed-loop evaluation mechanism of "scoring-basis-interpretation," ensuring the objectivity of the evaluation indicators for children's learning activities while enhancing the transparency and guiding value of educational evaluation. Taking the stability index as an example, some evidence vectors (obtained through step 2 above) are listed below: A1: Number of times and proportion of connectors used; A2: Repetition rate of reinforcement actions; B1: Confidence level of detection of the presence or absence of connectors in the finished product; B2: Structural gap rate or component fit; C1: Shaking amplitude during the display of the work; C2: Collapse or loosening rate during the display.
[0090] In the feature fusion stage, the multimodal learning-based fusion model, on the one hand, inputs the handcrafted feature vectors into the cross-modal attention fusion unit for weighted processing to generate a comprehensive stability score; on the other hand, the aggregated vector output by the cross-modal attention fusion unit... The data is simultaneously input into a separate interpretation head. This interpretation head, composed of a trainable attention mechanism and a rule generation module, can generate corresponding explanatory outputs based on the importance distribution of different features. For example, when A1 and C1 scores are high while B2 is low, the interpretation head can automatically generate the following comment: "The child frequently uses the connectors and the work remains stable during display, but there are many structural gaps, and the stability still has room for improvement." Thus, it not only outputs a quantitative stability score but also simultaneously outputs a natural language explanation based on the chain of evidence. This design addresses the problem of existing educational assessment methods that only provide scores and lack source attribution and explanation, ensuring the comprehensibility and operability of evaluation results in educational settings.
[0091] Through data collection, scale mapping, modal feature extraction, and cross-modal fusion, a quantitative assessment report on children's creative and developmental levels is generated, providing teachers and parents with a scientifically based and visualized analysis of children's learning and growth. Simultaneously, an explanation of the scoring evidence is output to ensure the traceability, comprehensibility, and operability of the evaluation results.
[0092] Example 2: Resource scheduling in multi-server collaborative operation When a single physical server (i.e., a single edge server) cannot meet the demand under full load, the single physical server can be expanded. This expansion can be achieved by adding computing modules (such as graphics cards and CPUs) to the single physical server, or by adding more physical servers to utilize multiple edge servers. Regardless of the method, for this invention, it can be considered a logical expansion of computing resources. Here, the total computing resources can be divided into several virtual service resources, referred to as multi-server in this invention. This leads to the C / MS architecture, where C (Client) represents the smart desk and MS (Multi-Server) represents multiple servers. Therefore, the Server-only proximity collaboration algorithm based on the C / MS architecture is designed for situations where "clients do not directly connect or interact, and multiple servers collaborate based on client load requirements within their neighborhood." The client continues its session with the designated server (Server) without participating in any collaboration logic; multiple servers form an adjacency graph and perform two-stage actions in each time slice: ① Guarantee processed requests: Each server first guarantees computing resources for clients whose requests have been processed, ensuring 99% of the service level agreement (SLA) coverage.
[0093] ② Allocate surplus resources: Distribute some of the overload gap to servers with surplus resources in the neighborhood in a way that minimizes the "neighborhood coordination cost", avoiding centralized scheduling and global synchronization.
[0094] Here, the problem of balancing server computing resources scheduling between "gap and surplus" is abstracted into an entropy-regularized discrete optimal transport (EOT) algorithm, whose objective function is the Wasserstein distance under the adjacency cost matrix.
[0095] In summary, the Server-only neighbor cooperative algorithm employs the SLA EOT NS algorithm, which combines SLA guarantees, neighborhood entropy regularization for optimal transmission, and the neighborhood Sinkhorn-Knopp algorithm. The neighborhood Sinkhorn-Knopp algorithm is the core iterative algorithm for solving the entropy regularization optimal transmission problem. It does not require central control; adjacent servers only exchange local dual variables, converging to a stable transmission plan within the neighborhood. It features low traffic and scalability. Furthermore, the Server-only neighbor cooperative algorithm includes risk constraints and order preservation mechanisms. Specifically, it embeds Conditional Value at Risk (CVaR) regularization into the cost to explicitly suppress the risk of extreme tail latency / congestion; and it incorporates order preservation cooling to avoid load oscillations and frequent migrations.
[0096] The following is the design content of the Server-only proximity collaboration algorithm: Input: Adjacency graph G=(S,E), load of edge servers at each time slice t. computing power Service Level Agreement And parameters τ, λ, α, β, γ. Here, the adjacency graph G is a data structure composed of a set of virtual server nodes S and adjacency relations E; adjacency relations E are neighborhood graphs whose radius R can be defined by the same domain. This represents the real-time load of each virtual server node i; This represents the computing power possessed by each virtual server node i; τ represents the service level agreement for each virtual server node i; τ represents the entropy regularization strength parameter; λ represents the risk weight; α represents the probability threshold; β represents the confidence level of CVaR; and γ represents the smoothing factor.
[0097] Output: The migration execution table M = {(i→j, qty)}. Here, i and j represent the server node numbers; qty represents the number of virtual server nodes migrated from virtual server node i to virtual server node j.
[0098] Subsequently, the Server-only neighbor collaboration algorithm is executed according to the following process: 1) Perform a loop check on the computing resources of each server.
[0099] 2) Obtaining guaranteed capacity : By estimating the satisfaction Minimum capacity to obtain guaranteed capacity ,in Represents probability. Indicates a delay. Indicates the probability threshold. This represents the guaranteed capacity of each virtual server node i. That is: guaranteed capacity. It is to satisfy Greater than The probability is less than or equal to The minimum capacity under this condition.
[0100] 3) Obtain the elastic pool : Configure the elastic pool (strategy-based or adaptive). Among them, This represents the elastic pool for each virtual server node i.
[0101] 4) Through Identify the computing power gap ,in This represents the computing power gap for each virtual server node i. (Through...) Determine surplus computing power ,in This represents the surplus computing power of each virtual server node i.
[0102] 5) Calculate adjacency cost (RTT / bandwidth / queueing margin), and apply risk correction. The formula for calculating the corrected adjacency cost is as follows:
[0103] In the formula, Let be the transmission cost matrix, representing the adjacency cost before correction, which is the coordination cost of a unit task spreading from virtual server node i to virtual server node j, taking into account factors such as response latency, migration overhead, bandwidth, and queuing margin. This represents the corrected adjacency cost; Risk weights; Conditional risk value, This represents the change in the tail risk value after the load is migrated to node j.
[0104] 6) Based on the corrected adjacency cost The kernel matrix is calculated using the following formula:
[0105] In the formula, This represents the kernel matrix that diffuses from virtual server node i to virtual server node j. This represents the entropy regularization strength parameter. The larger the value, the more dispersed the solutions. When the value approaches 0, the solution approaches the traditional hard allocation.
[0106] 7) Initialization Synchronous iteration from k=1 to K. Wherein, and It is a pair of dynamically adjusted variables in the neighborhood Sinkhorn-Knopp algorithm, which iterates in a distributed manner across the entire adjacency graph, eventually converging to a set of values that both satisfies the constraints and is as close as possible to the values obtained from the Sinkhorn-Knopp algorithm. Define a cost-optimal load migration scheme.
[0107] 8) Each virtual server node i pulls data from its neighboring servers. And based on the computing power gap and kernel matrix Calculate the row scaling factor:
[0108] In the formula, This represents the row scaling factor, which indicates the pressure of unmet demand at demand-side node i. A value greater than 1 means that demand exceeds effective supply; if =1, then supply and demand are exactly in balance; if A value less than 1 means that demand is less than supply. Therefore, according to... renew .
[0109] 9) Each virtual server node j pulls data from its neighboring servers. And based on surplus computing power and kernel matrix Calculate column scaling factor:
[0110] In the formula, This represents the column scaling factor, which indicates the resource scarcity level of supply-side node j. A value greater than 1 means that supply exceeds demand; if =1, then supply and demand are exactly in balance; if A value less than 1 means that supply is less than demand. Therefore, according to... renew .
[0111] 10) Convergence check: If adjacent servers and When that time comes, it stops. Among them, Indicates the convergence precision, used to stop iteration; , .
[0112] 11) Calculate the optimal transmission plan:
[0113] In the formula, This indicates the optimal transmission plan.
[0114] 12) Calculate the execution cost of smooth migration:
[0115] In the formula, This represents the amount of time slice t that is used for smooth transition execution. This represents a smoothing factor that combines historical execution plans with the currently calculated optimal plan using an exponentially weighted moving average. This smoothing factor can enhance system stability, improve robustness, and reduce system overhead.
[0116] 13) Migration Execution and Rollback Mechanism: ① Execution Conditions: ≥ Execution granularity; ② Triggering rollback conditions: Migration is only issued to entries that meet the execution conditions, updating the ledger and monitoring; when the rollback condition is triggered, a two-hop or upper-level rollback is triggered.
[0117] In summary, this invention addresses the issues of weak information infrastructure, subjective teaching evaluation, and low resource utilization in early childhood STEM education, particularly in areas such as network management, computing power allocation, teaching assessment, and server collaboration. The integrated assessment and teaching smart desk described in this invention is not only suitable for early childhood STEM teaching, but its distributed networking, dynamic computing resource implementation plan, multimodal data fusion framework, and dynamic assessment model also possess horizontal and vertical scalability. Horizontally, it can be seamlessly migrated to more types of STEM experimental platforms, such as science inquiry experimental platforms, robot programming workstations, 3D printing creation platforms, virtual reality engineering training stations, and open electronic circuit experimental platforms, while vertically covering STEM teaching for different age groups. Furthermore, the containerized resource deployment mechanism allows the assessment system to be flexibly embedded into various smart education hardware ecosystems, providing a solid technical foundation for future expansion into home STEM education scenarios, community maker spaces, and even K-12 interdisciplinary teaching. Therefore, this invention possesses high commercial value and a promising market prospect.
[0118] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A control and management method for a smart desk integrating STEM assessment and teaching for preschoolers, characterized in that, Includes the following steps: (1) Set the network scale according to teaching needs and number of students. The smart desk analyzes the required resources based on the identified teaching aid type or selected target course role, and sends a resource request to the edge server to obtain the corresponding course role container image, so as to dynamically reorganize the teaching area and change the original teaching area to the target course teaching area. (2) The smart desk calculates the load fusion score through adaptive dynamic weight calculation of teaching scenarios, and adopts a two-level triggering mechanism to select to enter the warning state or the critical state, and executes resource pre-application and full resource application respectively; The edge server dynamically adjusts the attention weights of the smart desk's resource request characteristics and the edge server's state characteristics through a cross-attention mechanism, and then integrates them with the resource request characteristics to generate the optimal resource scheduling strategy. (3) Collect multimodal data, extract features and fuse them across modalities through a multimodal learning fusion model, obtain the scores, binary classification values and contribution of each modal data and natural language interpretation of the quantitative indicators of educational theory, so as to generate an evaluation report and output the evaluation basis and improvement suggestions. (4) Based on the Client / Multi-Serve architecture, a Server-only proximity collaboration algorithm is built, in which multiple edge servers form an adjacency graph to ensure the computing resources of the processed requests in each time slice. Then, the surplus resources are allocated to the overloaded server through the discrete optimal transmission algorithm with entropy regularization. The neighborhood Sinkhorn-Knopp algorithm is used to iteratively solve the optimal transmission plan and generate a smooth migration execution volume to perform the migration and realize resource scheduling of multi-server collaboration. The Server-only proximity collaboration algorithm specifically includes: In each time slice, based on the adjacency graph, the guaranteed capacity is obtained according to the service level agreement of the edge servers; the computing power gap and surplus computing power are determined based on the load, guaranteed capacity, and elastic pool of the edge servers; the transmission cost matrix of the servers is calculated as the adjacency cost, and risk correction is performed on it according to the risk weight parameter; the kernel matrix is calculated based on the corrected adjacency cost and the entropy regularization strength parameter; based on the kernel matrix, computing power gap, and surplus computing power, the neighborhood Sinkhorn-Knopp algorithm is used to iteratively solve the dynamic adjustment variable pairs; the optimal transmission plan is calculated based on the dynamic adjustment variable pairs and the kernel matrix; the smooth migration execution amount for this time slice is calculated based on the optimal transmission plan and the smoothing factor; the migration is executed based on the smooth migration execution amount to achieve resource scheduling for multi-server collaboration.
2. The control and management method according to claim 1, characterized in that, The smart desk is equipped with an edge computing device, a sound acquisition sensor, four high-definition cameras, and a line enclosure box. The line enclosure box contains a sub-router and a switch. The high-definition cameras and the sound acquisition sensor are connected to the switch, the switch is connected to the sub-router, the sub-router is connected to both the edge computing device and the main router, and the main router is connected to the edge server. The sound acquisition sensor is used to acquire voice signals to enable voice interaction; The high-definition camera is used to capture children's movement sequences and automatically identify the types of teaching aids placed on the table; The edge server is configured with a container pool, which includes a GPU instance pool, a course resource repository, and a model repository, and stores course role container images in a containerized manner. The edge computing device is used to realize local image recognition and voice interaction. It is used to analyze the required resources according to the type of teaching aid or the selected target course role, and send a resource request to the edge server to obtain the corresponding course role container image, so as to dynamically switch the smart desk to the target course teaching area.
3. The control and management method according to claim 1, characterized in that, The upper limit of the network scale is one network with eight desks, that is, eight smart desks are contained in the same local area network. The edge server is configured with a course resource repository and a model repository, which store course content, evaluation system, AI model, computing resources and storage space in a containerized manner to form a course role container image; The edge server is pre-loaded with container images of all course roles, so that after receiving a resource request from the smart desk, the corresponding course role container image can be distributed to the smart desk through elastic distribution.
4. The control and management method according to claim 1, characterized in that, The formula for calculating the load fusion score is as follows: In the formula, Indicates the load fusion score. Indicates local GPU utilization. Indicates the current delay. This indicates a delay in the target. Indicates the number of video frames to be processed. Indicates the maximum tolerable number of video frames. , , These represent the parameter weights for different teaching stages; The two-level triggering mechanism specifically includes: when the load fusion score is greater than or equal to a preset first threshold for a duration of not less than a first set duration, a warning state is entered and the smart desk performs a resource pre-request; when the load fusion score is greater than or equal to a preset second threshold for a duration of not less than a second set duration, a severe state is entered and the smart desk performs a complete resource request.
5. The control and management method according to claim 1, characterized in that, The generation of the optimal resource scheduling strategy specifically includes: The resource request features of the smart desk are encoded into an n-dimensional query vector Q, whose dimensions include scenario type, request level, course attributes, and resource requirements; the status features of the edge server are encoded into an m-dimensional key vector K and value vector V, whose dimensions include hardware utilization, container status, and cache information. The query vector Q, key vector K, and value vector V are fed into the cross-attention mechanism to calculate the attention weights; By fusing resource request characteristics with attention weights, an optimal resource scheduling strategy is generated, including container allocation instructions, resource reservation strategies, priority adjustment, and preloading of model lists.
6. The control and management method according to claim 1, characterized in that, The multimodal data includes children's postural key points, movement sequences, facial expressions and attention levels, language and dialogue, and artwork displays; The quantitative indicators of the educational theory include quantitative indicators of creation level and quantitative indicators of development level. The quantitative indicators of creation level include stability, structure, playability, completeness and narrative. The quantitative indicators of development level include health level, living level, exploration level, perception level and learning level.
7. The control and management method according to claim 1, characterized in that, Step (3) specifically includes: Collect multimodal data and preprocess it to obtain preprocessed multimodal data, including motion trajectories, visual images, speech text, and facial expression data; Based on the preprocessed multimodal data, key features of each modality are extracted by quantization transformation method and aggregated into manual feature vectors according to time windows; Construct a multimodal learning fusion model and collect training data samples for training. During the training process, with minimizing the total loss function as the optimization objective, adjust the parameters of the multimodal learning fusion model until the preset training rounds are reached to obtain the trained multimodal learning fusion model. The reasoning process utilizes a trained multimodal learning fusion model to obtain scores, binary classification values, and the contribution of each modality's data to the scores of each educational theory quantitative indicator, along with natural language interpretation, in order to generate an evaluation report and output evaluation criteria and improvement suggestions.
8. The control and management method according to claim 7, characterized in that, The multimodal learning-based fusion model includes a modal feature encoder, a temporal aggregation layer, a cross-modal attention fusion unit, and an output head. The modal feature encoder includes a motion trajectory encoder, a visual image encoder, a speech-text encoder, and an expression encoder. The motion trajectory encoder encodes preprocessed motion trajectories to obtain corresponding evidence vectors; the visual image encoder encodes preprocessed visual images to obtain corresponding evidence vectors; the speech-text encoder encodes preprocessed speech-text to obtain corresponding evidence vectors; and the expression encoder encodes preprocessed expression data to obtain corresponding evidence vectors. The temporal aggregation layer performs positional encoding and average pooling on the evidence vectors corresponding to each modality to obtain the temporal data for each modality. A temporal-aware token sequence is generated. The temporal-aware token sequence of each modality, along with handcrafted feature vectors, is input into a cross-modal attention fusion unit. Connection confidence graph tokens are inserted into the input temporal-aware token sequence to encode the topological information of visually detected connection points into graph vectors, which guide attention weight allocation. The final output is an aggregated vector. The output head includes a regression head, a classification head, and an interpretation head. The aggregated vector is input into these three heads respectively. The regression head outputs the scores of the quantitative indicators of educational theory. The classification head outputs the binary classification values of the quantitative indicators of educational theory. The interpretation head, through cross-analysis of attention weights and evidence vectors from each modality, outputs the contribution of each modality's data to the scores of the quantitative indicators of educational theory and provides a natural language interpretation. The total loss function is specifically calculated as follows: the score loss is calculated based on the predicted score of the quantitative indicator of educational theory output by the multimodal learning fusion model and its corresponding real score label; the explanatory consistency loss is calculated based on the contribution of each modal data output by the multimodal learning fusion model to the score of the quantitative indicator of educational theory and its corresponding real evidence; and the total loss function is obtained by weighted summation of the score loss and the explanatory consistency loss.
Citation Information
Patent Citations
Dynamic expansion video analysis desk based on edge calculation and intelligent identification method
CN114359816A
Classroom real-time analysis method and device based on containerization and Internet of Things technology and medium
CN121258741A