Intelligent blackboard interaction control method and system based on multi-modal interaction

Through multimodal data collection and spatial behavior map construction, combined with the LSTM-MLP model to predict teacher intentions and generate structured content trees, the problem of single signal perception, lagging control responses, and inefficient content organization of smart blackboard systems is solved, and accurate prediction of teacher intentions and intelligent self-organization and optimization display of blackboard content is achieved, which improves teaching efficiency and classroom interactivity.

CN120143989AActive Publication Date: 2025-06-13GUANGDONG ACAD OF EDUCATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510607565.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-13
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing smart blackboard system has shortcomings in terms of single signal perception, lagging control response, and inefficient content organization. It is difficult to accurately predict teachers' intentions in a dynamic teaching environment and realize intelligent self-organization and optimization display of blackboard content.

Method used

By collecting multimodal data, such as teacher's voice signals, gesture signals, location data, environmental information and blackboard content, a joint feature matrix is ​​constructed, and a spatial behavior map is constructed based on student classroom behavior, which is transformed into spatial behavior characteristics through graph convolutional layers. Then, through feature fusion and pre-constructed LSTM-MLP model, a probability distribution of teacher intention is generated, and the maximum probability intention is selected as the teacher intention. Based on the teacher's intention and spatial behavior characteristics, a structured content tree is generated, and combined with the teacher's intention to bind to the nodes in the content tree, and a control strategy vector is generated to realize the intelligent display and organization of blackboard content.

Benefits of technology

It realizes accurate prediction of teachers' real control intentions, actively triggers blackboard control logic, reduces teachers' operational burden, improves classroom interaction fluency and teaching efficiency, and gives blackboard system stronger active service capabilities and adapts to the changing teaching environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120143989A_ABST
    Figure CN120143989A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent blackboard interaction control method and system based on multi-modal interaction, and the method comprises the steps: collecting classroom multi-modal data, and constructing a joint feature matrix; collecting classroom behaviors of students, constructing a spatial behavior map, and converting the spatial behavior map into spatial behavior features through a map convolution layer; performing feature fusion on the joint feature matrix and the spatial behavior features to generate joint features, inputting the joint features into a pre-constructed LSTM-MLP, outputting teacher intention probability distribution, selecting a maximum probability intention as a teacher intention, and representing a real-time control intention of the teacher; generating a structured content tree according to the teacher intention and the spatial behavior characteristics; the teacher intention is bound with nodes in the structured content tree, meanwhile, the spatial behavior characteristics serve as triggering conditions, and a control strategy vector is generated; the intelligent blackboard can adapt to changeable teaching environments, and the intelligent level and the application value of the intelligent blackboard are comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of smart blackboard interactive control, and in particular to a smart blackboard interactive control method and system based on multimodal interaction. Background Art

[0002] With the rapid development of emerging educational models such as smart education and smart classrooms, smart blackboards, as an important part of information-based classrooms, have gradually become an indispensable interactive device in modern teaching. Smart blackboards integrate multiple functions such as touch, display, and handwriting recognition, which greatly improves teaching efficiency and classroom interactivity. However, the current interactive control methods of smart blackboards still have significant limitations. Traditional systems mostly rely on a single interactive signal input, such as controlling the blackboard through gesture recognition or issuing operation instructions through voice recognition. This single interactive mode is often difficult to cope with complex and changing dynamic scenes in the classroom in actual teaching environments. For example, during the teacher's lecture, frequent walking, changes in standing position, noise interference in voice signals, and the complexity of blackboard content will significantly reduce the system's accurate recognition of the teacher's true intentions. In addition, most existing systems adopt a passive response control strategy, which can only trigger corresponding actions after receiving clear control signals, making it difficult to achieve early perception and active service of teachers' teaching behaviors, resulting in limited classroom processes and low interaction efficiency. At the same time, the smart blackboard also faces another core problem, that is, the real-time organization and presentation of the blackboard content relies on manual management by teachers. Especially when the amount of information is large and the content structure is complex (such as multi-step formula derivation and multi-level knowledge point explanation), teachers must manually switch screens, partition displays, or manually collect and expand content, which can easily increase the operational burden and affect the teaching rhythm and fluency.

[0003] Therefore, how to integrate multimodal signals, accurately predict teacher intentions, and achieve intelligent self-organization and optimized display of blackboard content in a dynamic teaching environment has become a technical bottleneck that needs to be urgently solved in the current smart blackboard interactive control system. Summary of the invention

[0004] The purpose of the present invention is to design a smart blackboard interactive control method and system based on multimodal interaction, which can effectively overcome the shortcomings of the existing smart blackboard system in terms of single signal perception, delayed control response, and inefficient content organization.

[0005] In order to achieve the above object, the present invention provides a first aspect of a smart blackboard interactive control method based on multimodal interaction, the method comprising:

[0006] S1. Collect classroom multimodal data and construct a joint feature matrix; the joint feature matrix includes the teacher's voice signal, gesture signal, location data, environmental information and blackboard content; the blackboard content is the semantic context of the current classroom;

[0007] S2. Collect the classroom behaviors of students. Combine the joint feature matrix as graph nodes and use the interaction relationship between the teacher and the nodes as graph edges in the graph to construct a spatial behavior graph, and transform the spatial behavior graph into spatial behavior features through a graph convolutional layer;

[0008] S3. Perform feature fusion on the joint feature matrix and the spatial behavior features to generate joint features, input them into a pre-constructed LSTM-MLP, output the teacher intention probability distribution, and select the intention with the highest probability as the teacher intention to represent the real-time control intention of the teacher;

[0009] S4. Generate a structured content tree according to the teacher intention and the spatial behavior features, including: the tree nodes represent semantic types, and the tree edges represent semantic relationships, where each tree node carries the teacher interaction weight and the logical hierarchical relationship;

[0010] S5. Bind the teacher intention to the nodes in the structured content tree, and at the same time use the spatial behavior features as the trigger condition to generate a control strategy vector; for each tree node in the structured content tree, generate a control strategy vector according to the teacher intention, the current spatial position of the teacher, and the conflict degree of the previous interaction control instructions. Each element of the control strategy vector corresponds to an interaction operation suggestion for a content node.

[0011] Preferably, S4 specifically includes:

[0012] S41. Use the update of the blackboard content as the current blackboard content pool;

[0013] S42. For all semantic blocks in the current blackboard content, first determine the intention-content relevance, which is used to measure the semantic matching degree between the current teacher intention and the content theme represented by the semantic block;

[0014] S43. According to the teacher's position and interaction trajectory in the spatial behavior features, calculate the attention tendency value of the teacher to this content area in the spatial behavior;

[0015] S44. Calculate the weights according to the semantic matching degree and the attention tendency value, and establish a tree structure from high to low according to the weights, and retain the logical connection between the tree nodes;

[0016] S45. After the structured content tree is generated, record the display suggestion attributes of each tree node; among them, each tree node includes semantic type, weight, and logical hierarchical relationship.

[0017] Preferably, the blackboard content is extracted in real time through OCR and transformed into structured text data;

[0018] Among them, the signal-to-noise ratio of the multimodal data is calculated, and the dynamic weight of each multimodal data is calculated according to the signal-to-noise ratio; the higher the signal-to-noise ratio, the greater the weight of the signal in the overall decision-making.

[0019] Preferably, the spatial behavior map is constructed as follows:

[0020] Taking the spatial position information of the teacher as the main node coordinates for map initialization, the blackboard area coordinates and the student distribution are preset from the classroom layout map;

[0021] Define the spatial interaction map , where represents the set of entity graph nodes in the classroom, represents the set of spatial relationship edges between entities; among them, the weight of the graph edge is calculated as:

[0022] ;

[0023] Among them, are the spatial position coordinates of the graph nodes and in the blackboard area respectively; is the spatial diffusion coefficient, which controls the sensitivity of distance perception; is the interaction intensity between the teacher and other graph nodes; is the behavior interaction weight factor;

[0024] Among them, the spatial behavior characteristics are generated through the activation function.

[0025] Preferably, by performing feature fusion on the joint feature matrix and the spatial behavior characteristics, and adjusting based on the interaction intensity between the teacher and other graph nodes as the dynamic regularization factor, the joint feature is generated.

[0026] Preferably, the display suggestion attributes include whether folding is required and whether split-screen highlighting is required.

[0027] Preferably, the control strategy vector is generated as follows:

[0028] ;

[0029] Among them, is the control strategy output for the tree node ; is the matching degree between the intention and the semantic type of the tree node; is the spatial proximity between the teacher's current spatial position and the content node; is the degree of conflict with the previous control strategy; and , is a regulation factor for the trade-off logic of the control strategy; is the Sigmoid function that normalizes the strategy strength to 0 - 1.

[0030] Preferably, a control action sequence list is generated according to the control strategy vector for the display module of the intelligent blackboard to call;

[0031] Among them, when the output of the control strategy is greater than 0.7, the system assigns a display action to the tree node; when the output of the control strategy is less than 0.3, the system marks the tree node as hidden; the middle area retains the original state.

[0032] Preferably, the interaction operation suggestions include the target content node number, the suggested action type, and the trigger condition.

[0033] In the second aspect of the present invention, an intelligent blackboard interaction control system based on multimodal interaction is provided, and the system includes:

[0034] A multimodal data acquisition unit for collecting classroom multimodal data and constructing a joint feature matrix; the joint feature matrix includes the teacher's voice signal, gesture signal, position data, environmental information, and the blackboard content; the blackboard content is the semantic context of the current class;

[0035] A spatial behavior feature generation unit for collecting students' classroom behaviors, using the joint feature matrix as graph nodes, and the interaction relationship between the teacher and the nodes as graph edges in the graph to construct a spatial behavior graph, and converting the spatial behavior graph into spatial behavior features through a graph convolutional layer;

[0036] A teacher intention analysis unit for fusing the joint feature matrix and the spatial behavior features to generate joint features, inputting them into a pre-constructed LSTM - MLP, outputting the teacher intention probability distribution, and selecting the intention with the maximum probability as the teacher intention to represent the teacher's real-time control intention;

[0037] A structured content tree generation unit for generating a structured content tree according to the teacher intention and the spatial behavior features, including: the tree nodes represent semantic types, and the tree edges represent semantic relationships, where each tree node carries the teacher interaction weight and the logical hierarchical relationship;

[0038] An interaction suggestion generation unit for binding the teacher intention to the nodes in the structured content tree, and using the spatial behavior features as trigger conditions to generate a control strategy vector; for each tree node in the structured content tree, a control strategy vector is generated according to the teacher intention, the teacher's current spatial position, and the conflict degree of the previous interaction control instructions, and each element of the control strategy vector corresponds to an interaction operation suggestion for a content node.

[0039] The beneficial technical effects of the present invention are at least as follows:

[0040] By integrating multiple interaction signals, including but not limited to teachers' voice instructions, gesture actions, spatial position changes, classroom rhythm characteristics, etc., and combining the real-time perception of the blackboard historical content and the classroom space environment, the present invention establishes a multi-modal deep fusion teaching intention understanding mechanism to accurately predict the real control intention of teachers, and then actively triggers the control logic of the blackboard. At the same time, aiming at the problems of heavy blackboard content management burden and chaotic hierarchical structure, the present invention innovatively designs a dynamic self-organization and display optimization method for blackboard content, which can automatically complete content stratification, storage, folding and display layout adjustment according to teachers' intentions and teaching processes, reduce teachers' operation burden, and improve classroom interaction fluency and teaching efficiency. Through the linkage control strategy of multi-modal perception, intention prediction and content intelligent organization, the present invention not only effectively solves the problems of slow response, easy misjudgment and complex operation of traditional intelligent blackboard systems, but also endows the blackboard system with stronger active service capabilities, can adapt to changing teaching environments, and comprehensively improves the intelligent level and application value of intelligent blackboards. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The present invention is further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to the following drawings without creative efforts.

[0042] Figure 1 It is a flowchart of the intelligent blackboard interaction control method based on multi-modal interaction of the present invention.

[0043] Figure 2 It is a framework diagram of the intelligent blackboard interaction control system based on multi-modal interaction of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0045] In one or more embodiments, as Figure 1 shown, an intelligent blackboard interaction control method based on multi-modal interaction is disclosed, and the method includes the following steps one to five:

[0046] S1. Collect multi-modal data in the classroom and construct a joint feature matrix; the joint feature matrix includes the teacher's voice signal, gesture signal, location data, environmental information, and blackboard content; the blackboard content is the semantic context of the current class.

[0047] Specifically, in a smart classroom, the teacher's behavior and blackboard content are the core signal sources for multi-modal interaction. To ensure the accuracy of subsequent intention prediction and the rationality of content display, the goal of this step is to jointly collect various signals (voice, gesture, location, environmental noise) in the classroom with the current blackboard content and perform preprocessing to form standardized features that can support multi-modal interaction decisions. It is mainly divided into the following steps:

[0048] Signal collection:

[0049] Multi-modal sensors are equipped in the classroom for data collection, specifically including:

[0050] Voice signal : Collect the teacher's voice through a microphone array and perform real-time streaming input at a sampling rate of 16 kHz. If the environmental noise is large, enhance the signal quality through sound source localization technology;

[0051] Gesture signal : Collect the teacher's gestures through a camera, with an image resolution of , and sample 30 frames per second. Extract gesture features through a convolutional neural network (CNN);

[0052] Location data : Use an indoor positioning system to obtain the teacher's position coordinates in real time , and calibrate their positions in the classroom. The sampling frequency is 10 Hz;

[0053] Environmental signal : Use an environmental noise sensor to monitor the background noise in the classroom (in decibels). This data is sampled once per second to facilitate judging the reliability of the current signal.

[0054] Blackboard content collection : The blackboard content is extracted in real time through an OCR (Optical Character Recognition) and layout analysis module, which recognizes the text, charts, formulas, etc. written on the blackboard and converts them into structured text data. Through OCR, the present invention can extract the written content on the blackboard (such as formulas, graphics, or text blocks) and assign them semantic labels (such as "definition", "formula", "step").

[0055] For example, when the teacher writes " " on the blackboard, OCR will recognize the formula and mark it as a physical formula type, generating structured information.

[0056] Furthermore, in order to make different signal sources (such as voice, gesture, location, etc.) have appropriate contribution degrees in different environments, a dynamic signal calibration mechanism (D-SCAM) is introduced in this step:

[0057] Calculate the signal-to-noise ratio (SNR) for each collected signal (voice, gesture, location, environment), and calculate its dynamic weight according to the SNR of each signal. The higher the SNR, the greater the weight of the signal in the overall decision-making; the formula is as follows:

[0058] ;

[0059] where, represents the signal-to-noise ratio of the -th signal modality, and is the total number of signals. This weight is used to dynamically adjust the weights of different signals.

[0060] Example: If the SNR of the voice signal is 12 dB, the SNR of the gesture signal is 25 dB, and the SNR of the environmental noise is 30 dB, the calculated weights are: , , , so the system will give priority to relying on gesture and location data in a noisy environment and reduce the reliance on voice signals.

[0061] Furthermore, based on the preprocessing results of each signal, a joint feature matrix is generated as the input in subsequent steps (such as intention prediction). Specifically: Each modality (voice, gesture, location, environment, blackboard content) is standardized by a feature extraction module and combined with the SNR weight; the blackboard content provides the semantic context of the current classroom, while other signals (voice, gesture, location) provide the immediate behavior information of the teacher. They are fused through weighting to form a unified input feature matrix for subsequent teacher intention prediction.

[0062] Furthermore, the output joint feature matrix is used as the input in subsequent steps (such as spatial behavior map construction, intention prediction, etc.). This feature matrix contains the joint information of teacher behavior, classroom space, and blackboard content, providing a complete context for intelligent decision-making.

[0063] S2. Collect students' classroom behaviors, use the combined feature matrix as graph nodes, and use the interaction relationship between the teacher and the nodes as graph edges in the graph to construct a spatial behavior graph, and transform the spatial behavior graph into spatial behavior features through a graph convolutional layer.

[0064] Specifically, this step is in the bridging position between "signal perception" and "intention prediction" in the overall patent solution. The main task is based on the combined feature matrix output in step 1 , fuse the dynamic behavior information in the classroom space, and construct a spatial behavior graph , providing a global spatial and interaction context for subsequent intention reasoning and blackboard content organization. Especially in a dynamic spatial environment such as a smart classroom, the teacher's behavior is restricted by spatial positions (such as the podium area, the blackboard area, the student area) and classroom behavior relationships (such as student interaction, group discussion, etc.). Therefore, the spatial behavior must be modeled as a dynamic feature input to effectively support the goal of "active control".

[0065] Furthermore, this step receives the combined feature matrix output from the previous step , where includes the teacher's voice signal , gesture signal , position data , environmental information and blackboard content ; these multi-modal data simultaneously contain the teacher's action features (gestures, voice, position information) and the context of the current teaching content (blackboard structure information ), but lack the overall modeling of "spatial dynamics" and "multi-agent interaction behavior".

[0066] Furthermore, the construction of the spatial behavior graph :

[0067] This step introduces a multi-agent spatial behavior graph modeling mechanism (MSGM mechanism), taking key entities such as "teacher, student, blackboard, podium" in the classroom as nodes, and the interaction relationship between the teacher and other entities (such as relative distance, attention direction, interaction actions) as edges in the graph;

[0068] The teacher's spatial position information is used as the main node coordinates for graph initialization, and the blackboard area coordinates and student distribution are preset from the classroom layout map;

[0069] Define the spatial interaction graph , where represents the set of entity graph nodes in the classroom, Represents the set of spatial relationship edges between entities. The weight of the graph edge is calculated by the following spatial perception function:

[0070] ;

[0071] where are the spatial position coordinates of graph nodes and respectively;

[0072] is the spatial diffusion coefficient, which controls the sensitivity of distance perception;

[0073] is the interaction intensity between the teacher and other graph nodes, such as whether the teacher is talking to the student or writing on the blackboard facing forward, and the value range is ;

[0074] is a custom behavior interaction weight factor, which is specially designed in the intelligent classroom scenario to strengthen the joint modeling of spatial distance and behavioral actions.

[0075] It can be understood that, compared with the traditional GNN graph structure, the graph edge weight in this graph spectrum introduces the interaction intensity as a regularization factor, which solves the problem that a single "physical distance" graph spectrum is difficult to reflect the "interaction intention intensity".

[0076] Further, the feature embedding of the spatial behavior graph spectrum:

[0077] Based on the above constructed , design a spatial context embedding model (SCEM model), and transform the graph structure into spatial behavior features through a graph convolutional layer (GCN) , and the output form is as follows:

[0078] ;

[0079] where represents the input feature vector of graph node in the graph spectrum (that is, the signal feature associated with node in );

[0080] is the learned graph convolutional weight matrix;

[0081] is the edge weight normalization factor of graph node ;

[0082] is the set of neighbors of the graph node .

[0083] Understandably, through this feature embedding, the system deeply fuses the "teacher behavior signal" and the "spatial relationship" to generate spatial behavior features that can be directly used for subsequent intention prediction .

[0084] The finally output includes information such as the interaction pattern, spatial dynamics, and distance relationship between the teacher and entities in the space (such as the blackboard, students, podium), and serves as the direct input for the subsequent step (intention prediction); has the following dimensional features: behavior signal features (teacher's voice, gesture, position) + spatial relationship features (dynamic edge weights with the blackboard, students, podium) + behavior interaction features ( current interaction activity reflected).

[0085] S3. Feature fusion is performed on the joint feature matrix and the spatial behavior features to generate joint features, which are input into a pre-constructed LSTM-MLP, and the teacher intention probability distribution is output. The intention with the highest probability is selected as the teacher intention, representing the teacher's real-time control intention

[0086] Specifically, the goal of this step is based on the multi-modal signal features output in step 1 and the spatial behavior features output in step 2 , jointly predicting the teacher's classroom control intention , as the key input for the subsequent active interaction control of the smart blackboard. This step serves as the "behavior decision layer" in the patent link, focusing on "understanding the teacher's intention" and not involving blackboard content processing or display

[0087] Further, the input is (teacher's voice, gesture, position signals and blackboard content features) and (spatial relationship features between the teacher and the blackboard, students, podium, etc.). The output is the teacher's control intention , such as labels like "prepare to write on the blackboard", "unfold new knowledge points", "clear part of the writing on the blackboard", etc

[0088] Further, a JIPM+ intention prediction model exclusive to smart classrooms is designed. The core innovation of this model lies in:

[0089] Spatial-behavior dynamic embedding module, which deeply fuses and features;

[0090] The classroom interaction intensity regularization term introduces the "interaction intensity" between the teacher and spatial entities as a dynamic adjustment factor in the intention prediction process, which is different from the general intention prediction model.

[0091] Furthermore, and generate joint features through the fusion layer , and introduce the spatial interaction intensity matrix (output from the spatial behavior graph in step 2, reflecting the interaction intensity between the teacher and the blackboard, students, and podium) as the dynamic regularization factor:

[0092] ;

[0093] Among them, is the fusion weight of each modal feature, is the feature fusion operation (such as concatenation or MLP mapping); is the "classroom interaction regularization coefficient" designed specifically for the smart classroom scenario, used to dynamically amplify or suppress the influence of spatial relationships on intention prediction; reflects the current interaction intensity between the teacher and other entities, such as "whether on the podium", "whether close to the students", etc.

[0094] Furthermore, input the fused feature into the LSTM-MLP structure to capture the temporal features and spatial dynamic changes of the teacher's behavior, and output the teacher intention probability distribution , and select the intention with the highest probability as the final result.

[0095] A "spatial dependence gating mechanism" is specially designed to dynamically adjust the influence of the spatial feature on the current intention state in the LSTM, ensuring that the output of the intention prediction model can be automatically optimized in different scenarios of "the teacher is close to the podium" and "the teacher is far from the blackboard".

[0096] Finally, output , representing the teacher's real-time control intention (such as "split screen", "new blackboard writing", "interaction mode", etc.), for direct invocation in subsequent steps (blackboard content organization and interaction control).

[0097] S4. Generate a structured content tree according to the teacher intention and spatial behavior characteristics, including: tree nodes represent semantic types, and tree edges represent semantic relationships, where each tree node carries the teacher interaction weight and logical hierarchical relationship.

[0098] Specifically, the goal of this step is to, based on the teacher intention output in step 3 , generate a representation of the blackboard content structure that matches the current teaching scenario . This step serves as the "knowledge organization layer", and the result is the "content display logic structure" understood by the system under the teacher's intention. For example: What are the key points? Which is the blackboard writing area? Is structural stratification (definition - derivation - conclusion) required, etc.

[0099] Furthermore, input: the teacher's current intention (such as "explain the definition", "display the chart", "expand on old content");

[0100] Teacher's spatial behavior characteristics (such as the teacher approaching the blackboard area, moving towards the student area);

[0101] Output: a structured content tree , representing the organizational units, logical levels, and spatial correlations of the content, for subsequent control strategy judgment.

[0102] Furthermore, the content structure generation framework (CTG mechanism): The present invention designs a content structure generation module based on intention - space drive (CTG: Contextualized Tree Generator) for dynamically generating a semantic structure tree.

[0103] Specifically, the system maintains an updatable "current blackboard content pool" , which includes the structured blackboard writing information (such as the text and image blocks after OCR recognition) extracted in step 1. Based on this, the present invention constructs a content semantic tree , with the structure as follows:

[0104] The tree nodes represent content units (such as definitions, examples, diagrams, derivations, etc.);

[0105] The tree edges represent semantic relationships (such as "belongs to", "derived from", "exemplified by", etc.);

[0106] Each node carries a teacher interaction weight , and and are used to calculate its current importance and display priority.

[0107] Furthermore, the structure weight calculation formula (core modeling formula):

[0108] To dynamically adjust the importance of each content unit in the structure, the present invention designs the following weight assignment mechanism:

[0109] ;

[0110] : content node Display weight in the structure tree;

[0111] : Intent-content relevance function, which measures the semantic matching degree between the current teacher's intent and the content theme represented by the semantic block;

[0112] : The attention tendency value of the teacher to this content area in spatial behavior, such as approaching this block, facing this area, etc.;

[0113] and : Adjustable coefficient, used to balance the structural modeling influence of intent guidance and spatial guidance.

[0114] Further, the structure tree generation process:

[0115] For all semantic blocks in, first calculate through means such as keyword matching and semantic embedding matching ;

[0116] Then, according to the teacher's position and interaction trajectory in, calculate ;

[0117] Calculate the weight comprehensively , establish a tree structure from high to low according to the weight, and retain the logical connection between tree nodes (such as: is 's extended description, etc.);

[0118] Structure tree After generation, record the display suggestion attributes of each tree node (such as: whether to fold, whether to split screen and highlight).

[0119] Output: The output is the content structure tree , where each tree node is marked with:

[0120] Semantic type (definition / formula / legend / interactive block, etc.);

[0121] Weight (affecting the subsequent display priority);

[0122] Logical hierarchical relationship (used for subsequent expansion, folding, and grouped display);

[0123] It will be used as the basis for executing control actions in step 5 to ensure that the content display is highly consistent with the teacher's intent and spatial state.

[0124] S5. Bind the teacher's intention to the nodes in the structured content tree, and at the same time use the spatial behavior characteristics as the trigger condition to generate a control strategy vector; for each tree node in the structured content tree, generate a control strategy vector according to the teacher's intention, the teacher's current spatial position, and the conflict degree of the previous interaction control instruction. Each element of the control strategy vector corresponds to an interaction operation suggestion for a content node.

[0125] Specifically, this step aims to establish a control strategy model to guide how the intelligent blackboard system generates an interactive control strategy according to the teacher's current teaching intention , spatial behavior state , and content structure tree actively, and transmit it to the execution module for display and feedback.

[0126] Furthermore, the input variables are: : The teacher's current intention; : The teacher's spatial behavior context; : The content structure tree, which contains all teaching content units and their hierarchical structures;

[0127] The goal of this step: Output a control strategy for the instruction execution system to complete the display behavior.

[0128] Furthermore, the present invention designs an "Intention-Space Collaborative Control Generation Model" (AICG: ActionIntent-Control Generator) to bind the intention with the nodes in the structure , and at the same time consider the spatial behavior context as the trigger condition to generate an interactive control instruction.

[0129] Furthermore, for each tree node in the structure tree , the system determines whether it is the current control focus according to . The system defines the following control strategy function:

[0130] ;

[0131] : The control strategy output for the tree node (the value is 0-1, indicating the confidence of action execution);

[0132] : The matching degree between the intention and the semantic type of the tree node (e.g., = display formula, then it highly matches the formula type node);

[0133] : The spatial proximity between the teacher's current spatial position and the content node (such as distance, orientation);

[0134] : The degree of conflict with the previous control strategy (such as repeated highlighting, overlapping display);

[0135] 、 、 : Adjustment factor, the trade-off logic of the control strategy;

[0136] : Sigmoid function, normalizing the strategy strength to 0 - 1.

[0137] Furthermore, strategy decoding and action definition:

[0138] When , the system assigns the "display" action to the tree node ;

[0139] When , the system marks the tree node as "hidden";

[0140] The middle area is "retain the original state".

[0141] The system outputs a list of control action sequences according to the entire sequence for the actuator (display module) to call.

[0142] The final output is the control strategy vector , and each element corresponds to an interaction operation suggestion for a content node; each interaction operation suggestion includes: target content node number, suggested action type (display / fold / highlight, etc.), trigger conditions (spatial behavior state, intent label);

[0143] Example output:

[0144] U_c = [{node: "t_3", action: "expand", trigger: "I_c=expand explanation && F_2=in front of the blackboard"},{node: "t_7", action: "hide", trigger: "I_c=switch topic && F_2=lectern area"},... ].

[0145] In one or more embodiments, as Figure 2 shown, a smart blackboard interaction control system based on multimodal interaction is disclosed, and the system includes:

[0146] The multimodal data acquisition unit 101 is used to acquire classroom multimodal data and construct a joint feature matrix; the joint feature matrix includes the teacher's voice signal, gesture signal, position data, environmental information, and blackboard content; the blackboard content is the semantic context of the current class.

[0147] The spatial behavior feature generation unit 102 is used to collect students' classroom behaviors, combine the joint feature matrix as graph nodes, use the interaction relationship between the teacher and the nodes as graph edges in the graph, construct a spatial behavior graph, and convert the spatial behavior graph into spatial behavior features through a graph convolutional layer.

[0148] The teacher intention analysis unit 103 is used to fuse the joint feature matrix and spatial behavior features to generate joint features, input them into a pre-constructed LSTM-MLP, output the teacher intention probability distribution, and select the intention with the highest probability as the teacher intention, representing the teacher's real-time control intention.

[0149] The structured content tree generation unit 104 is used to generate a structured content tree according to the teacher intention and spatial behavior features, including: the tree nodes represent semantic types, and the tree edges represent semantic relationships, where each tree node carries the teacher interaction weight and logical hierarchical relationship.

[0150] The interaction suggestion generation unit 105 is used to bind the teacher intention to the nodes in the structured content tree, and at the same time use the spatial behavior features as the trigger condition to generate a control strategy vector; for each tree node in the structured content tree, generate a control strategy vector according to the teacher intention, the teacher's current spatial position, and the conflict degree of the previous interaction control instructions, and each element of the control strategy vector corresponds to an interaction operation suggestion for a content node.

[0151] It should be noted that the specific working process of the intelligent blackboard interaction control system based on multimodal interaction provided in the embodiments of the present invention is the same as the process of the intelligent blackboard interaction control method based on multimodal interaction described in the above embodiments, and will not be elaborated here.

[0152] The embodiments of the present invention also provide an intelligent blackboard interaction control device based on multimodal interaction, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps in the embodiments of the intelligent blackboard interaction control method based on multimodal interaction as described above, such as Figure 1 the steps S1~S5 described therein; or, when the processor executes the computer program, it implements the functions of each module in the above system embodiments.

[0153] Exemplarily, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the intelligent blackboard interaction control device based on multimodal interaction.

[0154] The intelligent blackboard interaction control device based on multimodal interaction may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The intelligent blackboard interaction control device based on multimodal interaction may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the intelligent blackboard interaction control device based on multimodal interaction may further include input / output devices, network access devices, a bus, etc.

[0155] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the intelligent blackboard interaction control device based on multimodal interaction, and uses various interfaces and lines to connect various parts of the entire intelligent blackboard interaction control device based on multimodal interaction.

[0156] The memory may be used to store the computer program and / or modules. The processor realizes various functions of the intelligent blackboard interaction control device based on multimodal interaction by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the operation of the air conditioner controller, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0157] Among them, if the modules integrated in the intelligent blackboard interaction control device based on multimodal interaction are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0158] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above-mentioned embodiment methods, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the storage medium can be a magnetic disk, optical disc, read-only memory (Read-Only Memory, ROM), or random access memory (Random Access Memory, RAM), etc.

[0159] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A smart blackboard interactive control method based on multimodal interaction, characterized in that: The method comprises: S1. Collect classroom multimodal data and construct a joint feature matrix; the joint feature matrix includes the teacher's voice signal, gesture signal, location data, environmental information and blackboard content; the blackboard content is the semantic context of the current classroom; S2. Collect students' classroom behaviors, combine the joint feature matrix as graph nodes, use the interaction between teachers and nodes as graph edges in the graph, construct a spatial behavior graph, and convert the spatial behavior graph into spatial behavior features through a graph convolution layer; S3, performing feature fusion on the joint feature matrix and the spatial behavior feature to generate a joint feature, inputting the joint feature into the pre-built LSTM-MLP, outputting the probability distribution of the teacher's intention, selecting the maximum probability intention as the teacher's intention, and indicating the teacher's real-time control intention; S4, generating a structured content tree according to the teacher's intention and spatial behavior characteristics, including: tree nodes representing semantic types, tree edges representing semantic relationships, wherein each tree node carries a teacher interaction weight and a logical hierarchical relationship; S5. Bind the teacher's intention with the nodes in the structured content tree, and use the spatial behavior characteristics as trigger conditions to generate a control strategy vector; for each tree node in the structured content tree, generate a control strategy vector based on the teacher's intention, the teacher's current spatial position, and the degree of conflict of the previous interactive control instructions, and each element of the control strategy vector corresponds to an interactive operation suggestion for a content node.

2. According to claim 1, the intelligent blackboard interactive control method based on multimodal interaction is characterized in that: The S4 specifically includes: S41, taking the update of the blackboard content as the current blackboard content pool; S42, for all semantic blocks in the current blackboard content, first determine the intention-content relevance, which is used to measure the semantic matching degree between the current teacher's intention and the content theme represented by the semantic block; S43, calculating the teacher's attention tendency value for the content area in the spatial behavior according to the teacher's position and interaction trajectory in the spatial behavior characteristics; S44, calculating weights according to the semantic matching degree and the attention tendency value, and establishing a tree structure from high to low according to the weights, and retaining the logical connections between the tree nodes; S45. After the structured content tree is generated, the display suggestion attribute of each tree node is recorded; wherein each tree node includes a semantic type, a weight, and a logical hierarchical relationship.

3. According to claim 1, the intelligent blackboard interactive control method based on multimodal interaction is characterized in that: The blackboard content is extracted in real time through OCR and converted into structured text data; The signal-to-noise ratio of the multimodal data is calculated, and the dynamic weight of each multimodal data is calculated according to the signal-to-noise ratio; the highest the signal-to-noise ratio, the greatest weight of the signal in the overall decision.

4. According to claim 1, the intelligent blackboard interactive control method based on multimodal interaction is characterized in that: The spatial behavior map is constructed as follows: The spatial location information of the teacher is used as the main node coordinates for initializing the graph, and the blackboard area coordinates and student distribution are preset from the classroom layout map; Defining a spatial interaction graph ,in Represents the set of entity graph nodes in the classroom, Represents the set of spatial relationship edges between entities; the weight of the graph edge Calculated as: ; in, They are the graph nodes in the blackboard area. and The spatial position coordinates of is the spatial diffusion coefficient, which controls the sensitivity of distance perception; is the interaction strength between the teacher and other graph nodes; is the behavior interaction weight factor; The spatial behavior characteristics are Activation function generation.

5. According to claim 4, the intelligent blackboard interactive control method based on multimodal interaction is characterized in that: The joint feature is generated by fusing the joint feature matrix and the spatial behavior feature, and adjusting it based on the interaction strength between the teacher and other graph nodes as a dynamic regularization factor.

6. The intelligent blackboard interactive control method based on multimodal interaction according to claim 2 is characterized in that: The display suggestion attributes include whether folding is required and whether split-screen highlighting is required.

7. The intelligent blackboard interactive control method based on multimodal interaction according to claim 2 is characterized in that: The control strategy vector is generated as follows: ; in, For tree nodes The control strategy output; is the matching degree between the intention and the semantic type of the tree node; is the spatial proximity between the teacher's current spatial position and the content node; is the degree of conflict with the previous control strategy; , , The trade-off logic of the control strategy is the adjustment factor; is the Sigmoid function, which normalizes the strategy strength to 0~1.

8. The intelligent blackboard interactive control method based on multimodal interaction according to claim 7 is characterized in that: Generate a control action sequence list according to the control strategy vector for calling by the display module of the smart blackboard; When the control strategy output is greater than 0.7, the system assigns a display action to the tree node; when the control strategy output is less than 0.3, the system marks the tree node as hidden; and the middle area retains the original state.

9. The intelligent blackboard interactive control method based on multimodal interaction according to claim 8 is characterized in that: The interactive operation suggestion includes a target content node number, a suggested action type, and a trigger condition.

10. The intelligent blackboard interactive control system based on multimodal interaction is characterized by: The system comprises: A multimodal data acquisition unit is used to collect multimodal data of the classroom and construct a joint feature matrix; the joint feature matrix includes the teacher's voice signal, gesture signal, position data, environmental information and blackboard content; the blackboard content is the semantic context of the current classroom; A spatial behavior feature generation unit is used to collect students' classroom behaviors, combine the joint feature matrix as graph nodes, use the interaction between teachers and nodes as graph edges in the graph, construct a spatial behavior graph, and convert the spatial behavior graph into spatial behavior features through a graph convolution layer; The teacher intention analysis unit is used to fuse the joint feature matrix and the spatial behavior feature to generate a joint feature, input the joint feature into the pre-built LSTM-MLP, output the probability distribution of the teacher intention, select the maximum probability intention as the teacher intention, and represent the teacher's real-time control intention; A structured content tree generating unit, used to generate a structured content tree according to the teacher's intention and spatial behavior characteristics, including: tree nodes representing semantic types, tree edges representing semantic relationships, wherein each tree node carries a teacher interaction weight and a logical hierarchical relationship; The interactive suggestion generating unit is used to bind the teacher's intention with the nodes in the structured content tree, and use the spatial behavior characteristics as trigger conditions to generate a control strategy vector; for each tree node in the structured content tree, a control strategy vector is generated according to the teacher's intention, the teacher's current spatial position and the degree of conflict of the previous interactive control instructions, and each element of the control strategy vector corresponds to an interactive operation suggestion for a content node.

Citation Information

Patent Citations

  • Virtual-real fusion teaching aid automatic generation method

    CN112230772A

  • Teaching video live broadcast control method based on mobile network

    CN118138794A

  • Multi-mode interaction control method and system for smart blackboard

    CN118655979A

  • Intelligent classroom paper screen synchronous control method and system

    CN118692268A

  • Information pushing method and apparatus based on human-computer interaction, and computer device

    WO2021068321A1