Classroom interaction analysis method and device, edge device and storage medium
By adjusting the projection weights of multimodal information and inputting it into a convolutional neural network, the problem of decreased accuracy in classroom interaction analysis is solved, and efficient classroom interaction analysis results are achieved, which is particularly suitable for edge devices.
Patent Information
- Application Number
- CN202510907535.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-17
AI Technical Summary
The accuracy of classroom interaction analysis based on multimodal information in existing technologies has decreased, and the contribution differences of data from different dimensions have not been considered.
By calculating the information entropy change rate of voice information and posture information, adjusting the projection weight of each modal information, and adjusting the projection weight of expression information according to lighting conditions, the interactive features are formed and input into the convolutional neural network for analysis.
It improves the accuracy of classroom interaction analysis, is suitable for edge devices with high time response requirements and limited computing power, and simplifies computational complexity.
Smart Images

Figure CN120808415A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data analysis, and in particular to a classroom interaction analysis method and device, an edge device and a storage medium. BACKGROUND
[0002] In the field of modern education, improving teaching effectiveness is one of the core goals of education, and the key lies in helping learners quickly and effectively master knowledge. The learning mood of students has a direct impact on the learning effect in classroom teaching. In the traditional offline teaching scene, teachers usually rely on intuitive observation to judge the learning mood of students, but lack systematic data support, making it difficult to comprehensively and accurately assess the emotional state of students and its relevance to the teaching process.
[0003] In the prior art, a classroom interaction analysis method based on multi-modal information fusion can be established to collect data in multiple dimensions such as student cognitive attention, learning mood, and physiological arousal in real time, monitor and analyze the learning state of students in the classroom, perfect the teaching process analysis, and improve the distinguishability of teaching effectiveness. However, the above method simply inputs multi-modal information as features into an AI model, and does not consider that data in different dimensions have different contribution differences, resulting in a decrease in the accuracy of classroom interaction analysis results. SUMMARY
[0004] The embodiments of the present application provide a classroom interaction analysis method, device, edge device and storage medium to solve the technical problem of decreased accuracy of classroom interaction analysis based on multi-modal information in the prior art.
[0005] In a first aspect, the embodiments of the present application provide a classroom interaction analysis method, comprising:
[0006] Collecting multi-modal data in the classroom, the multi-modal data including voice information, expression information and posture information;
[0007] Calculating a first information entropy change rate corresponding to the voice information, calculating a second information entropy change rate corresponding to the posture information when the first information entropy change rate exceeds a preset first change rate threshold, and adjusting the projection weight of the voice information according to the first information entropy change rate and adjusting the projection weight of the posture information according to the second information entropy change rate when the second information entropy change rate is in a preset second change rate threshold interval range;
[0008] Adjusting the projection weight of the expression information according to the light conditions of the classroom;
[0009] Projecting the voice information sequence, expression information sequence and posture information sequence to a feature plane based on the adjusted projection weights to form an interaction feature;
[0010] The interactive feature is input into the trained convolutional neural network to obtain a classroom interaction analysis result.
[0011] In a second aspect, the embodiment of the present application further provides a classroom interaction analysis device, comprising:
[0012] The collection module is configured to collect multi-modal data in the classroom, the multi-modal data comprising voice information, expression information and posture information.
[0013] The adjustment module is configured to calculate a first information entropy change rate corresponding to the voice information, calculate a second information entropy change rate corresponding to the posture information when the first information entropy change rate exceeds a preset first change rate threshold, and adjust a projection weight of the voice information according to the first information entropy change rate and adjust a projection weight of the posture information according to the second information entropy change rate when the second information entropy change rate is in a preset second change rate threshold interval range.
[0014] The adjustment module is configured to adjust the projection weight of the expression information according to a light condition of the classroom.
[0015] The projection module is configured to project the voice information sequence, the expression information sequence and the posture information sequence to a feature plane based on the adjusted projection weights to form an interactive feature.
[0016] The input module is configured to input the interactive feature into the trained convolutional neural network to obtain a classroom interaction analysis result.
[0017] In a third aspect, the embodiment of the present application further provides an edge device, comprising:
[0018] One or more processors;
[0019] A storage device configured to store one or more programs,
[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement the classroom interaction analysis method provided by the above embodiment.
[0021] In a fourth aspect, the embodiment of the present application further provides a storage medium containing computer executable instructions, which, when executed by a computer processor, are used to execute the classroom interaction analysis method provided by the above embodiment.
[0022] The classroom interaction analysis method, device, edge device and storage medium provided by the embodiments of the present invention collect multimodal data in the classroom, wherein the multimodal data includes voice information, expression information and posture information; calculates the first information entropy change rate corresponding to the voice information, and when the first information entropy change rate exceeds a preset first change rate threshold, calculates the second information entropy change rate corresponding to the posture information, and when the second information entropy change rate is within a preset second change rate threshold interval, adjusts the projection weight of the voice information according to the first information entropy change rate, and adjusts the projection weight of the posture information according to the second information entropy change rate; adjusts the projection weight of the expression information according to the lighting conditions of the classroom; projects the voice information sequence, the expression information sequence and the posture information sequence onto the feature plane based on the adjusted projection weight to form interaction features; and inputs the interaction features into the trained convolutional neural network to obtain classroom interaction analysis results. By counting the entropy change rate of voice information and the entropy change rate of posture information, it is possible to determine whether the current scene is a classroom interaction scene. When it is determined to be an interactive scene, the projection weight of the modal information that can reflect the degree of interaction is adjusted, and the features of each modality are jointly projected onto a feature plane according to the projection weight, so that the convolutional neural network can be used to quickly obtain the classroom interaction analysis results. The interaction analysis is performed using the adjusted projection weights of multimodal features, which can enhance the accuracy of the analysis and can be achieved without more complex calculations. It is particularly suitable for edge devices with high time response requirements and low computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0024] Figure 1 1 is a flow chart of a classroom interaction analysis method provided in Example 1 of the present invention;
[0025] Figure 2 is a structural diagram of a classroom interaction analysis device provided in Example 2 of the present invention;
[0026] Figure 3 This is a structural diagram of the edge device provided in Example 3 of the present invention. DETAILED DESCRIPTION
[0027] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0028] Example 1
[0029] Figure 1 is a flowchart of a classroom interaction analysis method provided by an embodiment of the present application. The embodiment can be applied to the analysis of classroom interaction. The method can be executed by a classroom interaction analysis device and integrated into an edge device. The method includes the following steps:
[0030] S110, collecting multi-modal data in a classroom, the multi-modal data including voice information, expression information, and posture information.
[0031] For example, the multi-modal data in the classroom can be obtained by distributed sensing devices. These distributed sensing devices can include a microphone array, a camera, etc., for collecting voice information, expression information, and posture information. In actual applications, the microphone array is arranged at different positions in the classroom to capture the voice signals of the students, thereby extracting voice information such as tone, speed, and voice intensity. The camera is used to capture the facial expressions and body movements of the students, from which expression features and posture features such as smile frequency, head tilt angle, and hand movement activity level are extracted. The collection of these multi-modal data needs to ensure that the entire classroom area is covered to fully reflect the emotional state and interaction behavior of the students. In order to achieve this goal, the arrangement of the distributed sensing devices needs to consider the spatial structure of the classroom and the seat distribution of the students, while avoiding incomplete data collection due to device obstruction or signal interference.
[0032] S120, calculating a first information entropy change rate corresponding to the voice information. When the first information entropy change rate exceeds a preset first change rate threshold, calculating a second information entropy change rate corresponding to the posture information. When the second information entropy change rate is in a preset second change rate threshold interval range, adjusting the projection weight of the voice information according to the first information entropy change rate, and adjusting the projection weight of the posture information according to the second information entropy change rate.
[0033] Information entropy is a basic concept of information theory. It describes the uncertainty of the occurrence of each possible event of an information source. In the classroom, most of the time is spent on teaching by the teacher, with classroom discussions and questions in between. Therefore, the change rate of voice information entropy can be used to determine whether it is a discussion or question session, or whether it is a transition from a question session to a discussion session.
[0034] Exemplarily, the entropy change rate of the voice information can be calculated in the following manner: dividing the sound signal into short-time frames; calculating the amplitude range in the short-time frames; dividing the amplitude range into multiple quantization spaces uniformly, and setting corresponding discrete values based on the quantization spaces; calculating the information entropy of each short-time frame according to the number of occurrences of the discrete values in the short-time frames; and calculating the information entropy change rate of the voice information based on the information entropy of the short-time frames. Compared with the traditional manner of calculating the energy of the power spectrum, the amplitude can better reflect the characteristics of the voice and filter out the influence of external noise. The quantization space can be a segment of the amplitude, and the discrete value can deviate from the mean value or the variance value by a large amount, i.e., a point exceeding a set value, so as to reflect the deviation degree of other elements as much as possible, so as to obtain a more accurate information entropy of the short-time frame. The information entropy of the time period is obtained based on the information entropy between the short-time frames, and the ratio of the difference between the information entropy values of different time periods to the information entropy value of the first time period is taken as the change rate. When the change rate exceeds a first change rate threshold, it can be determined that the current classroom interaction is relatively active. On this basis, a second information entropy change rate corresponding to the posture information is calculated. During the classroom discussion, the student's actions are relatively frequent, and the posture will change frequently along with the discussion focus and thinking. At the same time, frequent changes in posture can also indicate a poor classroom discipline, so the change rate needs to be comprehensively considered, and therefore needs to be limited within a reasonable space.
[0035] Exemplarily, the second information entropy change rate corresponding to the posture information can be calculated in the following manner:
[0036] The human body contour is segmented from the image, a local skeleton model is constructed based on the human body contour, a gray value joint probability of the local skeleton model of the human body contour is calculated by using a gray co-occurrence matrix, a second information entropy is calculated according to the gray value joint probability, and a second information entropy change rate is calculated according to the second information entropy of the image corresponding to a preset time length.
[0037] Since the collected image has complex components and is easily disturbed by the outside world, a large amount of noise is generated, and the information entropy thereof cannot be calculated. Therefore, in the present embodiment, the human body contour is first segmented from the image, and a local skeleton model is constructed according to the general height of the students. The gray co-occurrence matrix is a method for describing texture by studying the spatial correlation characteristics of gray values. The essence is to statistically calculate the joint probability distribution of two pixel points with a certain distance (d) and a certain direction (θ) in the image. The second information entropy is calculated by using the joint probability distribution of the gray values i and j, and the second information entropy change rate is calculated according to the second information entropy of the image corresponding to a preset time length.
[0038] After the second information entropy change rate is calculated, it can be determined whether the second information entropy change rate is in a preset interval range. Unlike the first information entropy change rate, if the second information entropy change rate changes too much, it is possible that the current is not a discussion state, and in addition, there can be strong noise interference. Therefore, it is necessary to determine whether it is in a preset interval, and when it is in the preset interval, it has the meaning of adjustment.
[0039] For example, the projection weight of the voice information is adjusted according to the first information entropy change rate, which includes: determining the corresponding first mapping projection weight according to the first information entropy change rate based on the mapping relationship; correspondingly, the projection weight of the posture information is adjusted according to the second information entropy change rate, which includes: determining the corresponding second mapping projection weight according to the second information entropy change rate based on the mapping relationship. That is, the preset mapping relationship can be used to set the corresponding mapping projection weight according to the change rate.
[0040] S130, adjusting the projection weight of the expression information according to the light condition of the classroom.
[0041] Similarly, the expression information also needs to be collected by the camera, and needs more detailed pictures. However, due to the influence of external conditions, for example: light condition, etc. Therefore, the projection weight of the expression information needs to be adjusted according to the light condition of the classroom.
[0042] S140, projecting the voice information sequence, the expression information sequence and the posture information sequence to the feature plane based on the adjusted projection weight to form interactive features.
[0043] Based on the collected voice information, expression information and posture information, the voice information sequence, expression information sequence and posture information sequence are generated. Taking the voice information sequence as an example, it can be projected to the feature plane in the following way:
[0044] The voice information is converted into a high-dimensional voice feature sequence; the high-dimensional voice feature sequence is uniformly divided into a plurality of voice feature blocks; a projection matrix corresponding to a randomly selected projection plane is weighted processed according to the attention weight of the first mapping projection weight to form a first weighted projection matrix; and the plurality of voice feature blocks are projected to a feature plane by using the first weighted projection matrix. The projection can be regarded as a matrix transformation. First, a projection plane is randomly selected, and the first mapping projection weight is used as a bias weight to perform weight bias processing on the projection plane according to the attention mechanism algorithm. Since the matrix corresponding to the projection plane is best fixed as a matrix of a basic size, the projection effect is better, so the high-dimensional voice feature sequence can be uniformly divided into a plurality of voice feature blocks to meet the matrix requirement corresponding to the projection plane, and the plurality of voice feature blocks are projected to the feature plane by using the matrix multiplication method. The posture information and the expression information can also be projected to the feature plane in the same way to form interactive features together.
[0045] In S150, the interactive features are input into the trained convolutional neural network to obtain a classroom interaction analysis result.
[0046] In the prior art, various complex models are usually needed for multi-modal information fusion, for example, encoding and decoding structures are needed to fuse the modal information. However, such a model needs stronger computing power and longer operation time. For edge computing devices, this is very difficult to achieve. Therefore, in the embodiment, the multi-modal information is projected by the above steps to map various features to the same plane, and the most basic convolutional neural network is used to obtain the evaluation result of the classroom interaction analysis by using the plane information through convolution, normalization and classification. Further, various possible analysis data can be given by the output result of the last convolution layer, which is convenient for intuitive presentation in the later stage.
[0047] The embodiment collects multi-modal data in a classroom, the multi-modal data including voice information, expression information and posture information; calculates a first information entropy change rate corresponding to the voice information, calculates a second information entropy change rate corresponding to the posture information when the first information entropy change rate exceeds a preset first change rate threshold, adjusts a projection weight of the voice information according to the first information entropy change rate when the second information entropy change rate is in a preset second change rate threshold interval range, adjusts a projection weight of the posture information according to the second information entropy change rate; adjusts a projection weight of the expression information according to a light condition of the classroom; projects the voice information sequence, the expression information sequence and the posture information sequence to a feature plane based on the adjusted projection weights to form interactive features; and inputs the interactive features into a trained convolutional neural network to obtain a classroom interaction analysis result. By statistically analyzing the entropy change rates of the voice information and the posture information, it can be determined whether the current scene is a classroom interaction scene, and when it is determined to be an interactive scene, the projection weights of the modal information that can reflect the degree of interaction are adjusted, and the features of each mode are projected to a feature plane according to the projection weights, so that the convolutional neural network can quickly obtain the classroom interaction analysis result. The use of the adjusted projection weights of the multi-modal features for interaction analysis can enhance the accuracy of the analysis, and the implementation does not require more complex operations, and is particularly suitable for edge devices with high time response requirements and weak computing power.
[0048] Embodiment two
[0049] Figure 2 is a structural schematic diagram of a classroom interaction analysis device provided by the embodiment two of the present application, as Figure 2 shown, the device comprises:
[0050] The acquisition module 210 is configured to acquire multi-modal data in a classroom, the multi-modal data including voice information, expression information and posture information.
[0051] The first adjustment module 220 is configured to calculate a first information entropy change rate corresponding to the voice information, calculate a second information entropy change rate corresponding to the posture information when the first information entropy change rate exceeds a preset first change rate threshold, adjust a projection weight of the voice information according to the first information entropy change rate when the second information entropy change rate is in a preset second change rate threshold interval range, and adjust a projection weight of the posture information according to the second information entropy change rate.
[0052] The second adjustment module 230 is configured to adjust a projection weight of the expression information according to a light condition of the classroom.
[0053] The projection module 240 is configured to project the voice information sequence, the expression information sequence and the posture information sequence to a feature plane based on the adjusted projection weights to form interactive features.
[0054] The input module 250 is configured to input the interaction feature into the trained convolutional neural network to obtain the classroom interaction analysis result.
[0055] The classroom interaction analysis device provided in the embodiment is configured to collect multi-modal data in a classroom, the multi-modal data including voice information, expression information and posture information; calculate a first information entropy change rate corresponding to the voice information, and when the first information entropy change rate exceeds a preset first change rate threshold, calculate a second information entropy change rate corresponding to the posture information, and when the second information entropy change rate is in a preset second change rate threshold interval range, adjust a projection weight of the voice information according to the first information entropy change rate and adjust a projection weight of the posture information according to the second information entropy change rate; adjust a projection weight of the expression information according to a light condition of the classroom; project the voice information sequence, the expression information sequence and the posture information sequence to a feature plane based on the adjusted projection weights to form an interaction feature; and input the interaction feature into a trained convolutional neural network to obtain a classroom interaction analysis result. By calculating the entropy change rates of the voice information and the posture information, it can be determined whether the current scene is a classroom interaction scene, and when it is determined to be an interaction scene, the projection weights of the modal information that can reflect the interaction degree are adjusted, and the features of various modalities are projected to a feature plane according to the projection weights, so that the convolutional neural network can quickly obtain the classroom interaction analysis result. The use of the adjusted projection weights of the multi-modal features for interaction analysis can enhance the accuracy of the analysis, and the analysis can be realized without more complex operations, which is particularly suitable for edge devices with high time response requirements and weak computing power.
[0056] On the basis of the above embodiments, the collection module is configured to acquire the multi-modal data in the classroom through a distributed sensing device.
[0057] On the basis of the above embodiments, the first adjustment module includes:
[0058] The segmentation unit is configured to segment the sound signal into short-time frames.
[0059] The amplitude range calculation unit is configured to calculate an amplitude range in the short-time frame.
[0060] The uniform division unit is configured to uniformly divide the amplitude range to obtain a plurality of quantization spaces, and set corresponding discrete values based on the quantization spaces.
[0061] The information entropy calculation unit is configured to calculate the information entropy of each short-time frame according to the number of occurrences of the discrete values in the short-time frame.
[0062] The information entropy change rate calculation unit is configured to calculate the information entropy change rate of the voice information based on the information entropy of the short-time frame.
[0063] On the basis of each of the above embodiments, the first adjustment module comprises:
[0064] Segmenting a human body contour from an image, and constructing a local skeleton model based on the human body contour;
[0065] Calculating a joint probability of a gray value of the human body contour using a gray co-occurrence matrix to construct the local skeleton model;
[0066] Calculating a second information entropy according to the joint probability of the gray value;
[0067] Calculating a second information entropy change rate according to the second information entropy of the corresponding image after a preset time length.
[0068] On the basis of each of the above embodiments, the adjustment module comprises:
[0069] The second determination module determines an installation carrier of the application program according to a maximum computing power and a storage space required by the application program.
[0070] On the basis of each of the above embodiments, the second determination module comprises:
[0071] The first determination unit is configured to determine a first mapping projection weight corresponding to the first information entropy change rate based on the mapping relationship;
[0072] The second determination unit is configured to determine a second mapping projection weight corresponding to the second information entropy change rate based on the mapping relationship.
[0073] On the basis of each of the above embodiments, the projection module comprises:
[0074] The conversion unit is configured to convert the voice information into a high-dimensional voice feature sequence;
[0075] The segmentation unit is configured to uniformly segment the high-dimensional voice feature sequence into a plurality of voice feature blocks;
[0076] The weighting processing unit is configured to perform weighting processing on a projection matrix corresponding to a randomly selected projection surface according to the first mapping projection weight to form a first weighted projection matrix;
[0077] The projection unit is configured to project the plurality of voice feature blocks to a feature plane using the first weighted projection matrix.
[0078] The classroom interaction analysis device provided in the embodiments of the present application can execute the classroom interaction analysis method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0079] Embodiment three
[0080] Figure 3This is a schematic diagram of the structure of an edge device provided in Example 3 of the present invention. This edge device can implement the cloud phone platform functions in the above embodiments. Figure 3 A block diagram of an exemplary edge device 12 suitable for implementing embodiments of the present invention is shown. Figure 3 The edge device 12 shown is merely an example and should not limit the functionality and scope of use of the embodiments of the present invention.
[0081] like Figure 3 As shown, edge device 12 is implemented as a general-purpose computing device. Components of edge device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 connecting various system components (including system memory 28 and processing unit 16).
[0082] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0083] The edge device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the edge device 12, including volatile and non-volatile media, removable and non-removable media.
[0084] The system memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The edge device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 3 Not shown, often called a "hard drive"). Although Figure 3 Not shown, a magnetic disk drive for reading and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.
[0085] Program / utility 40 having a set of program modules 42 can be stored in memory 28 by way of example, such program modules 42 include an operating system, one or more application programs, other program modules, and program data, each or some combination thereof, which may
[0086] Edge device 12 can also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with edge device 12; and / or any devices (e.g., network card, modem, etc.) that enable edge device 12 to communicate with one or more other computing devices. Such communication can occur via input / output (I / O) interface(s) 22. Still yet, edge device 12 can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), and / or the Internet, through network adapter 20. As depicted, network adapter 20 communicates with the other components of edge device 12 via bus 18. It should be appreciated that although not shown, other hardware and / or software modules could be used in conjunction with edge device 12. Such as, but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0087] Processing unit 16 executes various program applications and data processing by running programs stored in system memory 28, such as implementing the classroom interaction analysis method provided by embodiments of the present application.
[0088] Embodiment Four
[0089] Embodiment Four of the present application also provides a storage medium containing computer executable instructions, which when executed by a computer processor, are used to perform the classroom interaction analysis method provided by the above embodiments.
[0090] The computer storage medium of the embodiments of the present application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of the computer-readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0091] The computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such propagated data signal can take a variety of forms, including but not limited to electro-magnetic, optical or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can be used to carry or propagate program code that is used by or in connection with an instruction execution system, apparatus, or device.
[0092] The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the above.
[0093] The computer program code for carrying out operations of the present application can be written in one or more programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or edge device. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments of the present application, electronic mail (email) can be utilized as the distrusting mechanism to effectuate exchange of information.
[0094] Note that the above merely describes preferred embodiments of the present application and the principles of the technology applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, modifications and substitutions can be made without departing from the scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.
Claims
1. A classroom interaction analysis method, characterized in that: include: Collecting multimodal data in the classroom, wherein the multimodal data includes voice information, expression information, and posture information; Calculating a first information entropy change rate corresponding to the voice information, and when the first information entropy change rate exceeds a preset first change rate threshold, calculating a second information entropy change rate corresponding to the posture information, and when the second information entropy change rate is within a preset second change rate threshold interval, adjusting the projection weight of the voice information according to the first information entropy change rate, and adjusting the projection weight of the posture information according to the second information entropy change rate; Adjust the projection weight of facial expression information according to the lighting conditions in the classroom; Based on the adjusted projection weights, the speech information sequence, expression information sequence, and posture information sequence are projected onto the feature plane to form interactive features; The interaction features are input into the trained convolutional neural network to obtain the classroom interaction analysis results.
2. The method according to claim 1, characterized in that The multimodal data collected in the classroom includes: Acquire multimodal data in the classroom through distributed sensing devices.
3. The method according to claim 1, characterized in that Calculating a first information entropy change rate corresponding to the voice information includes: Split the sound signal into short time frames; Calculate the amplitude range in a short time frame; Multiple quantization spaces are obtained based on the even division of the amplitude range, and corresponding discrete values are set based on the quantization spaces; Calculate the information entropy of each short time frame according to the number of occurrences of discrete values in the short time frame; The information entropy change rate of the speech information is obtained based on the information entropy calculation of the short time frame.
4. The method according to claim 3, characterized in that The calculating the second information entropy change rate corresponding to the posture information includes: Segment the human body contour from the image and build a local skeleton model based on the human body contour; The gray-level co-occurrence matrix is used to calculate the joint probability of gray values of the human body contour to construct the local skeleton model; Calculating a second information entropy according to the gray value joint probability; The second information entropy change rate is calculated according to the second information entropy of the corresponding image after the preset time period.
5. The method according to claim 4, characterized in that The adjusting the projection weight of the speech information according to the first information entropy change rate includes: Based on the mapping relationship, determining a corresponding first mapping projection weight according to the first information entropy change rate; Accordingly, adjusting the projection weight of the posture information according to the second information entropy change rate includes: Based on the mapping relationship, a corresponding second mapping projection weight is determined according to the second information entropy change rate.
6. The method according to claim 1, characterized in that The projecting of the speech information sequence, the expression information sequence, and the posture information sequence onto the feature plane based on the adjusted projection weights includes: Convert speech information into a high-dimensional speech feature sequence; Evenly divide the high-dimensional speech feature sequence into multiple speech feature blocks; The projection matrix corresponding to the randomly selected projection surface is weighted according to the attention weight set by the first mapping projection weight to form a first weighted projection matrix; The plurality of speech feature blocks are projected onto a feature plane using a first weighted projection matrix.
7. A classroom interaction analysis device, characterized in that: include: A collection module for collecting multimodal data in the classroom, wherein the multimodal data includes voice information, expression information, and posture information; A first adjustment module is configured to calculate a first information entropy change rate corresponding to the voice information, and when the first information entropy change rate exceeds a preset first change rate threshold, calculate a second information entropy change rate corresponding to the posture information, and when the second information entropy change rate is within a preset second change rate threshold interval, adjust the projection weight of the voice information according to the first information entropy change rate, and adjust the projection weight of the posture information according to the second information entropy change rate; The second adjustment module is used to adjust the projection weight of the expression information according to the lighting conditions in the classroom; A projection module is used to project the voice information sequence, the expression information sequence and the posture information sequence onto the feature plane based on the adjusted projection weights to form interactive features; The input module is used to input the interaction features into the trained convolutional neural network to obtain the classroom interaction analysis results.
8. An edge device, characterized in that: The edge device includes: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the classroom interaction analysis method as described in any one of claims 1 to 6.
9. A storage medium comprising computer-executable instructions, wherein the computer-executable instructions are used to perform the classroom interaction analysis method according to any one of claims 1 to 6 when executed by a computer processor.