Student recognition method, apparatus, and device based on speech and text classification
By combining voiceprint recognition and text conversion technologies with deep learning, the system accurately evaluates students' cognitive abilities during classroom discussions, solving the problem of teachers struggling to grasp students' discussion progress and improving the teaching effectiveness of discussion-based courses.
Patent Information
- Application Number
- CN202210905870.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2042-07-29
AI Technical Summary
In discussion-based classrooms, teachers often struggle to keep track of each group's discussion progress. Random grouping can lead to varying cognitive levels among members, affecting student participation and critical thinking. Furthermore, students' states and personalities may not be conducive to in-depth communication.
By employing voiceprint recognition and text conversion technologies, and combining voice data and classroom text data with deep learning methods, we can accurately evaluate students' cognitive abilities and construct a chart showing changes in cognitive levels.
It provides teachers with comprehensive information about students' cognitive abilities, improves the teaching quality of discussion-based courses, and helps teachers adjust their teaching strategies in a timely manner.
Smart Images

Figure CN115358300B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a student cognitive recognition method and device based on voice and text classification, equipment and storage medium. BACKGROUND
[0002] Discussion teaching is a widely used teaching mode in middle school classroom. In the discussion with students as the main body, students can exert their subjective initiative and actively construct their own cognition through a series of thinking activities. At the same time, classroom discussion can establish closer contact between students and teachers, improve classroom performance, help students explore potential, promote the development of students' healthy character, and realize the joint innovation of teaching and learning.
[0003] However, there are many deficiencies in the discussion class that need to be solved. First, the huge difference between the number of teachers and students exists all over the world. In the discussion class, the teacher cannot grasp the discussion status of each group member in time, and cannot give help in time. Secondly, due to the random grouping of group members in the discussion class, the cognitive level of group members will affect the discussion participation and thinking degree of students, which will eventually lead to poor effect of discussion class. Thirdly, there are many factors, such as poor state of students, difficult to adapt between students, and difficult to support students to speak actively, which lead to the situation that students cannot communicate and learn deeply. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a student cognitive recognition method and device based on voice and text classification, equipment and storage medium, which applies voiceprint recognition technology and text conversion technology to the discussion classroom, obtains the identity of the student corresponding to the voice data and the classroom text data, and uses deep learning method to accurately and efficiently evaluate the cognitive situation of students in the discussion classroom, provides more comprehensive information for teachers, and helps to improve the teaching quality of future discussion courses.
[0005] In a first aspect, the embodiments of the present application provide a student cognitive recognition method based on voice and text classification, comprising the following steps:
[0006] Obtain the voice data set of each student in the discussion classroom, wherein the voice data set includes voice data at different time moments;
[0007] Input the voice data set of each student into a preset voiceprint recognition model to obtain voiceprint recognition data corresponding to the voice data at different time moments, and obtain the identity of the student corresponding to the voice data at different time moments according to the voiceprint recognition data and a preset voiceprint feature library;
[0008] input the speech data set of each student into a preset text conversion model to obtain classroom text data corresponding to the speech data at each different time point;
[0009] input the classroom text data corresponding to the speech data at each different time point into a preset cognitive recognition model to obtain cognitive level data corresponding to the speech data at each different time point;
[0010] obtain cognitive change conditions of each student on the discussion classroom according to the student identity corresponding to the speech data at each different time point and the cognitive level data corresponding to the speech data at each different time point.
[0011] In a second aspect, an embodiment of the present application provides a student cognitive level recognition device based on speech and text classification, comprising:
[0012] a speech data obtaining module configured to obtain a speech data set of each student on a discussion classroom, wherein the speech data set comprises speech data at a plurality of different time points;
[0013] an identity recognition module configured to input the speech data set of each student into a preset voiceprint recognition model to obtain voiceprint recognition data corresponding to the speech data at each different time point, and obtain student identity corresponding to the speech data at each different time point according to the voiceprint recognition data and a preset voiceprint feature library;
[0014] a classroom text conversion module configured to input the speech data set of each student into a preset text conversion model to obtain classroom text data corresponding to the speech data at each different time point;
[0015] a cognitive level recognition module configured to input the classroom text data corresponding to the speech data at each different time point into a preset cognitive recognition model to obtain cognitive level data corresponding to the speech data at each different time point;
[0016] a display module configured to obtain cognitive change conditions of each student on the discussion classroom according to the student identity corresponding to the speech data at each different time point and the cognitive level data corresponding to the speech data at each different time point.
[0017] In a third aspect, an embodiment of the present application provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor, and the processor implements steps of the student cognitive recognition method based on speech and text classification according to the first aspect when executing the computer program.
[0018] In a fourth aspect, an embodiment of the present application provides a storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the student cognitive recognition method based on voice and text classification.
[0019] In the embodiment of the present application, a student cognitive recognition method based on voice and text classification, an apparatus, a device and a storage medium are provided. By applying voiceprint recognition technology and text conversion technology to a discussion class, the identity of a student corresponding to voice data and classroom text data are obtained, and a deep learning method is used to accurately and efficiently evaluate the cognitive condition of the student in multiple time intervals during the discussion, and a corresponding cognitive level change chart is constructed, which can reflect the performance and learning condition of the student during the discussion, provide more comprehensive information for teachers, and help to improve the teaching quality of future discussion courses.
[0020] In order to better understand and implement, the present application is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 A flowchart of the student cognitive recognition method based on voice and text classification provided by the first embodiment of the present application;
[0022] Figure 2 A flowchart of the student cognitive recognition method based on voice and text classification provided by the second embodiment of the present application;
[0023] Figure 3 A flowchart of the student cognitive recognition method based on voice and text classification provided by the third embodiment of the present application;
[0024] Figure 4 A flowchart of S8 in the student cognitive recognition method based on voice and text classification provided by the third embodiment of the present application;
[0025] Figure 5 A flowchart of S82 in the student cognitive recognition method based on voice and text classification provided by the third embodiment of the present application;
[0026] Figure 6 A flowchart of S83 in the student cognitive recognition method based on voice and text classification provided by the third embodiment of the present application;
[0027] Figure 7 A flowchart of the student cognitive recognition method based on voice and text classification provided by the fourth embodiment of the present application;
[0028] Figure 8A structure schematic diagram of a student cognitive level recognition device based on voice and text classification provided by the fifth embodiment of the present application;
[0029] Figure 9 A structure schematic diagram of a computer device provided by the sixth embodiment of the present application. DETAILED DESCRIPTION
[0030] The exemplary embodiments will be described in detail herein with reference to the attached drawings. In the following description, unless otherwise indicated, like numbers refer to like elements throughout the description and drawings. The following description is not meant to limit the application to all of the embodiments described herein. Rather, the following description is meant to provide examples of apparatus and methods consistent with the application as detailed in the following claims.
[0031] The terminology used in the present application is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used in the present application and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0032] It is to be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0033] In a discussion class, usually a teacher gives a student a question for discussion, and the student makes a speech around the question for discussion, and the cognition is the understanding degree of the student to the question when making the speech.
[0034] Please refer to Figure 1 , Figure 1 A flowchart of a student cognitive level recognition method based on voice and text classification provided by the first embodiment of the present application, the method comprising the following steps:
[0035] S1: obtaining a voice data set of each student in a discussion class, wherein the voice data set comprises voice data at different time instants.
[0036] The execution subject of the student recognition method based on voice and text classification is a recognition device (hereinafter referred to as a recognition device) of the student recognition method based on voice and text classification. In an optional embodiment, the recognition device can be a computer device, a server, or a server cluster formed by multiple computer devices.
[0037] In the embodiment, the recognition device can obtain a voice data set of each student in the discussion classroom. Specifically, before each speech of the student, a preset voice collection device is started to collect the voice data of the speech, and the time of the speech is recorded. When the student finishes the speech, the voice collection device is stopped, so as to obtain a plurality of voice data at different time points as the voice data set.
[0038] In an optional embodiment, the recognition device can also obtain the pre-stored voice data set of each student in the discussion classroom in a preset database.
[0039] S2: input the voice data set of each student into a preset voiceprint recognition model to obtain voiceprint recognition data corresponding to the voice data at different time points, and obtain the identity of each student corresponding to the voice data at different time points according to the voiceprint recognition data and a preset voiceprint feature library.
[0040] The voiceprint feature library contains the voiceprint feature data of all students. In the embodiment, the recognition device inputs the voice data into a preset voiceprint recognition model to obtain voiceprint recognition data corresponding to the voice data at different time points, and performs cosine similarity matching between the voiceprint recognition data and the voiceprint feature data of each student in the voiceprint feature library to obtain the identity of the student corresponding to the voice data.
[0041] The identity is a unique identity of each student. The identity can be the name of the student, the number of the student, or the like, which is not limited in detail herein.
[0042] In an optional embodiment, the recognition device adopts a ResNet34 convolutional neural network structure to construct the voiceprint recognition model, wherein the ResNet34 convolutional neural network structure includes a convolution layer, a pooling layer, a weight calculation layer and a full connection layer connected in sequence, the convolution layer includes an algorithm for calculating a feature vector of voice data, the pooling layer is used for dimension reduction processing to prevent gradient disappearance and network degradation, and the weight calculation layer includes a weight calculation algorithm and a weight assignment algorithm for assigning weights to the feature vector of the voice data, wherein the weight assignment algorithm is:
[0043] Xc=Fscale (Uc,Sc)
[0044] In the formula, Xc is the feature vector of the voice data after the weight is given, Uc is the Embedding corresponding to the feature vector, and Sc is the weight parameter corresponding to the feature vector.
[0045] The full connection layer functions as a classifier, and is configured to perform a weighted sum on the feature vector of the voice data after the weight is given, which is output by the weight calculation layer.
[0046] The recognition device takes the zhaishell data set in the zhvoice Chinese corpus as voiceprint recognition model training voice data, performs frame division, windowing, short-time Fourier transform, and standard deviation normalization processing on the voiceprint recognition model training voice data, obtains processed voice spectrum features, and finally inputs the voiceprint recognition model for training. Specifically, the training process of the voiceprint recognition model uses an SGD optimizer, the learning rate adjustment strategy selects StepLR, and the loss function selects cross-entropy.
[0047] Please refer to Figure 2 , Figure 2 The flowchart of the student recognition method based on voice and text classification provided by the second embodiment of the application also includes step S6, which is before step S2, and specifically as follows.
[0048] S6: Preprocess the voice data of each student to obtain preprocessed voice data of each student, wherein the preprocessing includes frame division, windowing, short-time Fourier transform, and standard deviation normalization.
[0049] Because voice data has short-time stationarity, the processing of analyzing the voice data must be based on "short time". In this embodiment, the recognition device preprocesses the voice data, including frame division, windowing, short-time Fourier transform, and standard deviation normalization. Specifically, frame division is performed by pre-setting an observation unit, which is a set of 512 sampling points, to avoid too large changes between two frames and to have an overlap area between the two frames. The overlap area is 0.001 of the sampling rate, that is, 160. A Hamming window is added to each frame of voice signal of the voice data to increase the continuity of the left end and the right end of each frame. The short-time Fourier transform is performed on each frame of the windowed signal, as follows.
[0050]
[0051] In the formula, STFT(t,f) is the frequency spectrum of the voice signal, and h(τ-t) is the Hamming window.
[0052] Subsequently, the standard deviation of the spectrum of the voice signal of the voice data is normalized to obtain preprocessed voice data.
[0053] S3: input the voice data set of each student into a preset text conversion model to obtain classroom text data corresponding to the voice data at each different time.
[0054] The text conversion model is one of ASR (Automated Speech Recognition) voice text recognition models, which is used to convert user voice data into corresponding text data.
[0055] In the embodiment, the recognition device inputs the voice data set of each student into a preset text conversion model to obtain classroom text data corresponding to the voice data at each different time. In an optional embodiment, the recognition device can perform stop word removal and word segmentation processing on the classroom text data according to a preset jieba Chinese word segmentation library to improve the accuracy and efficiency of recognizing the knowledge of students.
[0056] S4: input the classroom text data corresponding to the voice data at each different time into a preset cognitive recognition model to obtain cognitive level data corresponding to the voice data at each different time.
[0057] The cognitive recognition model is one of SVM (Support Vector Machine) classifiers, which is a generalized linear classifier (generalized linear classifier) for binary classification of data in a supervised learning (supervised learning) manner.
[0058] In the embodiment, the recognition device inputs the classroom text data corresponding to the voice data at each different time into a preset cognitive recognition model to obtain cognitive level data corresponding to the voice data at each different time. The cognitive level data can include cognitive level A, cognitive level B, cognitive level C, cognitive level D, etc., to reflect the cognitive situation of students at different times. Using deep learning method, the cognitive situation of students in the discussion classroom is accurately and efficiently evaluated.
[0059] Please refer to Figure 3 , Figure 3 The flowchart of the student cognitive recognition method provided by the third embodiment of the present application based on voice and text classification also includes training the cognitive recognition model, which includes steps S7-S8, as follows:
[0060] S7: Obtain a plurality of sample classroom text data and cognitive label data corresponding to the plurality of sample classroom text data.
[0061] The cognitive label data corresponds to the cognitive level data, and the evaluation label labeled by a person is used to reflect the cognitive situation of the sample classroom text data. In an optional embodiment, the cognitive label data can include high cognitive level, medium cognitive level, and low cognitive level.
[0062] In the embodiment, the recognition device obtains the plurality of sample classroom text data input by the user and the cognitive label data corresponding to the plurality of sample classroom text data, wherein the sample classroom text data includes a plurality of sample words.
[0063] S8: Input the plurality of sample classroom text data into a preset sentence vector representation calculation model to obtain the sentence vector representation of the plurality of sample classroom text data, and input the sentence vector representation of the plurality of sample classroom text data and the corresponding cognitive label data into a neural network model to be trained to obtain the cognitive recognition model.
[0064] The sentence vector representation calculation model includes an algorithm for calculating the sentence vector representation of the sample classroom text data.
[0065] The neural network model to be trained is an SVM classifier, including a penalty coefficient and a kernel function coefficient. In the embodiment, the recognition device inputs the plurality of sample classroom text data into a preset sentence vector representation calculation model to obtain the sentence vector representation of the plurality of sample classroom text data, and inputs the sentence vector representation of the plurality of sample classroom text data and the corresponding cognitive label data into a neural network model to be trained, iteratively trains the neural network model to be trained until the neural network model obtains the optimal penalty coefficient and the kernel function coefficient, and uses the neural network model as the cognitive recognition model.
[0066] The sentence vector representation calculation model includes a word embedding vector calculation module, a fusion word vector calculation module, and a sentence vector calculation module. Please refer to Figure 4 , Figure 4 The third embodiment of the present application provides a flowchart of the student cognitive recognition method based on speech and text classification, which is a schematic diagram of S8 in the flowchart, including steps S81-S83.
[0067] S81: According to the plurality of sample classroom text data and the word embedding vector calculation module, obtain the multi-dimensional word embedding vector representation of the plurality of sample words in the plurality of sample classroom text data output by the word embedding vector calculation module.
[0068] The word embedding vector calculation module adopts a Word2Vec word vector model. In an optional embodiment, the recognition device acquires a pre-trained Word2Vec word vector model, which is a pre-trained word vector model trained by researchers of Beijing Normal University and Renmin University of China on media such as Weibo and Sina News, and the recognition device performs transfer learning by inputting the plurality of sample classroom text data into the pre-trained word vector model as the word embedding vector calculation module.
[0069] In the embodiment, the recognition device acquires, according to the plurality of sample classroom text data and the word embedding vector calculation module, a multi-dimensional word embedding vector representation of a plurality of sample words in the plurality of sample classroom text data output by the word embedding vector calculation module according to a preset dimension in the word embedding vector calculation module.
[0070] S82: Acquire, according to the multi-dimensional word embedding vector representation of a plurality of sample words in the plurality of sample classroom text data and the fusion word vector calculation module, a fusion word vector representation of a plurality of sample words in the plurality of sample classroom text data.
[0071] The fusion word vector calculation module includes an algorithm associated with calculating a fusion word vector representation. In the embodiment, the recognition device acquires, according to the multi-dimensional word embedding vector representation of a plurality of sample words in the plurality of sample classroom text data and the fusion word vector calculation module, a fusion word vector representation of a plurality of sample words in the plurality of sample classroom text data, for training the cognitive recognition model, so as to more comprehensively analyze the cognitive situation of the student in each time interval.
[0072] Please refer to Figure 5 , Figure 5 The flowchart of the student cognitive recognition method based on speech and text classification provided by the third embodiment of the present application is a schematic diagram of S82, which includes steps S821-S823, and specifically as follows:
[0073] S821: Acquire, according to a plurality of sample words in the plurality of sample classroom text data and a preset term frequency-inverse document frequency calculation algorithm, a term frequency-inverse document frequency value corresponding to a plurality of sample words in the plurality of sample classroom text data.
[0074] The term frequency-inverse document frequency calculation algorithm is as follows:
[0075]
[0076] In the formula, C i is the term frequency-inverse document frequency value of the ith sample word, TF i,jtfij is the term frequency of the ith sample term in the jth corpus text data in the preset corpus; IDF i IDF is the inverse document frequency of the ith sample term; n i,j tfij is the number of times that the ith sample term appears in the jth corpus text data, ∑ k n k,j k is the number of different sample terms in the sample classroom text data; D is the total number of corpus text data, |j: tfij i ∈d j | is the number of corpus text data in the corpus that contains the sample term t i .
[0077] In this embodiment, the recognition device obtains the term frequency-inverse document frequency values of the sample terms in the sample classroom text data according to the preset corpus, the sample terms in the sample classroom text data, and the preset term frequency-inverse document frequency calculation algorithm.
[0078] S822: Obtain the fusion word vector representation of the sample terms in the sample classroom text data according to the multi-dimensional word embedding vector representation of the sample terms in the sample classroom text data, the corresponding term frequency-inverse document frequency values, and the preset fusion word vector calculation algorithm.
[0079] The fusion word vector calculation algorithm is:
[0080] W2V-TFIDF i =C i *W2V i
[0081] In the formula, W2V-TFIDF i is the fusion word vector of the ith sample term, C i is the term frequency-inverse document frequency value of the ith sample term, W2V i is the multi-dimensional word embedding vector representation of the ith sample term.
[0082] In this embodiment, the recognition device obtains the fusion word vector representation of the sample terms in the sample classroom text data according to the multi-dimensional word embedding vector representation of the sample terms in the sample classroom text data, the corresponding term frequency-inverse document frequency values, and the preset fusion word vector calculation algorithm. By combining the term frequency-inverse document frequency values and the multi-dimensional word embedding vector representation, the corresponding fusion word vector representation is obtained to improve the accuracy and generality of the knowledge recognition model.
[0083] S823: Obtain the sentence vector representation of the plurality of sample classroom text data according to the fused word vector representation of the plurality of sample words of the plurality of sample classroom text data and the sentence vector calculation module.
[0084] The sentence vector calculation module includes an algorithm associated with calculating a sentence vector representation. In this embodiment, the recognition device obtains the sentence vector representation of the plurality of sample classroom text data according to the fused word vector representation of the plurality of sample words of the plurality of sample classroom text data and the sentence vector calculation module.
[0085] Please refer to Figure 6 , Figure 6 The schematic diagram of S83 in the process of the student recognition method based on speech and text classification provided by the third embodiment of the present application includes step S831, which is specifically as follows:
[0086] S831: Obtain the sentence vector representation of the plurality of sample classroom text data according to the fused word vector representation of the plurality of sample words of the plurality of sample classroom text data and the preset sentence vector calculation algorithm.
[0087] The sentence vector calculation algorithm is as follows:
[0088]
[0089] In the formula, SV p is the sentence vector representation of the pth sample classroom text data, ∑ k W2V-TFIDF k is the cumulative vector of the fused word vector of all sample words of the pth sample classroom text data, and n is the dimension of the fused word vector.
[0090] In this embodiment, the recognition device averages each dimension of the fused word vector representation of the sample words contained in each of the plurality of sample classroom text data according to the fused word vector representation of the plurality of sample words of the plurality of sample classroom text data and the preset sentence vector calculation algorithm, to obtain the sentence vector representation of the plurality of sample classroom text data.
[0091] S5: Obtain the cognitive change of each student in the discussion classroom according to the student identity corresponding to the speech data at each different time and the cognitive level data corresponding to the speech data at each different time.
[0092] In this embodiment, the recognition device obtains the cognitive change of each student in the discussion classroom according to the student identity corresponding to the speech data at each different time and the cognitive level data corresponding to the speech data at each different time.
[0093] Specifically, the recognition device can combine the cognitive level data corresponding to the speech data of each different time corresponding to the student identity, as the cognitive level data set corresponding to each student, and sort a plurality of cognitive level data in the cognitive level data set corresponding to each student in chronological order, obtain the cognitive level data set corresponding to each student after sorting, and construct a cognitive level change chart or a cognitive level change table corresponding to each student according to the cognitive level data set corresponding to each student after sorting, so as to obtain the cognitive change of each student in the discussion classroom, and save to a preset storage space.
[0094] Please refer to Figure 7 , Figure 7 The flowchart of the student cognitive recognition method based on speech and text classification provided by the fourth embodiment of the present application also includes step S9, which is specifically as follows:
[0095] S9: In response to the display instruction, the display instruction includes the student identity of the student to be displayed, the cognitive change of the student to be displayed on the discussion classroom is obtained according to the student identity of the student to be displayed, and the display interface is returned to display and mark.
[0096] The display instruction is issued by the user and received by the recognition device.
[0097] The recognition device obtains the display instruction sent by the user and responds. The recognition device obtains the cognitive change of the student to be displayed on the discussion classroom from the preset storage space according to the obtained student identity of the student to be displayed, returns to the preset display interface, and displays and marks, which reflects the change of the cognitive level of the student with time, can reflect the performance and learning situation of the student during the discussion, provides more comprehensive information for the teacher, and helps to improve the teaching quality of future discussion courses.
[0098] Please refer to Figure 8 , Figure 8 The structural schematic diagram of the student cognitive level recognition device based on speech and text classification provided by the fifth embodiment of the present application. The device can realize all or part of the student cognitive level recognition device based on speech and text classification through software, hardware or combination of both. The device 8 includes:
[0099] The speech data obtaining module 81 is configured to obtain a speech data set of each student in a discussion classroom, wherein the speech data set includes speech data at a plurality of different times;
[0100] The identity recognition module 82 is configured to input the speech data set of each student into a preset voiceprint recognition model, to obtain voiceprint recognition data corresponding to the speech data at each different time, and to obtain a student identity corresponding to the speech data at each different time according to the voiceprint recognition data and a preset voiceprint feature library.
[0101] The classroom text conversion module 83 is configured to input the speech data set of each student into a preset text conversion model, to obtain classroom text data corresponding to the speech data at each different time.
[0102] The cognitive level recognition module 84 is configured to input the classroom text data corresponding to the speech data at each different time into a preset cognitive recognition model, to obtain cognitive level data corresponding to the speech data at each different time.
[0103] The display module 85 is configured to obtain cognitive change conditions of each student on the discussion classroom according to the student identity corresponding to the speech data at each different time and the cognitive level data corresponding to the speech data at each different time.
[0104] In the embodiments of the present application, the speech data set of each student on the discussion classroom is obtained through the speech data obtaining module, wherein the speech data set includes speech data at a plurality of different times; the speech data set of each student is input into a preset voiceprint recognition model through the identity recognition module, to obtain voiceprint recognition data corresponding to the speech data at each different time, and to obtain a student identity corresponding to the speech data at each different time according to the voiceprint recognition data and a preset voiceprint feature library; the speech data set of each student is input into a preset text conversion model through the classroom text conversion module, to obtain classroom text data corresponding to the speech data at each different time; the classroom text data corresponding to the speech data at each different time is input into a preset cognitive recognition model through the cognitive level recognition module, to obtain cognitive level data corresponding to the speech data at each different time; and the cognitive change conditions of each student on the discussion classroom are obtained through the display module according to the student identity corresponding to the speech data at each different time and the cognitive level data corresponding to the speech data at each different time. By applying voiceprint recognition technology and text conversion technology to the discussion classroom, the student identity corresponding to the speech data and the classroom text data are obtained, and a deep learning method is used to accurately and efficiently evaluate the cognitive conditions of students on the discussion classroom, and a corresponding cognitive level change condition chart is constructed, which can reflect the performance and learning conditions of students during the discussion, provide more comprehensive information for teachers, and help to improve the teaching quality of future discussion courses.
[0105] Please refer to Figure 9 , Figure 9 The structural schematic diagram of the computer device provided in the sixth embodiment of the present application is shown in FIG. 9. The computer device 9 comprises a processor 91, a memory 92, and a computer program 93 stored in the memory 92 and executable on the processor 91. The computer device can store a plurality of instructions, which are suitable for being loaded by the processor 91 and executing the method steps of the above-mentioned embodiments one to five. The specific execution process can refer to the specific description of embodiments one to five, which will not be described here.
[0106] The processor 91 can comprise one or more processing cores. The processor 91 connects various parts in the server through various interfaces and lines, executes various functions and processes data of the student cognitive level recognition device 8 based on voice and text classification by running or executing the instructions, programs, code sets or instruction sets stored in the memory 92, and calling the data in the memory 92. Optionally, the processor 91 can be realized in the form of at least one of digital signal processing (Digital Signal Processing, DSP), field-programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programble Logic Array, PLA). The processor 91 can be integrated with one or a combination of central processing unit (Central Processing Unit, CPU), graphics processing unit (Graphics Processing Unit, GPU), and modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the touch display screen; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 91, but be realized by a separate chip.
[0107] The memory 92 can include a random access memory (RAM) and a read-only memory (ROM). Optionally, the memory 92 includes a non-transitory computer-readable storage medium. The memory 92 can be configured to store instructions, programs, codes, code sets, or instruction sets. The memory 92 can include a program storage area and a data storage area, where the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as touch instructions, etc.), instructions for implementing the above-described various method embodiments, etc., and the data storage area can store data involved in the above-described various method embodiments, etc. The memory 92 can optionally be at least one storage device located away from the processor 91.
[0108] The embodiments of the present application further provide a storage medium, which can store a plurality of instructions. The instructions are suitable for being loaded and executed by a processor to perform the method steps of the above-described embodiments 1 to 5. For details, refer to the specific description of the embodiments 1 to 5, which will not be repeated here.
[0109] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example for illustration. In actual applications, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the above method embodiments, which will not be repeated here.
[0110] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.
[0111] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0112] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are only schematic. The division of the modules or units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0113] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0114] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0115] The integrated module / unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by computer programs instructing related hardware, and the computer programs can be stored in a computer readable storage medium. The computer program, when executed by a processor, can implement the steps of each method embodiment described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms.
[0116] The present application is not limited to the above-described embodiments, and various modifications or alterations can be made to the present application without departing from the spirit and scope of the present application, and it is intended that such modifications and alterations be included within the scope of the present application recited in the following claims and their equivalents.
Claims
1. A student recognition method based on speech and text classification, characterized in that, The method comprises the following steps: obtaining a speech data set of each student in a discussion class, wherein the speech data set comprises speech data at different time points; inputting the speech data set of each student into a preset voiceprint recognition model to obtain voiceprint recognition data corresponding to the speech data at different time points, and obtaining a student identity corresponding to the speech data at different time points according to the voiceprint recognition data and a preset voiceprint feature library; inputting the speech data set of each student into a preset text conversion model to obtain classroom text data corresponding to the speech data at different time points; inputting the classroom text data corresponding to the speech data at different time points into a preset cognitive recognition model to obtain cognitive level data corresponding to the speech data at different time points; obtaining cognitive changes of each student in the discussion class according to the student identity corresponding to the speech data at different time points and the cognitive level data corresponding to the speech data at different time points, comprising combining the cognitive level data corresponding to the speech data at different time points corresponding to the student identity as a cognitive level data set corresponding to each student, sorting a plurality of cognitive level data in the cognitive level data set corresponding to each student in chronological order to obtain the cognitive level data set corresponding to each student after sorting, and constructing a cognitive level change chart or a cognitive level change table corresponding to each student according to the cognitive level data set corresponding to each student after sorting to obtain the cognitive changes of each student in the discussion class; obtaining a plurality of sample classroom text data and cognitive label data corresponding to the plurality of sample classroom text data, wherein the sample classroom text data comprises a plurality of sample words; inputting the plurality of sample classroom text data into a preset sentence vector representation calculation model to obtain sentence vector representations of the plurality of sample classroom text data, and inputting the sentence vector representations of the plurality of sample classroom text data and the corresponding cognitive label data into a neural network model to be trained to obtain the cognitive recognition model.
2. The student recognition method based on voice and text classification according to claim 1, characterized in that, Before inputting the speech data of each student into a preset voiceprint recognition model to obtain voiceprint recognition data corresponding to each speech data, the method comprises the following steps: preprocessing the speech data of each student to obtain preprocessed speech data of each student, wherein the preprocessing comprises framing, windowing, short-time Fourier transform, and standard deviation normalization.
3. The student cognitive recognition method based on speech and text classification according to claim 1, characterized in that: the sentence vector representation calculation model comprises a word embedding vector calculation module, a fused word vector calculation module, and a sentence vector calculation module; the inputting of the plurality of sample classroom text data into a preset sentence vector representation calculation model to obtain sentence vector representations of the plurality of sample classroom text data comprises the following steps: According to the plurality of sample classroom text data and the word embedding vector calculation module, the multi-dimensional word embedding vector representation of the plurality of sample words in the plurality of sample classroom text data output by the word embedding vector calculation module is obtained; According to the multi-dimensional word embedding vector representation of the plurality of sample words in the plurality of sample classroom text data and the fusion word vector calculation module, the fusion word vector representation of the plurality of sample words in the plurality of sample classroom text data is obtained; According to the multi-dimensional word embedding vector representation of the plurality of sample words in the plurality of sample classroom text data and the sentence vector calculation module, the sentence vector representation of the plurality of sample classroom text data is obtained.
4. The student recognition method based on voice and text classification according to claim 3, characterized in that, According to the multi-dimensional word embedding vector representation of the plurality of sample words in the plurality of sample classroom text data and the fusion word vector calculation module, the fusion word vector representation of the plurality of sample words in the plurality of sample classroom text data is obtained, comprising the steps of: According to the plurality of sample words in the plurality of sample classroom text data and the preset term frequency-inverse document frequency calculation algorithm, the corresponding term frequency-inverse document frequency value of the plurality of sample words in the plurality of sample classroom text data is obtained, wherein the term frequency-inverse document frequency calculation algorithm is: wherein C i is the term frequency-inverse document frequency value of the ith sample term, TF i,j is the term frequency of the ith sample term in the jth corpus text data in the preset corpus; IDF i is the inverse document frequency of the ith sample term; n i,j is the number of times the ith sample term appears in the jth corpus text data, ∑ k n k,j is the total number of times all corpus terms appear in the jth corpus text data, k is the number of different sample terms in the sample classroom text data; D is the total number of corpus text data, |j:t i ∈d j | is the number of corpus text data in the corpus that contains the sample term t i According to the multi-dimensional word embedding vector representation of the plurality of sample words in the plurality of sample classroom text data, the corresponding term frequency-inverse document frequency value, and the preset fusion word vector calculation algorithm, the fusion word vector representation of the plurality of sample words in the plurality of sample classroom text data is obtained, wherein the fusion word vector calculation algorithm is: W2V-TFIDF i = C i * W2V i where W2V-TFIDF is the TFIDF-weighted word vector representation of the ith sample word i is the fused word vector representation of the ith sample word, i is the term frequency-inverse document frequency value of the ith sample word, i is the multi-dimensional word embedding vector representation of the ith sample word.
5. The student recognition method based on voice and text classification according to claim 4, characterized in that, According to the multi-dimensional word embedding vector representation of the plurality of sample words in the plurality of sample classroom text data and the sentence vector calculation module, the sentence vector representation of the plurality of sample classroom text data is obtained, comprising the steps of: According to the multi-dimensional word embedding vector representation of the plurality of sample words in the plurality of sample classroom text data and the preset sentence vector calculation algorithm, the sentence vector representation of the plurality of sample classroom text data is obtained, wherein the sentence vector calculation algorithm is: where SV p is the sentence vector representation of the pth sample classroom text data,∑ k W2V-TFIDF k is the accumulated vector of the fused word vectors of all sample words of the pth sample classroom text data, n is the dimension of the fused word vectors.
6. The student recognition method based on voice and text classification according to claim 1, characterized in that, Further comprising the steps of: In response to a display instruction, the display instruction includes a student identity of a student to be displayed, according to the student identity of the student to be displayed, the cognitive change of the student to be displayed on the discussion classroom is obtained, returned to the preset display interface, and displayed and marked.
7. A student cognitive level recognition apparatus based on voice and text classification, characterized by, Comprise: The voice data obtaining module is used for obtaining a voice data set of each student on the discussion classroom, wherein the voice data set comprises voice data at different time moments; The identity recognition module is used for inputting the voice data set of each student into a preset voiceprint recognition model to obtain voiceprint recognition data corresponding to the voice data at each different time moment, and obtaining a student identity corresponding to the voice data at each different time moment according to the voiceprint recognition data and a preset voiceprint feature library; The classroom text conversion module is used for inputting the voice data set of each student into a preset text conversion model to obtain classroom text data corresponding to the voice data at each different time moment; The cognitive level recognition module is configured to input the classroom text data corresponding to the speech data at the different time moments into a preset cognitive recognition model, and obtain cognitive level data corresponding to the speech data at the different time moments. The display module is configured to obtain cognitive change conditions of each student in the discussion classroom according to the student identity corresponding to the speech data at the different time moments and the cognitive level data corresponding to the speech data at the different time moments, including combining the cognitive level data corresponding to the speech data at the different time moments corresponding to the student identity as a cognitive level data set corresponding to each student based on the student identity, sorting a plurality of cognitive level data in the cognitive level data set corresponding to each student in a time sequence to obtain the cognitive level data set corresponding to each student after sorting, and constructing a cognitive level change condition chart or a cognitive level change condition table corresponding to each student according to the cognitive level data set corresponding to each student after sorting to obtain the cognitive change conditions of each student in the discussion classroom. A plurality of sample classroom text data and cognitive label data corresponding to the sample classroom text data are obtained, wherein the sample classroom text data includes a plurality of sample words. The plurality of sample classroom text data are input into a preset sentence vector representation calculation model to obtain sentence vector representations of the plurality of sample classroom text data, and the sentence vector representations of the plurality of sample classroom text data and the corresponding cognitive label data are input into a neural network model to be trained to obtain the cognitive recognition model.
8. A computer device, comprising: The storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the student cognitive recognition method based on speech and text classification according to any one of claims 1 to 6.
9. A storage medium characterized by: The storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the student cognitive recognition method based on speech and text classification according to any one of claims 1 to 6.
Citation Information
Patent Citations
Learning condition evaluation method and device, storage medium and electronic equipment
CN110600033A
Cognitive function evaluation device, cognitive function evaluation system, and cognitive function evaluation method
US20200229752A1