Voice-based group learning atmosphere analysis method and system

By constructing speech enhancement and voiceprint recognition models, the problems of uneven participation and high evaluation difficulty in group learning are solved, enabling efficient learning atmosphere analysis and personalized feedback, and improving the accuracy of learning atmosphere assessment.

CN119446152BActive Publication Date: 2025-11-25SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411417885.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-11
Publication Date
2025-11-25
Estimated Expiration
2044-10-11

AI Technical Summary

Technical Problem

In traditional group learning, uneven participation among members can lead to conflicts and contradictions, making it difficult for teachers to understand the learning situation of each group and making classroom assessment challenging.

Method used

By constructing a target speech enhancement model and a target voiceprint recognition model, we obtain group speech data from the classroom, perform data enhancement and voiceprint recognition, calculate group learning participation, and combine emotion prediction analysis to understand the learning atmosphere.

Benefits of technology

It improves the clarity and accuracy of voice signals, can identify the voice of each group member, provide personalized feedback, quickly capture classroom learning, and achieve a comprehensive assessment of the group learning atmosphere.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446152B_ABST
    Figure CN119446152B_ABST
Patent Text Reader

Abstract

The application discloses a voice-based group learning atmosphere analysis method and system, which comprises the following steps: constructing a target voice enhancement model and a target voiceprint recognition model; inputting group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information; inputting the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data; calculating the target voice enhancement information and the target voice character information to obtain group learning participation; performing feature extraction processing on the target voice enhancement information to obtain target voice feature information; inputting the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information; and analyzing and processing the group learning participation and the target emotion information to obtain group learning atmosphere data. The application can obtain comprehensive group learning atmosphere data and can be widely applied to the technical field of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a group learning atmosphere analysis method and system based on voice. BACKGROUND

[0002] At present, group learning is an important education and teaching method. Students complete learning tasks in the form of grouping. In the traditional group learning process, there may be uneven participation among members, and even conflicts and contradictions may be caused for the group members. For teachers, it is difficult to evaluate the classroom. Therefore, it is difficult for teachers to understand the learning situation of each group in the group learning process.

[0003] In summary, the technical problems in the related art need to be improved. SUMMARY

[0004] Embodiments of the present application aim to at least partially solve one of the technical problems in the related art. To this end, the main purpose of the embodiments of the present application is to propose a group learning atmosphere analysis method and system based on voice, which can obtain comprehensive group learning atmosphere data.

[0005] To achieve the above-mentioned purpose, one aspect of the embodiments of the present application proposes a group learning atmosphere analysis method based on voice, which comprises the following steps:

[0006] Obtain a voice information processing scheme, and construct a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme;

[0007] Obtain group voice data in a teaching classroom;

[0008] Input the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information;

[0009] Input the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data;

[0010] Calculate the target voice enhancement information and the target voice character information to obtain group learning participation degree;

[0011] Perform feature extraction processing on the target voice enhancement information to obtain target voice feature information;

[0012] Input the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information;

[0013] Analyze and process the group learning participation degree and the target emotion information to obtain group learning atmosphere data.

[0014] In some embodiments, the voice information processing scheme includes a local processing scheme, the voice information processing scheme is obtained, and a target voice enhancement model and a target voiceprint recognition model are constructed based on the voice information processing scheme, including:

[0015] The voice information processing scheme and the to-be-trained voice data are obtained.

[0016] According to the teaching environment information, a target voice processing scheme is selected from the voice information processing scheme.

[0017] If the target voice processing scheme is the local processing scheme in the voice information processing scheme, the target voice enhancement model and the target voiceprint recognition model are constructed according to the to-be-trained voice data.

[0018] In some embodiments, if the target voice processing scheme is the local processing scheme in the voice information processing scheme, the target voice enhancement model and the target voiceprint recognition model are constructed according to the to-be-trained voice data, including:

[0019] If the target voice processing scheme is the local processing scheme in the voice information processing scheme, information extraction is performed on the to-be-trained voice data to obtain clean voice data.

[0020] Reverb effects are added to the clean voice data to obtain reverb voice data.

[0021] The reverb voice data and a plurality of target noise segments are integrated to obtain target training data.

[0022] The target training data is input into an initial voice enhancement model for training to obtain the target voice enhancement model.

[0023] If the target voice processing scheme is the local processing scheme in the voice information processing scheme, feature extraction processing is performed on the to-be-trained voice data to obtain target training feature information.

[0024] The target training feature information is input into a linear time series model for training to obtain the target voiceprint recognition model; the target voiceprint recognition model is used to obtain voiceprint information of a target voice person, and the voiceprint information is used to construct a voiceprint model library.

[0025] In some embodiments, the target training feature information is input into a linear time series model for training to obtain the target voiceprint recognition model, including:

[0026] inputting the target training feature information into the linear time sequence model for training to obtain an initial embedding vector;

[0027] performing regularization processing on the initial embedding vector to obtain a target embedding vector;

[0028] calculating a centroid of the target embedding vector to obtain a target vector centroid;

[0029] constructing a similarity matrix based on the target embedding vector and the target vector centroid;

[0030] training the linear time sequence model based on the similarity matrix and a target voiceprint loss function to obtain the target voiceprint recognition model.

[0031] In some embodiments, the calculating the target voice enhancement information and the target voice person information to obtain a group learning participation degree comprises:

[0032] calculating a target active person quantity and a target voice person quantity in the target voice person information within a target time period;

[0033] taking a proportion of the target active person quantity and the target voice person quantity as the group learning participation degree.

[0034] In some embodiments, the feature extraction processing on the target voice enhancement information to obtain target voice feature information comprises:

[0035] performing audio feature extraction processing on the target voice enhancement information to obtain an audio feature vector;

[0036] performing text feature extraction processing on the target voice enhancement information to obtain a text feature vector corresponding to the audio feature vector;

[0037] performing splicing processing on the audio feature vector and the text feature vector to obtain the target voice feature information.

[0038] In some embodiments, the voice information processing scheme comprises a cloud processing scheme, and the method further comprises:

[0039] selecting a target voice processing scheme from the voice information processing scheme according to teaching environment information;

[0040] if the target voice processing scheme is the cloud processing scheme in the voice information processing scheme, constructing a target cloud processing model based on the cloud processing scheme;

[0041] obtaining the group voice data in the teaching classroom;

[0042] input the group voice data into the target cloud processing model for data processing to obtain target cloud voice information and the target voice character information corresponding to the group voice data;

[0043] perform calculation on the target cloud voice information and the target voice character information to obtain the group learning participation degree;

[0044] perform feature extraction processing on the target cloud voice information to obtain the target voice feature information.

[0045] To achieve the above-mentioned purpose, another aspect of the embodiment of the present application proposes a group learning atmosphere analysis system based on voice, which comprises the following modules:

[0046] a model construction module, configured to obtain a voice information processing scheme, and construct a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme;

[0047] a group voice data acquisition module, configured to acquire group voice data in a teaching classroom;

[0048] a data enhancement processing module, configured to input the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information;

[0049] a voiceprint recognition processing module, configured to input the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data;

[0050] a group learning participation degree calculation module, configured to perform calculation on the target voice enhancement information and the target voice character information to obtain a group learning participation degree;

[0051] a feature extraction processing module, configured to perform feature extraction processing on the target voice enhancement information to obtain target voice feature information;

[0052] an emotion information prediction module, configured to input the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information;

[0053] a group learning atmosphere analysis module, configured to perform analysis processing on the group learning participation degree and the target emotion information to obtain group learning atmosphere data.

[0054] To achieve the above-mentioned purpose, another aspect of the embodiment of the present application proposes an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the method described above when executing the computer program.

[0055] To achieve the above object, another aspect of the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method described above.

[0056] The embodiment of the present application has at least the following beneficial effects: the present application provides a voice-based group learning atmosphere analysis method and system, which obtains a voice information processing scheme, constructs a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme, obtains group voice data in a teaching classroom, inputs the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information, inputs the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data, calculates the target voice enhancement information and the target voice character information to obtain a group learning participation degree, extracts features from the target voice enhancement information to obtain target voice feature information, inputs the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information, and analyzes and processes the group learning participation degree and the target emotion information to obtain group learning atmosphere data. The embodiment of the present application enhances the group voice data by using the target voice enhancement model, improves the clarity and purity of the voice signal, and thus improves the accuracy of subsequent emotion prediction and group learning participation degree analysis; the target voiceprint recognition model is used to recognize the voiceprint of the group voice data, which can accurately identify the voice of each group member and provide the possibility for subsequent personalized feedback and intervention; the group learning participation degree is combined with the emotion prediction result for analysis, which can obtain comprehensive group learning atmosphere data, so that the group members and teachers can quickly capture the learning situation in the classroom. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 is a step flowchart of the voice-based group learning atmosphere analysis method provided by the embodiment of the present application;

[0058] Figure 2 is a framework schematic diagram of a voiceprint recognition model training and inference process provided by the embodiment of the present application;

[0059] Figure 3 is a framework schematic diagram of a voice enhancement process provided by the embodiment of the present application;

[0060] Figure 4 is a framework schematic diagram of a group learning participation degree calculation process provided by the embodiment of the present application;

[0061] Figure 5 is a structure schematic diagram of an emotion analysis and recognition module provided by the embodiment of the present application;

[0062] Figure 6 is a schematic diagram of interface content of a visualization module provided by an embodiment of the present application;

[0063] Figure 7 is a schematic diagram of a process of a group learning atmosphere analysis method based on voice provided by an embodiment of the present application;

[0064] Figure 8 is a schematic diagram of a neural network application implementation of a group learning atmosphere analysis method based on voice provided by an embodiment of the present application;

[0065] Figure 9 is a schematic diagram of a structure of a group learning atmosphere analysis system based on voice provided by an embodiment of the present application;

[0066] Figure 10 is a schematic diagram of a hardware structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application. When the following description refers to the accompanying drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementation described in the following exemplary embodiments does not represent all the implementations consistent with the embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0068] It can be understood that the terms "first", "second", and the like used in the present application can be used herein to describe various concepts, but unless specifically stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "when" or "in response to determining".

[0069] The terms "at least one", "multiple", "each", "any" and the like used in the present application include one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application.

[0071] Before the embodiments of the present application are explained in detail, the description of the embodiments of the present application will be described with some terms and terminologies involved in the embodiments of the present application, which are applicable to the following explanations.

[0072] At present, group learning is an important education and teaching method, and students complete learning tasks through grouping. In the traditional group learning process, there may be uneven participation among members, and even conflicts and contradictions may be caused for the group members. For teachers, it is difficult to evaluate the classroom. Therefore, it is difficult for teachers to understand the learning situation of each group in the group learning process.

[0073] Therefore, in the embodiments of the present application, a group learning atmosphere analysis method and system based on voice are provided. The scheme acquires a voice information processing scheme, constructs a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme, acquires group voice data in a teaching classroom, inputs the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information, inputs the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data, calculates the target voice enhancement information and the target voice character information to obtain a group learning participation degree, extracts features from the target voice enhancement information to obtain target voice feature information, inputs the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information, and analyzes and processes the group learning participation degree and the target emotion information to obtain group learning atmosphere data. The embodiments of the present application enhance the group voice data through the target voice enhancement model, improve the clarity and purity of the voice signal, and thus improve the accuracy of subsequent emotion prediction and group learning participation degree analysis; the target voiceprint recognition model is used to recognize the voiceprint of the group voice data, which can accurately identify the voice of each group member, and provides the possibility for subsequent personalized feedback and intervention; the group learning participation degree is combined with the emotion prediction result for analysis, and comprehensive group learning atmosphere data can be obtained, so that the group members and teachers can quickly capture the learning situation in the classroom.

[0074] The method for analyzing a group learning atmosphere based on voice provided by the embodiments of the present application relates to the technical field of data processing. The method for analyzing a group learning atmosphere based on voice provided by the embodiments of the present application can be applied to a terminal, can be applied to a server, and can also be software running in the terminal or the server. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, and the like, but is not limited thereto; the server end can be configured as a stand-alone physical server, can be configured as a server cluster or a distributed system composed of multiple physical servers, can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs (Content Delivery Networks), and big data and artificial intelligence platforms, and the server can also be a node server in a blockchain network; and the software can be an application for implementing the method for analyzing a group learning atmosphere based on voice, and the like, but is not limited to the above forms.

[0075] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs (Personal Computers), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0076] Please refer to Figure 1 , Figure 1 is an optional step flowchart of the method for analyzing a group learning atmosphere based on voice provided by the embodiments of the present application, Figure 1 The method in the step S101 to the step S108 can include but is not limited to the steps.

[0077] In step S101, a voice information processing scheme is acquired, and a target voice enhancement model and a target voiceprint recognition model are constructed based on the voice information processing scheme;

[0078] Optionally, the speech information processing scheme includes a local processing scheme and a cloud processing scheme. The local processing scheme is a scheme including a training data preprocessing process and a model construction process, and after training data preprocessing and model construction, a target speech enhancement model and a target voiceprint recognition model can be obtained. The target speech enhancement model and the target voiceprint recognition model are used to process real-time teaching classroom speech data to obtain a group learning atmosphere. The cloud processing scheme is a scheme of directly deploying a cloud processing model (a cloud online speaker diarization model), and then processing real-time teaching classroom speech data according to the cloud processing model to obtain a group learning atmosphere. The cloud processing model can accurately detect the speaking state of the target speaker in a complex audio environment through multi-layer feature extraction and sequence modeling. Specifically, if there is no condition or requirement for deploying a local voiceprint recognition model in the teaching implementation environment, the training data preprocessing process is not required, and the next implementation step S102 is directly entered. If there is a condition or requirement for deploying a local voiceprint recognition model in the teaching implementation environment, the training data needs to be preprocessed.

[0079] It should be noted that the cloud processing model is a mature model in the related art, such as a TS-VAD model. The structure of the TS-VAD model usually includes an input feature extraction layer, a target speaker feature encoding layer, a speaker correlation analysis layer, and a VAD decision layer. It is easy to understand that for the selection of the cloud processing model, a related model can be selected with reference to the TS-VAD model, and the present embodiment is not limited in this regard. In addition, for the specific implementation principle of the TS-VAD model, reference can be made to the specific description of the related technical scheme, and the present embodiment will not be repeated here.

[0080] In some embodiments, step S101 can include: obtaining a speech information processing scheme and training speech data; selecting a target speech processing scheme from the speech information processing scheme according to teaching environment information; and if the target speech processing scheme is a local processing scheme in the speech information processing scheme, constructing a target speech enhancement model and a target voiceprint recognition model according to the training speech data.

[0081] The target voice processing scheme is a voice information processing scheme selected according to whether a local voiceprint recognition model is deployed in the teaching implementation environment or a requirement thereof. The target voice processing scheme can be a local processing scheme or a cloud processing scheme. The target voice processing scheme is a voice information processing scheme selected according to whether a local voiceprint recognition model is deployed in the teaching implementation environment or a requirement thereof. The target voice processing scheme can be a local processing scheme or a cloud processing scheme. In some embodiments, if the target voice processing scheme is the local processing scheme in the voice information processing scheme, constructing the target voice enhancement model and the target voiceprint recognition model according to the training voice data can include: if the target voice processing scheme is the local processing scheme in the voice information processing scheme, extracting information from the training voice data to obtain clean voice data; adding a reverberation effect to the clean voice data to obtain reverberation voice data; integrating the reverberation voice data and a plurality of target noise segments to obtain target training data; inputting the target training data into an initial voice enhancement model for training to obtain the target voice enhancement model; if the target voice processing scheme is the local processing scheme in the voice information processing scheme, extracting features from the training voice data to obtain target training feature information; inputting the target training feature information into a linear time series model for training to obtain the target voiceprint recognition model; and the target voiceprint recognition model is used to obtain voiceprint information of a target voice person, and the voiceprint information is used to construct a voiceprint model library.

[0082] The teaching environment information refers to various parameters and conditions related to a specific teaching environment, including but not limited to the size, layout, noise level, echo condition, number of students, group division, etc. of a classroom.

[0083] The target voice processing scheme is a voice information processing scheme selected according to whether a local voiceprint recognition model is deployed in the teaching implementation environment or a requirement thereof. The target voice processing scheme can be a local processing scheme or a cloud processing scheme.

[0084] In some embodiments, if the target voice processing scheme is the local processing scheme in the voice information processing scheme, constructing the target voice enhancement model and the target voiceprint recognition model according to the training voice data can include: if the target voice processing scheme is the local processing scheme in the voice information processing scheme, extracting information from the training voice data to obtain clean voice data; adding a reverberation effect to the clean voice data to obtain reverberation voice data; integrating the reverberation voice data and a plurality of target noise segments to obtain target training data; inputting the target training data into an initial voice enhancement model for training to obtain the target voice enhancement model; if the target voice processing scheme is the local processing scheme in the voice information processing scheme, extracting features from the training voice data to obtain target training feature information; inputting the target training feature information into a linear time series model for training to obtain the target voiceprint recognition model; and the target voiceprint recognition model is used to obtain voiceprint information of a target voice person, and the voiceprint information is used to construct a voiceprint model library.

[0085] In some embodiments, inputting the target training feature information into the linear time series model for training to obtain the target voiceprint recognition model can include: inputting the target training feature information into the linear time series model for training to obtain an initial embedding vector; performing regularization processing on the initial embedding vector to obtain a target embedding vector; calculating a centroid of the target embedding vector to obtain a target vector centroid; constructing a similarity matrix based on the target embedding vector and the target vector centroid; and training the linear time series model based on the similarity matrix and a target voiceprint loss function to obtain the target voiceprint recognition model.

[0086] wherein the pure speech data is speech data without noise and reverberation effects. Reverberation refers to the phenomenon of sound waves reflecting off reflective surfaces (such as walls, ceilings, floors, etc.) and superimposing onto the direct sound. The reverberation effect can cause the sound to continue for a period of time during transmission, making the original sound more blurred and diffuse. It is easily understood that the reverberant speech data refers to the speech signal processed by simulating the acoustic characteristics (and reverberation) in the teaching environment based on the pure speech data.

[0087] In a specific implementation, for the training of the target speech enhancement model, using speech segments with reverberation can improve the robustness and effectiveness of the model in actual application; and adding noise enables the model to learn to identify and separate the speech components in the noise signal during the training process, thereby providing clearer and higher quality speech output in actual application.

[0088] wherein the target noise segment refers to different types of noise samples used to mix with the speech signal to generate noisy speech, such as white noise, traffic noise, etc. The target training data is a speech data set used to train the speech enhancement model or the voiceprint recognition model, which includes pure speech data, reverberant speech data, and noisy speech data, etc.

[0089] The initial speech enhancement model refers to the initial version of the speech enhancement model used at the beginning of the training phase. The initial speech enhancement model will be optimized through the training process to obtain the target speech enhancement model. The target speech enhancement model is the final speech enhancement model that has been trained and optimized to effectively remove noise and reverberation and improve speech quality.

[0090] Optionally, the target training feature information is a feature vector extracted from the to-be-trained voice data, and the target training feature information is used for training the voiceprint recognition model. The linear time series model is a machine learning model for voiceprint recognition, which can process time series data and learn the characteristics of a speaker therefrom. By inputting the target training feature information into the linear time series model for training, a target voiceprint recognition model can be obtained. The target voice person can refer to a student or a teacher in a teaching classroom. The voiceprint model library is a database containing voiceprint information corresponding to multiple speakers.

[0091] For the initial embedding vector, which is an initial vector extracted from the target training feature information for representing the voiceprint characteristics of a speaker in the training process of the voiceprint recognition model, the initial embedding vector can be processed by L2 regularization to obtain a target embedding vector with stronger representation ability and generalization ability, which can more accurately reflect the identity characteristics of the speaker. The target vector centroid is a center point calculated in a group of target embedding vectors, and the target vector centroid can be used to represent the overall characteristics or central tendency of the group of vectors. In the voiceprint recognition task, the target vector centroid is used to compare with the target embedding vector to calculate the similarity therebetween, so as to obtain a similarity matrix defining the cosine similarity between each embedding vector and all centroids.

[0092] Optionally, the target voiceprint loss function is defined based on the target embedding vector and the target vector centroid, and is used for training and optimizing the linear time series model to obtain the target voiceprint recognition model.

[0093] In a specific implementation, the process of constructing the target voiceprint recognition model is as follows: firstly, a certain amount of speaker corpus data (to-be-trained voice data) is obtained, the obtained corpus data is subjected to feature extraction to obtain corpus features (target training feature information); the obtained corpus features are processed by a linear time series model with a selective state space (also referred to as a Mamba model) to obtain an embedding vector, and then the embedding vector is subjected to L2 regularization; the centroids of the embedding vectors of all speakers are calculated, and a similarity matrix is constructed based on the embedding vectors and the vector centroids, wherein the similarity matrix defines the cosine similarity between each embedding vector and all centroids; then a voiceprint loss function is defined based on the embedding vector and the vector centroid to train the linear time series model, and the training purpose is to make each embedding vector closer to the centroid of the real speaker and farther away from the centroids of other speakers; finally, the linear time series model is trained based on the similarity matrix and the loss function to obtain a voiceprint recognition model, and the voiceprint recognition model comprises a student voiceprint encoder, the student voiceprint encoder is used to obtain student voiceprint information, and multiple voiceprint information is used to construct a voiceprint information library.

[0094] In a specific implementation, the process of constructing the target speech enhancement model is as follows: first, information extraction is performed on the to-be-trained speech data to obtain student clean speech data, reverberation effects are added to the clean speech data to obtain speech with reverberation, and training data is generated using the speech with reverberation and different noise segments; then, a convolutional layer is used to extract features of the training data, and a dual-path recurrent neural network module is used to model time and frequency dependencies of the training data, that is, the speech signal of the training data is modeled in two dimensions of time and frequency to effectively capture the time-frequency features of the speech signal; next, a speech enhancement loss function is defined, and an initial speech enhancement model is optimized using a loss function combining negative signal-to-noise ratio (SNR) and mean square error (MSE); the model parameters of the initial speech enhancement model are updated through back propagation during the training process, and finally the target speech enhancement model is obtained, which uses batch normalization and instantaneous layer normalization to process data.

[0095] In the embodiments of the present application, the implementation process of step S101 includes the following two steps (steps A and B):

[0096] Step A, according to the teaching environment information, a speech information processing scheme is selected.

[0097] Specifically, if there is no condition or requirement for deploying a local voiceprint recognition model in the teaching implementation environment (i.e., a cloud processing scheme is adopted), no preprocessing is required, and the next implementation step B is directly entered. If there is a condition or requirement for deploying a local voiceprint recognition model in the teaching implementation environment (i.e., a local processing scheme is adopted), the following preprocessing procedure is performed, specifically: first, different corpus data of multiple speakers are obtained, then the corpus data is feature extracted using mel frequency cepstral coefficients, and the extracted features are input into a Mamba model for training to obtain a series of voiceprint information libraries.

[0098] For example, if there is a condition or requirement for deploying a local voiceprint recognition model in the teaching implementation environment, the data preprocessing procedure needs to be performed, please refer to Figure 2 , Figure 2 is a schematic diagram of a voiceprint recognition model training and inference process framework provided by the embodiments of the present application, wherein the data preprocessing procedure (i.e., the training phase of the voiceprint recognition model) is as follows: Figure 2 as shown, first, K speaker training speech information is obtained, and each speaker has S training speech information; then, the features of the training speech information are extracted using mel frequency cepstral coefficients to obtain feature vectors, each feature vector x ij (1≤i≤K,1≤j≤S) represents the features extracted from the corpus data j of speaker i. f(x ij; w) represents the output of a neural network (Mamba model), where w represents all the parameters of the neural network, and the L2 normalized output of the neural network is the speaker voiceprint embedding vector, represents the embedding vector of the jth speech information of the ith speaker, and the embedding vector of the ith speaker is [e 1j ,…,e Mj ]. The similarity matrix S ij,k is defined as the scaled cosine similarity between each embedding vector e ij and all the center points o k . where w and b are learnable parameters. The training objective is to make the corpus embedding of each speaker similar to the center point of all the embeddings of this speaker, while being far away from the embedding centers of other speakers. The loss of each embedding vector e ih can be defined as This loss function represents pulling each embedding vector closer to its center point and further away from all other center points. After training, the target voiceprint recognition model is obtained, which contains a student voiceprint encoder for obtaining student voiceprint information. Multiple voiceprint information (such as Figure 2 "Voiceprint information A" and "Voiceprint information B" in

[0099] After completing the training of the voiceprint recognition model and obtaining the target voiceprint recognition model, the inference stage process of the voiceprint recognition model is as follows: as shown in Figure 2 , first, the test speech is processed for feature extraction to obtain test speech feature information; then the extracted test speech feature information is input into the already trained target voiceprint recognition model; in the target voiceprint recognition model, the test speech feature information is compared with the voiceprint features of known speakers stored in the model library, and scoring / judgment is performed, i.e., scoring is performed according to the similarity of the input feature and the known voiceprint feature, and through comparison and calculation, it is determined which (or which) known speaker's voiceprint feature is most matched with the test speech; finally, according to the result of scoring / judgment, the system determines the identity of the speaker to which the test speech belongs, and outputs this result.

[0100] Step B, according to the selection of the speech information processing scheme in step A, deploy the cloud online speaker diarization model or train the local speech enhancement model and the local voiceprint recognition model. Specifically: if a non-local voiceprint recognition scheme (i.e., a cloud processing scheme) is selected, the cloud online speaker diarization model is deployed; if a local voiceprint recognition scheme (i.e., a local processing scheme) is selected, the local model is trained.

[0101] Exemplarily, if the local deployment voiceprint recognition scheme (i.e., the local processing scheme) is selected, the process of local model training is as follows: first, collect pure speech data, add reverberation effects to the pure speech data to obtain reverberation speech, and generate training data using the reverberation speech and different noise segments; in the training stage, the pure audio is randomly segmented into 5-second segments, and the pure audio is convolved with the room impulse responses (RIRs) randomly selected from openSLR26 and openSLR28. Then, the noisy speech is generated by mixing the reverberation speech with the noise, and the signal-to-noise ratio (SNR) is set to be between -5dB and 5dB; then, the features of the noisy speech are extracted using the convolution layer, and the time and frequency dependence is modeled using the dual-path recurrent neural network module; wherein, the training target is the complex ratio mask (CRM), and the real part and the imaginary part of the complex ratio mask output by the decoder are respectively taken as two streams; in the training stage, the learning target is optimized by signal approximation, that is, the spectrum of the noisy speech is multiplied by the estimated mask to obtain the enhanced spectrum, and the time-domain signal is reconstructed by inverse STFT (Short-Time Fourier Transform); then, the loss function is defined, and the loss function combining the negative SNR and the mean square error (MSE) is used to optimize the model. The first loss function is the negative SNR, which is defined as: The first loss function can constrain the amplitude of the output, avoiding the level offset between the input and the output; the second loss function combines the negative SNR and the mean square error (MSE) of the spectrum, which is defined as: The MSE loss function includes the errors of the real part, the imaginary part and the amplitude, and by taking the logarithm, it is ensured to be in the same order of magnitude as the negative SNR; the model parameters of the initial speech enhancement model are updated through back propagation in the training process, and finally the target speech enhancement model is obtained, which uses batch normalization and instantaneous layer normalization to process data.

[0102] Please refer to Figure 3 , Figure 3 is a framework schematic diagram of a speech enhancement process provided by an embodiment of the present application; as shown in Figure 3 , the framework of the speech enhancement process includes a spectrum decomposition module, an encoding module, a neural network module, a decoding module and a spectrum reconstruction module, and the functions of each module are as follows:

[0103] Spectrum decomposition module: convert the time-domain signal of the input speech into frequency-domain representation through STFT, so that the model can more conveniently process complex speech signals;

[0104] Encoding module: extracts useful features from the noisy spectrum, reduces the dimensionality of the data while preserving important speech information;

[0105] Neural network module: uses a bi-path recurrent neural network module to capture and enhance complex features in the speech signal, improving the performance of the speech enhancement model;

[0106] Decoding module: restores the low-dimensional features output by the network to the size of the original spectrum to perform inverse short-time Fourier transform (iSTFT) and reconstruct the time-domain signal;

[0107] Spectrum reconstruction module: reconstructs the enhanced frequency domain features into time domain signals through iSTFT to ensure that the processed signals can be actually used, and finally obtains the enhanced speech data.

[0108] Step S102, obtaining group speech data in a teaching classroom;

[0109] Among them, the group speech data refers to the collection of speech information generated by all participants (mainly teachers and students) in a teaching classroom. The group speech data can include the teaching content of the teacher, the questions, answers, discussions of the students, and other sound activities in the classroom.

[0110] In the embodiments of the present application, the general mode of group learning is group learning. For example, in a teaching classroom, group members of the same group sit together, an array microphone device is placed in the center of the table, and real-time speech collection is performed using the microphone array to obtain group speech data in the teaching classroom.

[0111] Step S103, inputting the group speech data into the target speech enhancement model for data enhancement processing to obtain target speech enhancement information;

[0112] Among them, the target speech enhancement information refers to the speech signal with higher quality obtained after processing by the target speech enhancement model.

[0113] For example, the array microphone is used to obtain speech data in each direction of the group members in real time, and the speech data in each direction is input into the target speech enhancement model as input, and the speech enhancement information of the group members is obtained by processing the target speech enhancement model.

[0114] Step S104, inputting the group speech data into the target voiceprint recognition model for voiceprint recognition processing to obtain target speech person information corresponding to the group speech data;

[0115] Among them, the target speech person information is the speaker identity information corresponding to the group speech data obtained by voiceprint recognition processing by the target voiceprint recognition model.

[0116] Step S105, calculating the target speech enhancement information and the target speech character information to obtain the group learning participation degree;

[0117] In some embodiments, step S105 can include: calculating the target active character quantity and the target speech character quantity in the target speech character information within the target time period; and taking the proportion of the target active character quantity and the target speech character quantity as the group learning participation degree.

[0118] The target time period refers to the time range considered when calculating the group learning participation degree, and can be set according to actual needs, such as the length of a class or the length of a discussion session. The setting of the target time period is conducive to focusing on analyzing the participation of students in a specific time period, so as to more accurately evaluate the group learning participation degree.

[0119] The target active character quantity is the number of students actively participating in classroom discussion or speaking identified through the target speech character information within the target time period. The target active character (student) usually speaks frequently, asks questions or participates in discussion in the classroom, showing high activity and participation. The target speech character quantity is the number of all students participating in the classroom obtained through the target speech character information within the target time period.

[0120] Optionally, the group learning participation degree can also be represented by the group activity degree. The group learning participation degree is an index for measuring the overall participation and interaction of students in a teaching classroom, which is of great significance for subsequent analysis of the group learning atmosphere.

[0121] In a specific implementation, for the calculation of the group learning participation degree, the proportion of the number of active members in the group within a period of time is used as the result. For example, set the analysis window length T, and calculate the group learning participation degree according to the group member speech information output by the speech enhancement model within the window (from t-T to the current time t). The analysis process is performed in real time, and each time point is analyzed with a T-time span as the analysis window. The analysis process can obtain a continuously changing group learning participation degree.

[0122] For example, refer to Figure 4 , Figure 4 is a framework schematic diagram of a group learning participation degree calculation process provided by an embodiment of the present application; as Figure 4 shown, the calculation process of the group learning participation degree is: inputting the group speech data into the voiceprint recognition model for voiceprint recognition processing to obtain the group member information corresponding to the group speech data; and calculating the speech enhancement information output by the speech enhancement model and the group member information to obtain the group learning participation degree.

[0123] Step S106, feature extraction processing is performed on the target speech enhancement information to obtain target speech feature information.

[0124] In some embodiments, step S106 can include: performing audio feature extraction processing on the target speech enhancement information to obtain an audio feature vector; performing text feature extraction processing on the target speech enhancement information to obtain a text feature vector corresponding to the audio feature vector; and performing splicing processing on the audio feature vector and the text feature vector to obtain the target speech feature information.

[0125] The audio feature vector is a numerical representation obtained after feature extraction processing of an audio signal. The text feature extraction processing is a process of extracting text data from an audio signal and then extracting features representing the content or semantic information of the text data, i.e., obtaining a text feature vector. The text feature vector is a numerical representation obtained after feature extraction processing of the text data. For the target speech feature information, it is the fusion data obtained by splicing the audio feature vector and the text feature vector.

[0126] In other embodiments, it can also include: selecting a target speech processing scheme from the speech information processing schemes according to the teaching environment information; if the target speech processing scheme is a cloud processing scheme in the speech information processing schemes, constructing a target cloud processing model based on the cloud processing scheme; obtaining group speech data in a teaching classroom; inputting the group speech data into the target cloud processing model for data processing to obtain target cloud speech information and target speech person information corresponding to the group speech data; calculating the target cloud speech information and the target speech person information to obtain group learning participation; and performing feature extraction processing on the target cloud speech information to obtain target speech feature information.

[0127] It is easy to understand that if there is no condition or requirement for deploying a local voiceprint recognition model in the teaching implementation environment, data preprocessing is not required, a cloud online speaker diarization model is directly deployed, and then real-time group speech data in a teaching classroom is input into the cloud online speaker diarization model for data processing to obtain target cloud speech information and target speech person information corresponding to the group speech data; the target cloud speech information and the target speech person information are calculated to obtain group learning participation; and feature extraction processing is performed on the target cloud speech information to obtain target speech feature information.

[0128] In a specific implementation, an audio feature vector in the target speech enhancement information is extracted, and a text feature vector corresponding to the target speech enhancement information is extracted. By way of example, the openSmile open source toolkit is used to extract the LLDS (Low Level Descriptors) feature, and then text recognition processing is performed on the obtained speech information to obtain text data, and a text feature vector corresponding to the speech information is extracted from the text data.

[0129] In step S107, the target speech feature information is input into a target emotion prediction model for emotion prediction to obtain target emotion information.

[0130] Optionally, the target emotion prediction model is a pre-trained model specially configured to receive speech feature information (such as a combination of an audio feature vector and a text feature vector) as input and output corresponding emotion prediction results. Through learning and understanding a large amount of speech data with emotion labels, the target emotion prediction model can capture the emotion features contained in the speech signal and perform emotion prediction on new speech input.

[0131] The target emotion information is an output result obtained by performing emotion prediction on the target speech feature information by the target emotion prediction model.

[0132] In a specific implementation, please refer to Figure 5 , Figure 5 is a structural schematic diagram of an emotion analysis and recognition module provided by an embodiment of the present application. The emotion analysis and recognition module includes a speech pickup unit, a feature extraction unit, and an emotion detection unit. The speech pickup unit is configured to pick up speech information in group learning. The feature extraction unit is configured to extract an audio feature vector in the speech information and extract a text feature vector corresponding to the speech information. The emotion detection unit is configured to detect the emotion of a group member in group learning according to the audio feature vector and the text feature vector, and determine an emotion type. Specifically, speech information output by a speech enhancement model is first obtained, then feature extraction is performed on the speech information, then the extracted audio feature vector and text feature vector are input into an emotion prediction model for prediction, and a prediction result is obtained. The emotion type of the student is determined according to the prediction result. The emotion type of the student can be summarized as a positive emotion type, a negative emotion type, and a neutral emotion type.

[0133] In step S108, the group learning participation and the target emotion information are analyzed and processed to obtain group learning atmosphere data.

[0134] The group learning atmosphere data is used to reflect the atmosphere of group (team) learning in a teaching classroom. The group learning atmosphere data is a multi-dimensional and quantitative evaluation result, which comprehensively reflects the atmosphere of group learning in a teaching classroom by integrating group learning participation and emotional information. The group learning atmosphere data is not only beneficial for teachers to understand the performance of students in the classroom, but also provides strong support for teaching improvement. Meanwhile, group members can also understand their learning state in time and make self-adjustment according to the group learning atmosphere data.

[0135] For group learning team atmosphere degree analysis, an open-source emotion prediction model is used to predict the emotion type of each group member. According to the proportion of members with positive, negative and neutral emotion types in the total members in the group, and in combination with learning participation, the group learning atmosphere of the group is obtained.

[0136] It should be noted that the emotion prediction model can refer to the implementation principle in the related art. The emotion prediction model selected by the embodiments of the present application is not limited, and the embodiments of the present application will not be repeated here.

[0137] Exemplarily, the group discussion atmosphere is obtained in combination with the group learning participation and the emotion type of the group members. The discussion atmosphere is only set to two grades, namely, excellent and poor, for intuitive and convenient classroom observation and adjustment. When the group learning participation is less than 50%, the discussion atmosphere is poor. When the group learning participation is greater than 50% and the number of members with positive and neutral emotions is less than the number of members with negative emotions, the discussion atmosphere is also poor. When the group learning participation is greater than 50% and the number of members with positive and neutral emotions is greater than the number of members with negative emotions, the discussion atmosphere is excellent. It should be noted that the number of atmosphere grades can be increased according to the actual situation and the corresponding calculation method is adapted, and the embodiments of the present application are not limited thereto.

[0138] It should be noted that after step S108, the following step content is further included: visualizing the data through a specific device to realize real-time feedback in the classroom. Specifically, through a color LED (Light-Emitting Diode) device, the group learning participation analysis result is mapped to a series of colors, the group participation change in the group learning classroom is represented by light change, and the group learning participation and group learning atmosphere degree analysis results in the classroom are displayed in real time through a display device. Based on the real-time feedback effect of voice-based group learning, the teacher can understand the classroom situation in real time and conveniently. Finally, the teacher can adjust and intervene the classroom teaching according to the feedback result. Specifically, the teacher can understand the teaching situation of the group learning classroom at any time according to the visual real-time feedback function, and intervene the classroom in real time according to the needs, for example, guiding the group with low participation to discuss, understanding the problems encountered by the group with poor atmosphere and giving suggestions, etc. At the same time, the group members can also understand their learning state in time and make self-adjustment according to the group learning atmosphere data.

[0139] Please refer to Figure 6 , Figure 6 is a schematic diagram of the interface content of a visualization module provided by the embodiment of the present application, as shown in Figure 6 The feedback page content in the visualization module includes a color LED device (indicated by letter A in Figure 6 ), which maps the group learning participation analysis result to a series of colors, represents the group participation change in the group learning classroom by light change, and displays the group learning participation and group learning atmosphere degree analysis results in the classroom in real time through a display device. As shown in Figure 6 , the yellow part in Figure 6 indicates low group participation, and the green part indicates high group participation. When the yellow color gradually deepens until the green part, it indicates that the group participation changes from low to high, as shown in Figure 6 , the LED device color in Figure 6 is green. According to the color change comparison standard, it indicates that the group participation is high at this time. Combined with the emotion information, the "discussion atmosphere: excellent" in Figure 6 can be obtained. The visualization interface content further includes the content of "number of members participating in the discussion: N people", "inactive members: Zhang San, Li Si". Through the visual real-time feedback function, the teacher can understand the teaching situation of the group learning classroom at any time, and intervene the classroom in real time according to the needs.

[0140] The steps S101 to S108 shown in the embodiments of the present application are to obtain a voice information processing scheme, and construct a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme; obtain group voice data in a teaching classroom; input the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information; input the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data; calculate the target voice enhancement information and the target voice character information to obtain group learning participation degree; perform feature extraction processing on the target voice enhancement information to obtain target voice feature information; input the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information; and analyze and process the group learning participation degree and the target emotion information to obtain group learning atmosphere data. The embodiments of the present application perform enhancement processing on the group voice data through the target voice enhancement model, improve the clarity and purity of the voice signal, and thus improve the accuracy of subsequent emotion prediction and group learning participation degree analysis; the target voiceprint recognition model is used to perform voiceprint recognition on the group voice data, which can accurately identify the voice of each group member, and provides the possibility for subsequent personalized feedback and intervention; the group learning atmosphere data obtained through analysis of the group learning participation degree combined with the emotion prediction result is comprehensive evaluation, so that the group members and teachers can quickly capture the learning situation in the classroom.

[0141] To explain the principle of the technical scheme of the present application in detail, the overall process of the present application will be described below in combination with some specific embodiments. It should be easily understood that the following is an explanation of the technical principle of the present application and cannot be regarded as a limitation of the present application.

[0142] Please refer to Figure 7 , Figure 7 is a flowchart of the voice-based group learning atmosphere analysis method provided by the embodiments of the present application; please refer to Figure 8 , Figure 8 is a neural network application implementation schematic diagram of the voice-based group learning atmosphere analysis method provided by the embodiments of the present application; in combination with Figure 7 and Figure 8 , the implementation process of the voice-based group learning atmosphere analysis method can be obtained, as shown in Figure 7 , the overall process of the voice-based group learning atmosphere analysis method includes the following six steps (steps 701-706):

[0143] Step 701, according to the teaching environment information, a voice information processing scheme is selected;

[0144] In a specific implementation, the speech information processing scheme includes a local processing scheme and a cloud processing scheme. The local processing scheme is a scheme including a training data preprocessing process and a model construction process. After training data preprocessing and model construction, a target speech enhancement model and a target voiceprint recognition model can be obtained. The target speech enhancement model and the target voiceprint recognition model process real-time teaching classroom speech data to obtain a group learning atmosphere. The cloud processing scheme is a scheme of directly deploying a cloud processing model (a cloud online speaker diarization model), and then processing real-time teaching classroom speech data according to the cloud processing model to obtain a group learning atmosphere.

[0145] Specifically, as shown in Figure 8 First, the selection of the speech information processing scheme is needed. It is necessary to determine whether the teaching implementation environment has the condition or requirement of deploying a local voiceprint recognition model. If the teaching implementation environment does not have the condition or requirement of deploying a local voiceprint recognition model (i.e., the cloud processing scheme is used), preprocessing is not needed, and the next implementation step 702 is directly entered. If the teaching implementation environment has the condition or requirement of deploying a local voiceprint recognition model (i.e., the local processing scheme is used), a data preprocessing process is needed. The data preprocessing process specifically includes: first, pre-recording student speech information, then using a mel-frequency cepstral coefficient to extract features of the student speech information, inputting the extracted speech features into a Mamba model for training to obtain a voiceprint recognition model, and the voiceprint recognition model including a student voiceprint encoder, the student voiceprint encoder being used to obtain student voiceprint information, and multiple voiceprint information being used to construct a voiceprint information library.

[0146] Step 702: According to the selection type of the speech information processing scheme, deploy a cloud online speaker diarization model or train a local speech enhancement model and a local voiceprint recognition model.

[0147] According to the selection of the speech information processing scheme in step 701, deploy a cloud online speaker diarization model or train a local speech enhancement model and a local voiceprint recognition model. Specifically: if a non-local voiceprint recognition scheme is selected (i.e., the cloud processing scheme is used), a cloud online speaker diarization model is deployed; if a local voiceprint recognition scheme is selected (i.e., the local processing scheme is used), a local speech enhancement model and a local voiceprint recognition model are trained.

[0148] Step 703, acquiring real-time speech in the classroom, according to the selection type of the speech information processing scheme, using a cloud processing model to process the real-time speech or using a local speech enhancement model and a local voiceprint recognition model to perform speech enhancement processing and voiceprint recognition processing, to obtain processed speech information;

[0149] Specifically, as shown in Figure 8 In a teaching classroom, members of the same student group sit together around a table, and an array microphone device is placed in the center of the table. Real-time speech collection is performed using the microphone array to obtain real-time group speech data in the teaching classroom.

[0150] As shown in Figure 8 If a local processing scheme is adopted, the real-time group speech data in the teaching classroom is input as input data into a local speech enhancement model, and speech enhancement information of the group members is obtained by processing the local speech enhancement model. At the same time, the group speech data is input into a local voiceprint recognition model for voiceprint recognition processing to obtain speech person information corresponding to the group speech data.

[0151] As shown in Figure 8 If a cloud processing scheme is adopted, the real-time group speech data in the teaching classroom is input into a cloud online speaker diarization model for data processing to obtain cloud speech information and target speech person information corresponding to the group speech data.

[0152] Step 704, according to the speech information output in step 703, using an algorithm to calculate the group learning participation degree, and performing emotional analysis on the group learning group members, combining the group learning participation degree and the emotional analysis result to calculate the group learning atmosphere;

[0153] Specifically, for the calculation of the group learning participation degree, the proportion of active members in the group to the total number of members in the group within a period of time is used as the result.

[0154] As shown in Figure 8 If a cloud processing scheme is adopted, the cloud speech information and target speech person information corresponding to the group speech data obtained by processing the cloud online speaker diarization model are calculated to obtain the group learning participation degree. The cloud speech information is processed by audio feature extraction and text feature extraction to obtain speech feature information.

[0155] As shown in Figure 8 If a local processing scheme is adopted, the speech enhancement information of the group members output by the speech enhancement model and the speech person information obtained by the voiceprint recognition model are used to calculate the group learning participation degree. The speech enhancement information is processed by audio feature extraction and text feature extraction to obtain speech feature information.

[0156] For group learning group atmosphere analysis, the speech feature information obtained by each scheme is subjected to emotion prediction using an open-source emotion prediction model, the emotion type of each group member is predicted, and according to the proportion of members of positive emotion type, negative emotion type and neutral emotion type in the total members in the group, combined with learning participation, the group learning atmosphere of the group is obtained.

[0157] Step 705, the data is visualized by a specific device to realize real-time feedback of the classroom;

[0158] Specifically, as shown in Figure 8 , the group learning participation analysis result is mapped to a series of colors by a color LED device, the group learning participation change in the classroom is represented by light change, and the classroom group learning participation and group learning atmosphere analysis result are displayed in real time by a display device. The real-time feedback effect of the speech-based group learning enables the teacher to understand the classroom situation in real time and conveniently.

[0159] Step 706, the teacher adjusts and intervenes in the classroom teaching according to the feedback result.

[0160] Specifically, as shown in Figure 8 , the teacher can adjust and intervene in the classroom teaching according to the feedback result. Specifically, the teacher can understand the teaching situation of the group learning classroom at any time according to the visual real-time feedback function, and can intervene in the classroom in real time according to the needs, for example, guiding the group with low participation to discuss, understanding the problems encountered by the group with poor group atmosphere and giving suggestions, etc. At the same time, the group members can also understand their learning state in time according to the group learning atmosphere data and make self-adjustment.

[0161] It should be pointed out that the present embodiment only briefly and schematically illustrates the general process of the speech-based group learning atmosphere analysis method, and the detailed description of each step can refer to the related content in the foregoing embodiments, which will not be repeated here. It can be understood that the present application does not limit this.

[0162] The embodiment of the application obtains a voice information processing scheme, constructs a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme, obtains group voice data in a teaching classroom, inputs the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information, inputs the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice person information corresponding to the group voice data, calculates the target voice enhancement information and the target voice person information to obtain group learning participation, performs feature extraction processing on the target voice enhancement information to obtain target voice feature information, inputs the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information, and analyzes and processes the group learning participation and the target emotion information to obtain group learning atmosphere data. The embodiment of the application performs enhancement processing on the group voice data by using the target voice enhancement model, improves the clarity and purity of the voice signal, and thus improves the accuracy of subsequent emotion prediction and group learning participation analysis. The target voiceprint recognition model is used to perform voiceprint recognition on the group voice data, which can accurately identify the voice of each group member and provides the possibility for subsequent personalized feedback and intervention. The group learning atmosphere data obtained by analyzing the group learning participation combined with the emotion prediction result is comprehensive, so that the group members and teachers can quickly capture the learning situation in the classroom. At the same time, the group members can understand their own learning state in time and make self-adjustment according to the group learning atmosphere data, and the teachers can also quickly capture the dynamic changes in the classroom and take timely intervention measures, which can assist the teachers in optimizing the teaching strategy and improve the classroom teaching effect.

[0163] Please refer to Figure 9 The embodiment of the application also provides a voice-based group learning atmosphere analysis system 900, which can implement the voice-based group learning atmosphere analysis method. The voice-based group learning atmosphere analysis system includes the following modules:

[0164] A model construction module 901 is configured to obtain a voice information processing scheme, and construct a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme.

[0165] A group voice data acquisition module 902 is configured to obtain group voice data in a teaching classroom.

[0166] A data enhancement processing module 903 is configured to input the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information.

[0167] A voiceprint recognition processing module 904 is configured to input the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice person information corresponding to the group voice data.

[0168] The group learning participation degree calculation module 905 is configured to calculate the target voice enhancement information and the target voice character information to obtain a group learning participation degree.

[0169] The feature extraction processing module 906 is configured to perform feature extraction processing on the target voice enhancement information to obtain target voice feature information.

[0170] The emotion information prediction module 907 is configured to input the target voice feature information into a target emotion prediction model to perform emotion prediction and obtain target emotion information.

[0171] The group learning atmosphere analysis module 908 is configured to analyze and process the group learning participation degree and the target emotion information to obtain group learning atmosphere data.

[0172] It can be understood that the content in the above method embodiments is applicable to the present system embodiment, the present system embodiment specifically implements the same functions as the above method embodiments, and achieves the same beneficial effects as the above method embodiments.

[0173] The present application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned group learning atmosphere analysis method based on voice when executing the computer program. The electronic device can be any intelligent terminal, such as a tablet computer or a vehicle-mounted computer.

[0174] It can be understood that the content in the above method embodiments is applicable to the present device embodiment, the present device embodiment specifically implements the same functions as the above method embodiments, and achieves the same beneficial effects as the above method embodiments.

[0175] Please refer to Figure 10 , Figure 10 The hardware structure of the electronic device of another embodiment is illustrated, which includes:

[0176] The processor 1001 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute related programs to implement the technical solutions provided by the present application.

[0177] The memory 1002 can be implemented in the form of a Read-Only Memory (ROM), a static storage device, a dynamic storage device, or a Random Access Memory (RAM), etc. The memory 1002 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1002 and are called and executed by the processor 1001 to perform the voice-based group learning atmosphere analysis method of the embodiments of the present application;

[0178] The input / output interface 1003 is configured to realize information input and output.

[0179] The communication interface 1004 is configured to realize the communication interaction between the device and other devices, and the communication can be realized by a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).

[0180] The bus 1005 is configured to transmit information between various components (for example, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004) of the device.

[0181] The processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are connected to each other through the bus 1005 to realize the communication connection between the devices.

[0182] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the voice-based group learning atmosphere analysis method.

[0183] It can be understood that the contents in the above method embodiments are applicable to the present storage medium embodiments, the functions realized by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved by the present storage medium embodiments are also the same as those of the above method embodiments.

[0184] The memory is a non-transitory computer readable storage medium, which can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0185] The method and system for analyzing group learning atmosphere based on voice provided by the embodiments of the present application obtain a voice information processing scheme, and construct a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme; obtain group voice data in a teaching classroom; input the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information; input the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data; calculate the target voice enhancement information and the target voice character information to obtain group learning participation; perform feature extraction processing on the target voice enhancement information to obtain target voice feature information; input the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information; and analyze and process the group learning participation and the target emotion information to obtain group learning atmosphere data. The embodiments of the present application improve the clarity and purity of the voice signal by performing enhancement processing on the group voice data through the target voice enhancement model, thereby improving the accuracy of subsequent emotion prediction and group learning participation analysis; the target voiceprint recognition model is used to perform voiceprint recognition on the group voice data, which can accurately identify the voice of each group member, thereby providing the possibility for subsequent personalized feedback and intervention; the group learning atmosphere data obtained by analyzing the group learning participation combined with the emotion prediction result can be comprehensively evaluated, thereby enabling the group members and the teacher to quickly capture the learning situation in the classroom.

[0186] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0187] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than the figures, or combine certain steps or different steps.

[0188] The system embodiments described above are only schematic, and the units described as separate components can or can not be physically separate, that is, can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0189] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the function modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.

[0190] The terms "first", "second", "third", "fourth", and the like in the description of this application and in the claims hereof, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so termed herein is solely for descriptive purposes and not for pronouncing the limitations of the application described except as described in the claims. Moreover, the terms "comprise", "have", "contain" and "include" and any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, system, product or apparatus that comprises, has, contains or includes a list of steps or elements, but not those not expressly listed or inherent to such process, method, system, product or apparatus, is not excluded from the scope of this application.

[0191] It should be understood that, in this application, "at least one" means one or more, "multiple" means two or more. "And / or" is used to describe the relationship between associated objects, which means that there can be three relationships, for example, "A and / or B" can mean: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including single or multiple combinations. For example, at least one of a, b or c, can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0192] In several embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the above-described system embodiments are only illustrative, for example, the division of the above-mentioned units is only a logical functional division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between systems or units, which can be electrical, mechanical or other forms.

[0193] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the embodiment scheme.

[0194] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0195] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions used to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various other media that can store programs.

[0196] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and are not limited to the scope of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.

Claims

1. A voice-based group learning atmosphere analysis method, characterized by, The method comprises the following steps: obtain a voice information processing scheme, and construct a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme; obtain group voice data in a teaching classroom; input the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information; input the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data; calculate the target voice enhancement information and the target voice character information to obtain group learning participation; perform feature extraction processing on the target voice enhancement information to obtain target voice feature information; input the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information; analyze and process the group learning participation and the target emotion information to obtain group learning atmosphere data.

2. The method of claim 1, wherein, The voice information processing scheme includes a local processing scheme, and the obtaining of the voice information processing scheme and the construction of the target voice enhancement model and the target voiceprint recognition model based on the voice information processing scheme comprises: obtain the voice information processing scheme and training voice data; select a target voice processing scheme from the voice information processing scheme according to teaching environment information; if the target voice processing scheme is the local processing scheme in the voice information processing scheme, construct the target voice enhancement model and the target voiceprint recognition model according to the training voice data.

3. The method of claim 2, wherein, If the target voice processing scheme is the local processing scheme in the voice information processing scheme, the construction of the target voice enhancement model and the target voiceprint recognition model according to the training voice data comprises: if the target voice processing scheme is the local processing scheme in the voice information processing scheme, perform information extraction on the training voice data to obtain pure voice data; add a reverberation effect in the pure voice data to obtain reverberation voice data; integrate the reverberation voice data and a plurality of target noise segments to obtain target training data; input the target training data into an initial voice enhancement model for training to obtain the target voice enhancement model; if the target voice processing scheme is the local processing scheme in the voice information processing scheme, perform feature extraction processing on the training voice data to obtain target training feature information; input the target training feature information into a linear time series model for training to obtain the target voiceprint recognition model; the target voiceprint recognition model is used to obtain voiceprint information of a target voice character, and the voiceprint information is used to construct a voiceprint model library.

4. The method of claim 3, wherein, The inputting of the target training feature information into the linear time series model for training to obtain the target voiceprint recognition model comprises: input the target training feature information into the linear time series model for training to obtain an initial embedding vector; perform regularization processing on the initial embedding vector to obtain a target embedding vector; calculate a centroid of the target embedding vector to obtain a target vector centroid; construct a similarity matrix based on the target embedding vector and the target vector centroid; train the linear time sequence model based on the similarity matrix and a target voiceprint loss function to obtain the target voiceprint recognition model.

5. The method of claim 1, wherein, The calculation of the target voice enhancement information and the target voice character information obtains a group learning participation degree, including: calculating a target active character quantity and a target voice character quantity in the target voice character information within a target time period; regarding the proportion of the target active character quantity and the target voice character quantity as the group learning participation degree.

6. The method of claim 1, wherein, The feature extraction processing of the target voice enhancement information obtains target voice feature information, including: performing audio feature extraction processing on the target voice enhancement information to obtain an audio feature vector; performing text feature extraction processing on the target voice enhancement information to obtain a text feature vector corresponding to the audio feature vector; performing splicing processing on the audio feature vector and the text feature vector to obtain the target voice feature information.

7. The method of claim 1, wherein, The voice information processing scheme includes a cloud processing scheme, and the method further includes: selecting a target voice processing scheme from the voice information processing scheme according to teaching environment information; if the target voice processing scheme is the cloud processing scheme in the voice information processing scheme, constructing a target cloud processing model based on the cloud processing scheme; obtaining the group voice data in the teaching classroom; inputting the group voice data into the target cloud processing model for data processing to obtain target cloud voice information and the target voice character information corresponding to the group voice data; calculating the target cloud voice information and the target voice character information to obtain the group learning participation degree; performing feature extraction processing on the target cloud voice information to obtain the target voice feature information.

8. A voice-based group learning atmosphere analysis system, characterized by, The system includes the following modules: a model construction module for obtaining a voice information processing scheme and constructing a target voice enhancement model and a target voiceprint recognition model based on the voice information processing scheme; a group voice data acquisition module for obtaining group voice data in a teaching classroom; a data enhancement processing module for inputting the group voice data into the target voice enhancement model for data enhancement processing to obtain target voice enhancement information; a voiceprint recognition processing module for inputting the group voice data into the target voiceprint recognition model for voiceprint recognition processing to obtain target voice character information corresponding to the group voice data; a group learning participation degree calculation module for calculating the target voice enhancement information and the target voice character information to obtain a group learning participation degree; a feature extraction processing module for performing feature extraction processing on the target voice enhancement information to obtain target voice feature information; an emotion information prediction module for inputting the target voice feature information into a target emotion prediction model for emotion prediction to obtain target emotion information; a group learning atmosphere analysis module for analyzing and processing the group learning participation degree and the target emotion information to obtain group learning atmosphere data.

9. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, the computer-readable storage medium comprising: The computer program is executed by the processor to implement the method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal emotion emergency decision-making system based on multi-stage long short-term memory network

    CN115393927A

  • Atmosphere detection method and device, electronic equipment and computer readable storage medium

    CN115965895A