Electronic whiteboard function separation method and system based on voice authority management
Through the electronic whiteboard function separation method based on voice permission management, using deep learning algorithms and natural language processing technology, precise classification and permission management of teachers, students and unknown personnel is realized, solving the problem of low authority management efficiency in the existing electronic whiteboard system, and improving the convenience of teaching efficiency and functional control.
Patent Information
- Application Number
- CN202510105364.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
The existing electronic whiteboard system has problems such as confusing operational permissions and low efficiency of traditional control methods in terms of functional management, which may cause teachers and students to cross the boundaries of permissions during use, affecting teaching order.
The electronic whiteboard function separation method based on voice permission management is adopted, and the speech features are extracted through deep learning algorithms, and the user identity judgment is determined using the Mel frequency cepspectral coefficient and Gaussian hybrid model, permission rules are loaded dynamically, and voice instructions are parsed through natural language processing to perform corresponding functions.
Accurate classification and authority management for teachers, students and unknown personnel is realized, classroom interference is reduced, teaching efficiency is improved, and the convenience and efficiency of functional control is improved through voice control.
Smart Images

Figure CN119993158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic whiteboards, and in particular to an electronic whiteboard function separation method and system based on voice authority management. Background Art
[0002] With the rapid development of modern educational informatization, electronic whiteboards, as an important tool for classroom teaching, have been widely used in various educational scenarios. Electronic whiteboards integrate functions such as multimedia display, interactive teaching and real-time feedback, which effectively improves the teaching quality and classroom interaction efficiency. However, in actual use, the functional management of electronic whiteboards faces the following problems: 1. Confusion of operating permissions: The existing electronic whiteboard system lacks effective distinction and management of operating permissions, which may lead to teachers and students crossing the boundaries of permissions when using it. For example, students' mistaken operation of high-authority functions (such as clearing content, screen sharing, etc.) may affect the normal teaching order. 2. Traditional control methods are inefficient: Many electronic whiteboards currently rely on physical buttons or touch screen control. This method is inefficient in complex teaching scenarios, especially when teachers need to switch functions frequently, which is easy to distract attention and reduce classroom management efficiency. Summary of the invention
[0003] To solve the above problems, the present invention provides an electronic whiteboard function separation method and system based on voice permission management, which further separates and responds to the functions of the electronic whiteboard by extracting voice features through a deep learning algorithm, thereby reducing classroom interference and improving teaching efficiency.
[0004] To achieve the above object, the technical solution adopted by the present invention is:
[0005] A method for separating functions of an electronic whiteboard based on voice authority management, comprising:
[0006] Capture user voice data in real time and perform data cleaning and standardization;
[0007] Extract features from the standardized user voice data based on the Mel-frequency cepstral coefficients to generate a feature vector;
[0008] Based on the feature vector, determining the user identity by using a Gaussian mixture model;
[0009] The electronic whiteboard function is loaded according to the user identity, the user instructions are parsed from the standardized voice data through natural language processing, and the electronic whiteboard functions are traversed and matched according to the user instructions and executed.
[0010] Furthermore, the real-time capture of user voice data, data cleaning and standardization includes the following steps:
[0011] Convert user voice data into digital signals;
[0012] The background noise of digital signal is eliminated by adaptive filtering algorithm;
[0013] The noise-reduced speech signal is divided into frames using a signal framing algorithm, with the frame length set to 25ms and the frame shift set to 10ms.
[0014] The speech signal after the frame division processing is subjected to windowing processing.
[0015] Furthermore, the feature extraction of the standardized user voice data based on the Mel-frequency cepstral coefficients includes the following steps:
[0016] Performing spectrum analysis on the standardized speech signal through Fourier transform algorithm to generate the spectrum of the speech signal;
[0017] The frequency spectrum of the speech signal is subjected to frequency scale mapping processing through a Mel filter bank, and the frequency distribution is mapped to the Mel frequency scale;
[0018] Applying logarithmic transformation algorithm to the Mel frequency scale for energy compression;
[0019] The compressed Mel frequency scale data is processed by discrete cosine transform algorithm to extract feature vectors, generate Mel frequency cepstrum coefficients and use them as feature vectors.
[0020] Furthermore, determining the user identity by using a Gaussian mixture model comprises the following steps:
[0021] The historical speech feature vectors of teachers and students are modeled using a Gaussian mixture model. The mean vector, covariance matrix, and mixture weight of each type of data are calculated using the expectation maximization algorithm to generate a teacher model and a student model respectively.
[0022] Input the user's speech feature vector into the teacher model and the student model, and use the probability distribution function of the Gaussian mixture model to calculate the posterior probability values of the feature vector in the two models respectively;
[0023] Compare the posterior probability values of the user's voice feature vector in the teacher model and the student model. When the probability value reaches the preset teacher threshold, it is judged as a teacher; when it reaches the preset student threshold, it is judged as a student; if both are lower than the preset teacher threshold and the preset student threshold, it is judged as an unknown person.
[0024] Furthermore, the calculation formula of the Gaussian mixture model is as follows:
[0025]
[0026] Where P(x|λ) is the probability of the feature vector x under the model parameter λ; K is the number of Gaussian components; w kis the weight of the kth Gaussian component; μ k is the mean vector of the kth Gaussian component; Σ k is the covariance matrix of the kth Gaussian component; d is the dimension of the eigenvector.
[0027] Furthermore, the electronic whiteboard function is loaded according to the user identity, including:
[0028] When the user is a teacher, teacher functions are loaded, including clearing content, screen sharing, and function locking; when the user is a student, student permission rules are loaded, including the use of annotation tools and answer submission functions; when the user is an unknown person, access is restricted and the administrator is prompted.
[0029] Furthermore, parsing the user instructions from the standardized voice data by natural language processing comprises the following steps:
[0030] The standardized speech data is processed through a speech recognition model to convert the speech signal into text data;
[0031] The text is segmented and vectorized based on the pre-trained natural language model, keywords are extracted through the named entity recognition algorithm, and the semantic relationship in the text is analyzed through the syntactic dependency analysis algorithm to generate a preliminary mapping of the user's operation intention;
[0032] The extracted keywords and feature vectors are input into the Softmax classifier model to determine the specific functional module corresponding to the instruction.
[0033] Furthermore, the traversal and execution of matching electronic whiteboard functions through user instructions includes the following steps:
[0034] Based on the hash search algorithm, the keywords in the user command are compared with the mapping table of the electronic whiteboard function module;
[0035] A Boolean logic verification algorithm is used to verify whether the user's instructions comply with the permission rules of the current identity; if the verification passes, the operation interface of the corresponding functional module is triggered and the function call is triggered; if the verification fails, the operation is terminated and a prompt of insufficient permissions is returned.
[0036] Furthermore, the insufficient authority prompt is given through voice notification and electronic whiteboard screen visualization.
[0037] An electronic whiteboard function separation system based on voice authority management, applied to any of the above-mentioned electronic whiteboard function separation methods based on voice authority management, comprising a voice data acquisition module, a feature extraction module, an identity determination module and a function matching module connected in sequence;
[0038] The voice data acquisition module is used to capture user voice data in real time and perform data cleaning and standardization;
[0039] The feature extraction module is used to extract features from the standardized user voice data based on the Mel frequency cepstral coefficients to generate a feature vector;
[0040] The identity determination module is used to determine the user identity through a Gaussian mixture model based on the feature vector;
[0041] The function matching module is used to load the electronic whiteboard function according to the user identity, parse the user instruction from the standardized voice data through natural language processing, and traverse and match the electronic whiteboard function according to the user instruction and execute it.
[0042] The beneficial effect of the present invention is that the present invention realizes accurate classification of teachers, students and unknown persons through the extraction of user voice features and identity determination mechanism. Specifically, the voiceprint features of the user are extracted by Mel frequency cepstral coefficient (MFCC), and the identity is determined in combination with Gaussian mixture model (GMM). The corresponding permission rules are loaded according to the user identity, such as teachers can access advanced functions such as clearing content and screen sharing, while students are limited to annotation tools and answer submission functions. Unknown persons are restricted from access, which fundamentally solves the problems of authority crossing and misoperation, and ensures the stability of teaching order. The use of voice command parsing technology significantly improves the convenience and efficiency of electronic whiteboard function control. The user's voice command is converted into text through the voice recognition model, and then the user's operation intention is analyzed in combination with natural language processing (NLP) technology, and the electronic whiteboard function module is dynamically matched and executed. Compared with traditional physical buttons or touch screen operations, voice control is more suitable for complex teaching scenarios, does not require frequent manual operations by teachers, reduces classroom interference, and improves teaching efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a flow chart of the steps of an electronic whiteboard function separation method based on voice authority management in the present invention.
[0044] Figure 2 It is a flow chart of the steps of extracting features from standardized user voice data based on Mel-frequency cepstral coefficients in the present invention.
[0045] Figure 3 It is a structural schematic diagram of an electronic whiteboard function separation system based on voice authority management in the present invention. DETAILED DESCRIPTION
[0046] See also Figure 1-3 As shown, the present invention relates to a method for separating electronic whiteboard functions based on voice authority management, comprising:
[0047] Capture user voice data in real time and perform data cleaning and standardization;
[0048] Extract features from the standardized user voice data based on the Mel-frequency cepstral coefficients to generate a feature vector;
[0049] Based on the feature vector, determining the user identity by using a Gaussian mixture model;
[0050] The electronic whiteboard function is loaded according to the user identity, the user instructions are parsed from the standardized voice data through natural language processing, and the electronic whiteboard functions are traversed and matched according to the user instructions and executed.
[0051] Specifically, the system first uses a microphone to capture the user's voice input data in real time through the voice data acquisition module. The collected voice signal contains environmental noise and other interference information. In order to improve the data quality, the system performs noise reduction processing on the voice data. Specifically, an adaptive filtering algorithm is used to remove background noise and enhance the main voice signal. In addition, the signal needs to be frame divided and windowed. The frame length is set to 25ms and the frame shift is 10ms. The Hamming window is applied to reduce spectrum leakage to ensure that the frequency domain characteristics of the signal are clear. The voice data after cleaning and standardization is input into the feature extraction module. The system performs spectrum analysis on the signal through fast Fourier transform (FFT) and maps the spectrum to the Mel frequency scale using the Mel filter bank, which is more in line with the human hearing characteristics. Subsequently, a logarithmic transformation is performed to compress the frequency energy, and finally the Mel frequency cepstral coefficient (MFCC) is extracted through discrete cosine transform (DCT) to generate a feature vector that characterizes the user's voice characteristics. The generated feature vector is input into the identity determination module, which determines the user's identity based on the Gaussian mixture model (GMM). The Gaussian mixture model is pre-trained with voice samples of teachers and students to generate a teacher model and a student model respectively. In actual operation, the system uses the expectation maximization (EM) algorithm to calculate the posterior probability value of the input feature vector in each model to determine whether the user is a teacher, student or unknown person. If the user matches a teacher or a student, the system loads the corresponding permission rules; if the user is determined to be an unknown person, the system restricts the user's access rights and issues a prompt. After loading the permissions, the system enters the user command parsing and function execution stage. The standardized voice data is converted into text instructions through the voice recognition model. Based on the pre-trained natural language processing (NLP) model, the system performs semantic parsing on the text data to extract the user's operation intention. Taking the instruction "start screen sharing" as an example, the system extracts the two core keywords "start" and "screen sharing" through named entity recognition, and combines the context syntactic relationship analysis to confirm the user's operation intention. The parsed intention is mapped to the functional module, and the system verifies the legitimacy of the instruction through the permission verification mechanism. If the user is a teacher and has screen sharing permissions, the system calls the screen sharing module and executes the instruction, and prompts the operation status through voice and visual feedback (for example, "screen sharing has been started"). If the permission check fails, the system will prompt "Insufficient permissions".
[0052] Furthermore, the real-time capture of user voice data, data cleaning and standardization includes the following steps:
[0053] Convert user voice data into digital signals;
[0054] The background noise of digital signal is eliminated by adaptive filtering algorithm;
[0055] The noise-reduced speech signal is divided into frames using a signal framing algorithm, with the frame length set to 25ms and the frame shift set to 10ms.
[0056] The speech signal after the frame division processing is subjected to windowing processing.
[0057] In some embodiments, first, the system captures the user's voice signal through a high-precision microphone. The signal is an analog signal and needs to be converted into a digital signal through an analog-to-digital converter (ADC). This process discretizes the voice signal using a fixed sampling rate (e.g., 16kHz) to generate equally spaced digital sampling points to ensure that the signal can retain sufficient time domain and frequency domain characteristics in subsequent processing. After conversion to a digital signal, the signal may contain significant environmental noise, such as background sound, echo, or mechanical noise. In order to improve the signal quality, the system applies an adaptive filtering algorithm to eliminate background noise on the digital signal. Specifically, the filter uses the LMS (Least Mean Squares) adaptive algorithm to dynamically adjust the filter weight according to the input reference noise signal, thereby minimizing the correlation between the signal and the target noise, thereby achieving the effect of enhancing the voice signal. This process not only significantly improves the signal-to-noise ratio (SNR) of the signal, but also maximizes the retention of the semantic and spectral characteristics of the voice. The voice signal after noise reduction is then divided into several frames for subsequent short-term feature analysis. The system uses a signal framing algorithm to divide the continuous voice data according to a frame length of 25ms and a frame shift of 10ms. The role of framing is to capture the short-term stationary characteristics of speech signals, and at the same time, frame overlap is achieved through frame shift settings, thereby enhancing the information of time continuity. Each frame of data contains a relatively static speech feature, which is suitable for subsequent spectrum and feature analysis. The window function uses the Hamming window. The design feature of the Hamming window is that the weight value at the boundary is small, while the weight value in the center is close to 1. This distribution allows the main information of the signal in the frame to be retained, and the discontinuity of the boundary is smoothed. Specifically, the system performs a windowing operation point by point on each frame signal, that is, multiplying each weight value of the window function by the corresponding speech signal sampling point to generate a windowed speech frame. For example, for a frame of speech signal containing 160 sampling points, the system calculates the corresponding 160 weight values according to the mathematical expression of the Hamming window, and performs a windowing operation point by point. Compared with the unwindowed signal, the amplitude transition of the processed frame signal at the boundary is smoother, avoiding the interference caused by the boundary effect during spectrum analysis, thereby providing a more stable basis for subsequent feature extraction. Through windowing, the frequency domain information of the speech signal can be more efficiently retained and expressed, ensuring that the frequency distribution characteristics of the speech can be accurately captured in the short-time Fourier transform (STFT), providing better input data for feature extraction and classification in subsequent steps. This process is a crucial link in the entire speech signal processing chain and has a significant effect on improving the overall performance of the system.
[0058] Furthermore, the feature extraction of the standardized user voice data based on the Mel-frequency cepstral coefficients includes the following steps:
[0059] Performing spectrum analysis on the standardized speech signal through Fourier transform algorithm to generate the spectrum of the speech signal;
[0060] The frequency spectrum of the speech signal is subjected to frequency scale mapping processing through a Mel filter bank, and the frequency distribution is mapped to the Mel frequency scale;
[0061] Applying logarithmic transformation algorithm to the Mel frequency scale for energy compression;
[0062] The compressed Mel frequency scale data is processed by discrete cosine transform algorithm to extract feature vectors, generate Mel frequency cepstrum coefficients and use them as feature vectors.
[0063] In some embodiments, first, the system performs spectrum analysis on the standardized speech signal. The speech signal is converted from the time domain to the frequency domain through the fast Fourier transform (FFT) algorithm to generate the spectrum of the speech signal. The core of the FFT algorithm is to decompose the frequency components of the signal, which is specifically implemented by calculating the discrete Fourier transform (DFT) using a fast recursive algorithm, significantly improving the computational efficiency. The generated spectrum data provides the frequency distribution and amplitude information of the signal, which is the basis for subsequent processing. Then, the system performs frequency scale mapping processing on the spectrum data through a Mel filter bank. The Mel filter bank is a group of bandpass filters whose center frequencies are evenly distributed according to the Mel scale, which conforms to the frequency resolution characteristics of human hearing. Specifically, the system divides the spectrum into several frequency bands (usually 20-40), and accumulates the energy of each frequency band to generate the energy distribution under the Mel frequency scale. This step effectively simulates the auditory perception characteristics of the human ear and enhances the characterization ability of speech features. Subsequently, the energy distribution of the Mel frequency scale is subjected to energy compression processing by applying a logarithmic transformation algorithm. The role of logarithmic transformation is to compress the high-amplitude energy in the spectrum and amplify the low-amplitude energy, making the dynamic range of the signal more suitable for subsequent processing. This nonlinear transformation helps to extract the structural features of the signal, while reducing the impact of environmental noise and changes in signal amplitude on feature extraction. Finally, the system uses the discrete cosine transform (DCT) algorithm to extract the feature vector of the compressed Mel-frequency scale data. The purpose of DCT is to further convert the frequency domain data to the cepstrum domain, remove the high-order redundant information in the feature vector, and retain the main low-order features. Specifically, DCT converts the input signal into a set of cepstrum coefficients by calculating the projection of a set of orthogonal bases. Finally, the system selects the first several cepstrum coefficients as Mel-frequency cepstrum coefficients (MFCCs), which highly summarize the characteristics of the speech signal. For example, suppose the system processes a speech frame containing 256 sampling points: the spectrum data generated by FFT is mapped to 24 Mel-frequency filters, 24 energy values are obtained after logarithmic transformation, and finally 13 Mel-frequency cepstrum coefficients are generated by DCT. These coefficients will be used as feature vectors and input into the identity determination module.
[0064] Furthermore, determining the user identity by using a Gaussian mixture model comprises the following steps:
[0065] The historical speech feature vectors of teachers and students are modeled using a Gaussian mixture model. The mean vector, covariance matrix, and mixture weight of each type of data are calculated using the expectation maximization algorithm to generate a teacher model and a student model respectively.
[0066] Input the user's speech feature vector into the teacher model and the student model, and use the probability distribution function of the Gaussian mixture model to calculate the posterior probability values of the feature vector in the two models respectively;
[0067] Compare the posterior probability values of the user's voice feature vector in the teacher model and the student model. When the probability value reaches the preset teacher threshold, it is judged as a teacher; when it reaches the preset student threshold, it is judged as a student; if both are lower than the preset teacher threshold and the preset student threshold, it is judged as an unknown person.
[0068] In some embodiments, during the model training phase, the system uses the historical speech data of teachers and students for feature modeling. After each voice data is extracted by Mel Frequency Cepstral Coefficient (MFCC) features, a corresponding feature vector is generated, which characterizes the acoustic features of the user's voice. The system initializes the parameters of the Gaussian mixture model, including the mean vector, covariance matrix and mixture weight of each Gaussian component. Subsequently, the model parameters are optimized by the expectation maximization (EM) algorithm. The EM algorithm performs two key steps in each iteration: first, the system calculates the probability distribution of each voice feature vector belonging to each Gaussian component, which is called the "posterior probability"; then, based on the posterior probability, the parameters of each component are updated to gradually improve the model's fit to the training data. After multiple iterations, the model finally converges and generates a teacher model and a student model respectively. These models can effectively represent the distribution of the voice features of teachers and students. In the identity determination phase, the system receives the user's voice data captured in real time and extracts its corresponding feature vector. The feature vector is input into the teacher model and the student model in turn. The system calculates the probability value of the feature vector in each model, which represents the degree of matching between the voice feature and the model. For example, if the matching probability of the input voice feature in the teacher model is higher than that in the student model and exceeds the preset threshold of the teacher model, the user is determined to be a teacher. Similarly, if the probability in the student model is higher and exceeds the threshold of the student model, the user is determined to be a student. If the matching probability of the input voice feature in both models is lower than their respective thresholds, the system determines the user's identity as an unknown person. For example, the voice feature input by a user is determined to have a probability of 85% in the teacher model and a probability of 40% in the student model. At this time, if the judgment threshold of the teacher model is 80%, the system identifies the user as a teacher. If the probabilities are both lower than the model thresholds (such as 40% and 30%), the system identifies the user as an unknown person, restricts their access rights, and prompts them to contact the administrator.
[0069] Furthermore, the calculation formula of the Gaussian mixture model is as follows:
[0070]
[0071] Where P(x|λ) is the probability of the feature vector x under the model parameter λ; K is the number of Gaussian components; w k is the weight of the kth Gaussian component; μ k is the mean vector of the kth Gaussian component; ∑ kis the covariance matrix of the kth Gaussian component; d is the dimension of the eigenvector.
[0072] Specifically, the calculation formula of the Gaussian mixture model is used to determine the probability distribution of the feature vector in the model to determine the identity category of the input speech data. The Gaussian mixture model represents a complex probability distribution by a weighted combination of multiple Gaussian components. Each Gaussian component is defined by three main parameters: a mean vector, a covariance matrix, and a weight. Among them, the mean vector represents the central position of a specific component, reflecting the central tendency of the data distribution of the component; the covariance matrix describes the distribution shape and range of the component, characterizing the correlation between the feature vectors in each dimension; the weight determines the importance of each Gaussian component in the overall distribution, and the sum of all weights is equal to 1. For the input feature vector, the model calculates its probability value in the overall distribution in the following way: First, according to the degree of deviation of the feature vector from the center of each Gaussian component, the probability value of the feature vector belonging to the component is calculated. The degree of deviation is measured by the Mahalanobis distance, which comprehensively considers the difference between the feature vector and the mean and the influence of the covariance matrix. Then, the probability value of each component is weighted and summed according to its weight to obtain the overall probability value of the feature vector in the model. This calculation process can effectively map the input feature vector to the probability distribution described by the model. By comparing the probability values of the input data in different models, the degree of matching with each identity category (such as teacher or student) can be determined. For example, if the probability value of the feature vector in the teacher model is higher than that in the student model and exceeds the preset threshold of the teacher model, the system determines that the speech data belongs to the teacher category.
[0073] Furthermore, the electronic whiteboard function is loaded according to the user identity, including:
[0074] When the user is a teacher, teacher functions are loaded, including clearing content, screen sharing, and function locking; when the user is a student, student permission rules are loaded, including the use of annotation tools and answer submission functions; when the user is an unknown person, access is restricted and the administrator is prompted.
[0075] Specifically, when the user identity is determined to be a teacher, the system loads the teacher permission rules. Specifically, the teacher permissions cover advanced functions such as clearing content, screen sharing, and function locking. The system enables the operation interface of these functions through the permission mapping module, and dynamically updates all available functions within the teacher's permission to the user interface. For example, a teacher can call the screen sharing module through the voice command "Start screen sharing", and the system will prompt the operation status through voice or visual feedback after the operation is completed, such as "Screen sharing has been started". In addition, the teacher's permission allows the electronic whiteboard content to be cleared directly, or certain key functions to be locked to prevent misoperation. When the user identity is determined to be a student, the system loads the student permission rules. Student permissions are usually limited to basic functions, such as the use of annotation tools and answer submission functions. This permission rule ensures that students can participate in classroom interaction, but cannot access high-permission functions, thereby avoiding interference with the teaching process. For example, a student can enable the annotation function through the voice command "Use annotation tools", and the system verifies the permissions and calls the corresponding module and displays the operation interface of the annotation tool. If a student tries to access a function that is not within his or her permission range (such as clearing content), the system will reject the request and prompt "Insufficient permissions". When the user's identity is determined to be an unknown person, the system restricts access to all functions for security reasons. At this time, the system will lock the operation permissions of the electronic whiteboard and generate a prompt message, requiring the user to re-verify the identity or contact the administrator. For example, when the system recognizes that the voiceprint of the input voice does not match the teacher or student model, it will display a prompt of "Unauthorized access, please contact the administrator" and record the relevant operation logs for subsequent security review.
[0076] Furthermore, parsing the user instructions from the standardized voice data by natural language processing comprises the following steps:
[0077] The standardized speech data is processed through a speech recognition model to convert the speech signal into text data;
[0078] The text is segmented and vectorized based on the pre-trained natural language model, keywords are extracted through the named entity recognition algorithm, and the semantic relationship in the text is analyzed through the syntactic dependency analysis algorithm to generate a preliminary mapping of the user's operation intention;
[0079] The extracted keywords and feature vectors are input into the Softmax classifier model to determine the specific functional module corresponding to the instruction.
[0080] Specifically, the training of the natural language model starts with the construction and preprocessing of a large-scale corpus. The corpus covers a variety of text data related to teaching and voice instructions, such as common instructions such as "start screen sharing" and "clear all content". After preprocessing, the text data includes word segmentation to decompose sentences into basic language units, remove stop words to reduce unnecessary semantic interference, and generate a vocabulary to map each word to a vector representation of fixed dimensions. For example, the sentence "start screen sharing" is decomposed into three keywords and mapped to a vector form that the model can understand. The model is trained using pre-training technology, using BERT (Bidirectional Encoder Representations from Transformers) as the basic model, focusing on learning the contextual dependencies of language. In the training task, the model grasps the semantic logic within the sentence through the masked language model (MLM) task. For example, after randomly masking "clear" in "clear all content", the model predicts the probability distribution of the masked word based on the context to learn the association between words. In addition, through the next sentence prediction (NSP) task, the model further understands the logical relationship between sentences. For example, given "Start screen sharing", the next sentence may be "Screen sharing has started". Through such tasks, the model learns how to identify the semantic integrity of instruction statements. In the process of model optimization, the parameters are iteratively updated using the AdamW optimization algorithm to minimize the loss function. In the process of model training, a large number of data inputs are used, the loss is calculated in batches (min i-batch), and the model weights are adjusted through back propagation. For example, in the early stage of training, the model may not accurately identify the association between "Start screen sharing" and "Screen sharing has started", but after dozens of iterations, the model's understanding of the semantic relationship is gradually enhanced. After pre-training, the model is fine-tuned on the instruction dataset of the electronic whiteboard operation scenario. The fine-tuning task is a supervised learning process, and the model is further optimized on the labeled corpus. For example, "Use annotation tool" is annotated as a label of "annotation module call". By learning these mapping relationships, the model gradually adapts to the needs of electronic whiteboard function calls. In actual applications, when the user voice data is converted into text, the trained natural language model can parse the semantics of complex instructions. For example, for the instruction "start screen sharing and clear annotations", the model can extract the four key semantic units of "start", "screen sharing", "clear" and "annotation" through word segmentation and vectorization, and judge their priorities through context analysis, and output the instruction execution sequence of "call the screen sharing module, and then clear the annotations". This multi-level parsing capability comes from the model's precise learning of context dependencies during pre-training and fine-tuning. After completing the preliminary semantic parsing, the system inputs the extracted keywords and vectorized features into the Softmax classifier model to determine the specific functional module corresponding to the user's instructions.The Softmax classifier matches the user intent with predefined functional categories (such as screen sharing, annotation tools, etc.) based on the probability distribution of the feature vector. For example, for the "Start screen sharing" instruction, the highest probability output by the classifier corresponds to the screen sharing module, and the system activates the relevant function accordingly. If the instruction does not match any functional category, the system prompts "Unable to recognize, please try again".
[0081] Furthermore, the traversal and execution of matching electronic whiteboard functions through user instructions includes the following steps:
[0082] Compare the keywords in the user's command with the mapping table of the electronic whiteboard function module based on the hash search algorithm;
[0083] A Boolean logic verification algorithm is used to verify whether the user's instructions comply with the permission rules of the current identity; if the verification passes, the operation interface of the corresponding functional module is triggered and the function call is triggered; if the verification fails, the operation is terminated and a prompt of insufficient permissions is returned.
[0084] In some embodiments, first, after receiving the user instruction that has been parsed, the system extracts the core keywords in the instruction. These keywords are usually generated by parsing the natural language model. For example, the user voice instruction "start screen sharing" will extract "start" and "screen sharing" as core keywords. The system then uses a hash search algorithm to quickly compare the extracted keywords with the predefined function module mapping table. The function mapping table is a structured data table that stores all the function modules of the electronic whiteboard and their corresponding keywords. For example, "screen sharing" is mapped to the screen sharing module, and "annotation tool" is mapped to the annotation tool module. By mapping the keywords through the hash algorithm, the corresponding function module can be quickly located in a constant time, avoiding the performance bottleneck caused by traditional linear search. After completing the preliminary matching of the function module, the system verifies the user identity and permissions. Using the Boolean logic verification algorithm, the system checks whether the permission rules of the current user allow the execution of the function. For example, a teacher identity user has the permission to "start screen sharing", while a student identity user does not. The verification rule is based on the predefined permission model and combines the user identity and the function module permission label for logical judgment. For example, for the "Start Screen Sharing" command, the system verification rules will verify whether the current user is a teacher with screen sharing permissions. If the verification passes, the system records the legitimacy of the operation, triggers the operation interface of the corresponding function module and executes the function call; if the verification fails, the system terminates the operation, returns a prompt message of "Insufficient Permission", and informs the user through voice or visual means. In the function call phase, the system interacts with the target module to perform the specific operation requested by the user. For example, when calling the screen sharing module, the system passes the command mapping result to the module interface, starts the screen sharing function, and provides real-time feedback on the operation status, such as prompting "Screen sharing has been started". If an exception occurs during the call (such as the module not responding), the system will capture the exception information and return a prompt of "Operation failed, please try again". Through the above steps, the system can not only efficiently match user instructions and function modules, but also ensure the compliance of operations through permission verification, thereby preventing unauthorized users from executing high-permission functions. For example, if a student user tries to execute the "Clear Content" command, the system will find that the function is not included in its permission rules during verification, and then terminate the operation and prompt "Insufficient Permission". This design not only ensures the safety of teaching equipment, but also improves the efficiency and reliability of user command execution.
[0085] Furthermore, the insufficient authority prompt is given through voice notification and electronic whiteboard screen visualization.
[0086] Specifically, when the system detects that the user attempts to execute an instruction that exceeds the scope of their authority, such as a student user trying to call the "clear content" function, the system will immediately trigger the prompt mechanism of insufficient authority. First, the system sends a real-time prompt to the user through the voice notification function. In the specific implementation, the system uses text-to-speech (TTS) technology to synthesize prompt information such as "Insufficient authority, please contact the administrator" or "This function is only available to teachers" into a voice signal and play it through the built-in speaker. This real-time voice feedback method can quickly attract the user's attention, especially in a dynamic classroom environment, and can effectively remind the user why their operation is restricted. At the same time, the electronic whiteboard screen synchronously displays the visual prompt information of insufficient authority. The system presents the prompt content consistent with the voice notification in the form of a pop-up window or floating text on the screen, for example: "Insufficient authority: The current user does not have the authority to perform the 'clear content' function, please contact the administrator". Visual prompts are usually designed with high-contrast colors and eye-catching icons (such as warning signs) to ensure that the information is clear and easy to read. In addition, the prompt information may contain more detailed instructions, such as "The current function is only accessible to teacher users, please contact the administrator if you need help", so that users can more fully understand the reasons for the restriction. For the interactive design of insufficient permission prompts, the system also provides a response mechanism. For example, when the user clicks the "Learn More" option in the screen prompt window, the system can further guide the user to view the permission rules, or display the administrator's contact information through a pop-up window. This multi-level feedback method not only improves the user experience, but also provides a clear correction path for incorrect operations in teaching scenarios. Through the combination of voice notifications and screen visual prompts, the system can take into account the user's auditory and visual perception channels to ensure that information about insufficient permissions is conveyed to the user in an intuitive and clear manner. This multimodal feedback method effectively reduces the confusion that may arise from the failure of operations due to limited permissions, while improving the user-friendliness and operability of the system. Through this mechanism, the system further consolidates the rigor of permission management and the security of teaching equipment.
[0087] The present invention also includes an electronic whiteboard function separation system based on voice authority management, which is applied to any of the above-mentioned electronic whiteboard function separation methods based on voice authority management, and includes a voice data acquisition module, a feature extraction module, an identity determination module and a function matching module connected in sequence;
[0088] The voice data acquisition module is used to capture user voice data in real time and perform data cleaning and standardization;
[0089] The feature extraction module is used to extract features from the standardized user voice data based on the Mel frequency cepstral coefficients to generate a feature vector;
[0090] The identity determination module is used to determine the user identity through a Gaussian mixture model based on the feature vector;
[0091] The function matching module is used to load the electronic whiteboard function according to the user identity, parse the user instruction from the standardized voice data through natural language processing, and traverse and match the electronic whiteboard function according to the user instruction and execute it.
[0092] The above implementation modes are merely descriptions of the preferred implementation modes of the present invention, and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary engineering and technical personnel in the field shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A method for separating electronic whiteboard functions based on voice authority management, characterized in that: include: Capture user voice data in real time and perform data cleaning and standardization; Extract features from the standardized user voice data based on the Mel-frequency cepstral coefficients to generate a feature vector; Based on the feature vector, determining the user identity by using a Gaussian mixture model; The electronic whiteboard function is loaded according to the user identity, the user instructions are parsed from the standardized voice data through natural language processing, and the electronic whiteboard functions are traversed and matched according to the user instructions and executed.
2. According to claim 1, a method and system for separating electronic whiteboard functions based on voice authority management, characterized in that: The real-time capture of user voice data, data cleaning and standardization comprises the following steps: Convert user voice data into digital signals; The background noise of digital signal is eliminated by adaptive filtering algorithm; The noise-reduced speech signal is divided into frames using a signal framing algorithm, with the frame length set to 25ms and the frame shift set to 10ms. The speech signal after the frame division processing is subjected to windowing processing.
3. The electronic whiteboard function separation method based on voice authority management according to claim 1 is characterized in that: The feature extraction of the standardized user voice data based on the Mel-frequency cepstral coefficients comprises the following steps: Performing spectrum analysis on the standardized speech signal through Fourier transform algorithm to generate the spectrum of the speech signal; The frequency spectrum of the speech signal is subjected to frequency scale mapping processing through a Mel filter bank, and the frequency distribution is mapped to the Mel frequency scale; Applying logarithmic transformation algorithm to the Mel frequency scale for energy compression; The compressed Mel frequency scale data is processed by discrete cosine transform algorithm to extract feature vectors, generate Mel frequency cepstrum coefficients and use them as feature vectors.
4. The electronic whiteboard function separation method based on voice authority management according to claim 1 is characterized in that: Determining the user identity by using a Gaussian mixture model comprises the following steps: The historical speech feature vectors of teachers and students are modeled using a Gaussian mixture model. The mean vector, covariance matrix, and mixture weight of each type of data are calculated using the expectation maximization algorithm to generate a teacher model and a student model respectively. Input the user's speech feature vector into the teacher model and the student model, and use the probability distribution function of the Gaussian mixture model to calculate the posterior probability values of the feature vector in the two models respectively; Compare the posterior probability values of the user's voice feature vector in the teacher model and the student model. When the probability value reaches the preset teacher threshold, it is judged as a teacher; when it reaches the preset student threshold, it is judged as a student; if both are lower than the preset teacher threshold and the preset student threshold, it is judged as an unknown person.
5. The electronic whiteboard function separation method based on voice authority management according to claim 4 is characterized in that: The calculation formula of the Gaussian mixture model is as follows: Where P(x|λ) is the probability of the feature vector x under the model parameter λ; K is the number of Gaussian components; w k is the weight of the kth Gaussian component; μ k is the mean vector of the kth Gaussian component; ∑ k is the covariance matrix of the kth Gaussian component; d is the dimension of the eigenvector.
6. The electronic whiteboard function separation method based on voice authority management according to claim 5 is characterized in that: The electronic whiteboard function loaded according to the user identity includes: When the user is a teacher, teacher functions are loaded, including clearing content, screen sharing, and function locking; when the user is a student, student permission rules are loaded, including the use of annotation tools and answer submission functions; when the user is an unknown person, access is restricted and the administrator is prompted.
7. The electronic whiteboard function separation method based on voice authority management according to claim 1 is characterized in that: The method of parsing user instructions from the standardized voice data by natural language processing comprises the following steps: The standardized speech data is processed through a speech recognition model to convert the speech signal into text data; The text is segmented and vectorized based on the pre-trained natural language model, keywords are extracted through the named entity recognition algorithm, and the semantic relationship in the text is analyzed through the syntactic dependency analysis algorithm to generate a preliminary mapping of the user's operation intention; The extracted keywords and feature vectors are input into the Softmax classifier model to determine the specific functional module corresponding to the instruction.
8. The electronic whiteboard function separation method based on voice authority management according to claim 7, characterized in that: The traversal and matching of electronic whiteboard functions by user instructions and execution include the following steps: Based on the hash search algorithm, the keywords in the user command are compared with the mapping table of the electronic whiteboard function module; A Boolean logic verification algorithm is used to verify whether the user's instructions comply with the permission rules of the current identity; if the verification passes, the operation interface of the corresponding functional module is triggered and the function call is triggered; if the verification fails, the operation is terminated and a prompt of insufficient permissions is returned.
9. The method for separating electronic whiteboard functions based on voice authority management according to claim 8, characterized in that: The insufficient authority prompt is given through voice notification and electronic whiteboard screen visualization.
10. An electronic whiteboard function separation system based on voice authority management, applied to an electronic whiteboard function separation method based on voice authority management as claimed in any one of claims 1 to 9, characterized in that: It includes a voice data acquisition module, a feature extraction module, an identity determination module and a function matching module connected in sequence; The voice data acquisition module is used to capture user voice data in real time and perform data cleaning and standardization; The feature extraction module is used to extract features from the standardized user voice data based on the Mel frequency cepstral coefficients to generate a feature vector; The identity determination module is used to determine the user identity through a Gaussian mixture model based on the feature vector; The function matching module is used to load the electronic whiteboard function according to the user identity, parse the user instruction from the standardized voice data through natural language processing, and traverse and match the electronic whiteboard function according to the user instruction and execute it.