Voice control method and device for conference
By performing deep noise reduction and voiceprint recognition in the conference system, the problem of environmental noise interference in traditional voice control methods is solved, efficient and accurate voice control is achieved, and the intelligence level of conference management is improved.
Patent Information
- Application Number
- CN202510578459.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The traditional voice control method of conference system has environmental noise that cannot be filtered out, resulting in the inability to accurately recognize voice commands, and there is a certain delay, resulting in low control efficiency.
In pure signaling mode, deep noise reduction is performed through the voice acquisition module, voiceprint features are extracted and pre-trained voiceprint recognition model is used for identity matching, and voice signals are parsed to perform tasks.
It improves the accuracy and efficiency of voice control, realizes intelligent management, reduces the misidentification rate, and improves the overall efficiency of the meeting.
Smart Images

Figure CN120108402B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent conference control, and in particular to a method and device for voice control of a conference. Background Art
[0002] With the rapid development of information technology, intelligent conference systems have become a key trend in improving meeting efficiency and optimizing the meeting experience. Traditional conference systems often rely on manual operations, such as manually controlling conference equipment and manually recording meeting content. These methods are not only inefficient but also prone to errors.
[0003] As one of the core technologies for intelligent conference systems, voice signal processing has made significant progress in recent years, providing strong technical support for the intelligentization of conference systems. Traditional voice control methods suffer from the inability to filter out ambient noise, resulting in inaccurate recognition of voice commands, and a certain delay, leading to low control efficiency. Summary of the Invention
[0004] The present invention provides a method and device for voice control of a conference, so as to solve the problems of low voice control efficiency and low recognition accuracy.
[0005] According to one aspect of the present invention, a method for voice control of a conference is provided, the method comprising:
[0006] When the target working mode is the pure signaling mode, the initial voice signal input by the participant is collected by the voice acquisition module, and the initial voice signal is subjected to deep noise reduction processing to obtain the target voice signal;
[0007] The voiceprint features in the target voice signal are extracted through the voiceprint recognition module, the voiceprint features are input into a pre-trained voiceprint recognition model, and the voiceprint recognition result is determined based on the model output result; when the voiceprint recognition result shows that there is a matching participant identity, the target voice signal is parsed through the signaling processing module to obtain the target task execution instruction, and the target task is executed based on the target task execution instruction.
[0008] According to another aspect of the present invention, a voice control method and apparatus is provided, the apparatus comprising:
[0009] A voice signal acquisition module is configured to, when the target working mode is the pure signaling mode, collect the initial voice signal input by the participant through the voice acquisition module, and perform deep noise reduction processing on the initial voice signal to obtain a target voice signal;
[0010] The voiceprint recognition module is used to extract the voiceprint features in the target voice signal through the voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine the voiceprint recognition result based on the model output result; the task execution module is used to parse the target voice signal through the signaling processing module when the voiceprint recognition result shows that there is a matching participant identity, so as to obtain a target task execution instruction, and execute the target task based on the target task execution instruction.
[0011] According to another aspect of the present invention, an electronic device is provided, comprising:
[0012] at least one processor; and
[0013] a memory communicatively connected to the at least one processor; wherein,
[0014] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the voice control method for a conference described in any embodiment of the present invention.
[0015] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the voice control method for a conference described in any embodiment of the present invention when executed.
[0016] The technical solution of the embodiment of the present invention involves, when the target operating mode is pure signaling mode, acquiring an initial voice signal input by a participant through a voice acquisition module, and performing deep noise reduction processing on the initial voice signal to obtain a target voice signal. The noise-reduced target voice signal is clearer, providing a high-quality data foundation for subsequent voiceprint recognition. The voiceprint recognition module then extracts voiceprint features from the target voice signal, inputs these features into a pre-trained voiceprint recognition model, and determines the voiceprint recognition result based on the model output. This more comprehensively captures the unique characteristics of the voice signal. Based on high-quality data, the pre-trained voiceprint recognition model can more accurately match participant identities, reducing the false recognition rate. Finally, if the voiceprint recognition result indicates a matching participant identity, the signaling processing module analyzes the target voice signal to obtain a target task execution instruction. The target task is then executed based on the target task execution instruction. This solves the problems of low voice control efficiency and low recognition accuracy, thereby improving meeting efficiency and achieving the beneficial effects of intelligent management.
[0017] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 This is a flow chart of a method for voice control of a conference provided in Embodiment 1 of the present invention;
[0020] Figure 2 This is a flow chart of a method for voice control of a conference provided according to a second embodiment of the present invention;
[0021] Figure 3 This is a structural diagram of a voice control device for a conference provided in Embodiment 3 of the present invention;
[0022] Figure 4 The present invention is a schematic diagram of the structure of an electronic device for implementing the voice control method for a conference according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] Example 1
[0026] Figure 1 A flowchart of a conference voice control method is provided for the first embodiment of the present invention. This embodiment is applicable to conference voice control situations. The method can be executed by a conference voice control device. The conference voice control device can be implemented in the form of hardware and / or software. The conference voice control device can be configured in an electronic device. Figure 1 As shown, the method includes:
[0027] S110 : When the target working mode is the pure signaling mode, the initial voice signal input by the participant is collected by acquiring the voice collection module, and the initial voice signal is subjected to deep noise reduction processing to obtain a target voice signal.
[0028] The initial voice signal can be understood as pulse code modulation signal data. The target voice signal can be understood as the voice signal data after noise removal. The pure signaling mode can be understood as an operating mode that focuses on the precise transmission of various control commands to ensure stable system operation and flexible control.
[0029] Specifically, in pure signaling mode, the voice acquisition module first acquires PCM (pulse code modulation) data, which carries critical voice information. To improve voice quality, this data is subjected to deep noise reduction processing by introducing automatic gain control, automatic echo cancellation, and automatic noise suppression. This effectively filters out interfering factors such as ambient noise and echo, allowing subsequent processing to be based on a clearer and purer voice signal.
[0030] S120. Extract voiceprint features from the target speech signal through a voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine a voiceprint recognition result based on the model output result.
[0031] Voiceprint features can be understood as a set of speech features that characterize and identify the speaker. Voiceprint recognition results include speaker identification results, including at least one of the speaker's identity, a matching score with each voiceprint template in the database, and a matching status (success or failure).
[0032] Specifically, the target speech signal, after noise reduction processing, is accurately fed into the voiceprint recognition module. Within this module, two advanced feature extraction techniques, Mel-frequency cepstral coefficients and linear prediction cepstral coefficients, are used to extract voiceprint features from the processed speech signal. The extracted voiceprint features are then fed into a pre-trained voiceprint recognition model for matching operations, effectively and accurately completing the voiceprint recognition task and providing a solid guarantee for the system's security and accuracy.
[0033] Optionally, the voiceprint features in the target voice signal are extracted through a voiceprint recognition module, including: preprocessing the target voice signal to obtain a target digital signal, and extracting the Mel-frequency cepstral coefficients and linear prediction cepstral coefficients of the target digital signal; the preprocessing includes at least two of signal conversion processing, enhancement processing, windowing processing and frame processing; and determining the voiceprint features in the target voice signal based on a feature fusion method, a pattern recognition algorithm, the Mel-frequency cepstral coefficients and the linear prediction cepstral coefficients.
[0034] Specifically, signal conversion converts the original speech signal into a digital signal suitable for subsequent processing. The analog speech signal is sampled and quantized using an analog-to-digital converter (ADC) to generate a digital signal. Enhancement uses noise reduction algorithms (such as spectral subtraction, Wiener filtering, and deep noise reduction) to reduce noise in the speech signal. Windowing multiplies the speech signal by a function such as a Hamming window to smooth the signal in the time domain. Framing divides the speech signal into frames using a fixed frame length and frame shift.
[0035] Specifically, voiceprint features that can characterize the speaker's identity are extracted from the target digital signal. Common voiceprint features include Mel-frequency cepstral coefficients and linear prediction cepstral coefficients. These features can be fused using serial, parallel, or weighted fusion methods. The fused voiceprint features are then classified and identified to determine the speaker's identity. Pattern recognition algorithms can include Gaussian mixture models, deep neural networks, and support vector machines.
[0036] Optionally, extracting Mel-frequency cepstral coefficients of the target digital signal includes:
[0037] Performing spectral analysis on the target digital signal to obtain a power spectrum corresponding to the target digital signal, and determining a spectrum based on the power spectrum; performing Mel filtering on the spectrum, performing discrete cosine transform on the filtered spectrum to obtain the Mel-frequency cepstral coefficients, and extracting the Mel-frequency cepstral coefficients.
[0038] Specifically, a fast Fourier transform is performed on the target digital signal to convert the time domain signal into a frequency domain signal. The complex spectrum of the signal is obtained, which contains amplitude and phase information. The amplitude of the complex spectrum is squared to obtain the power spectrum of the signal. The power spectrum reflects the energy distribution of the signal at different frequencies. The logarithm of the power spectrum is taken to compress the dynamic range and enhance the visibility of low-amplitude components. A set of bandpass filters is designed on the Mel frequency scale, with each filter covering a certain frequency range. The power spectrum or logarithmic power spectrum is passed through the Mel filter bank to obtain the energy at the output of each filter, and then a Mel frequency energy spectrum is obtained, where each element corresponds to the output energy of a Mel filter. The Mel frequency energy spectrum is converted into cepstral coefficients using a discrete cosine transform to remove the correlation between filters and obtain a more compact feature representation.
[0039] For example, a first-order finite difference equation is used , enhance the high frequency part, make the signal spectrum flatter, compensate for the natural attenuation of the target digital signal in the high frequency part, highlight the high frequency information in the speech signal, which is conducive to subsequent feature extraction; use function , perform frame processing to decompose the target digital signal into a series of short-term stable signals, which is convenient for subsequent feature extraction on each frame; use the Hamming window function Perform windowing on each frame of signal to reduce spectrum leakage caused by frame truncation and make the energy of each frame of signal in the frequency domain more concentrated; perform fast Fourier transform on each frame of signal after windowing to obtain the spectrum , convert the time domain signal into the frequency domain signal to obtain the spectrum information of the speech signal; construct a set of Mel filters with a center frequency The conversion relationship between (unit: Hz) and Mel frequency m (unit: Mel) is For example, suppose P Mel filters are constructed, and for each filter (p is between 20-40), calculate the spectrum after Mel filtering , simulating the human ear's frequency perception characteristics. Taking the logarithm gives , converting the multiplication operation of the spectrum into addition operation, compressing the dynamic range, reducing the error of numerical calculation, and simulating the human ear's perception of sound intensity to a certain extent; taking the logarithm of the Mel spectrum Perform DCT to obtain Mel frequency cepstral coefficients (The value of C is between 12-16).
[0040] Optionally, extracting the linear prediction cepstral coefficients of the target digital signal includes: performing linear prediction analysis on the target digital signal to obtain linear prediction coefficients corresponding to the target digital signal; converting the linear prediction coefficients into linear prediction cepstral coefficients, and extracting the linear prediction cepstral coefficients.
[0041] Specifically, the target digital signal is segmented into short time frames, each typically containing a preset number of milliseconds of speech data. The autocorrelation function of each frame is calculated to estimate the signal's periodicity. The autocorrelation coefficients are used to solve the linear prediction equation. These coefficients are then converted to the cepstral domain using a recursive formula, resulting in the linear prediction cepstral coefficients. These coefficients offer improved robustness and discriminability in speech and voiceprint recognition.
[0042] For example, the function , perform frame processing to decompose the speech signal into a series of short-term stable signals, which is convenient for subsequent feature extraction on each frame;
[0043] Use the Hamming window function Perform windowing on each frame of signal. With window function Multiply to get the windowed frame ; Using the autocorrelation function (p is the linear prediction order, ranging from 10-16); linear prediction equation ,in is the target digital signal is the linear prediction coefficient, p is the prediction order, Is the prediction error. When solving by the autocorrelation method, the autocorrelation matrix R and vector r are constructed according to the calculated autocorrelation function, and then the equation is solved Get the linear prediction coefficient According to the linear prediction coefficient , the prediction error energy can be calculated , and thus gain ; According to the linear prediction coefficient Calculate the linear prediction cepstral coefficients. By recursive formula , (C is the order of the linear prediction cepstral coefficient, ranging from 12 to 16).
[0044] Optionally, before inputting the voiceprint features into a pre-trained voiceprint recognition model and determining the voiceprint recognition results based on the model output results, the method further includes: constructing a sample set, and dividing the sample set into a training set and a test set based on a preset ratio; wherein the sample set includes a preset number of sample voiceprint features and identity labels corresponding to the sample voiceprint features; performing iterative training on a pre-established deep neural network for a preset number of times based on the training set, and in each iteration, calculating the loss based on the probability prediction value output by the model and the waiting identity label, and adjusting the weight of the network through the back propagation algorithm; and testing each trained model based on the test set to obtain multiple model performance indicators, and determining the voiceprint recognition model based on the multiple model performance indicators.
[0045] Specifically, according to the task requirements and data characteristics, choose a suitable network architecture. Determine parameters such as the number of neurons in each layer, activation function, convolution kernel size and stride (for CNN), and the number of recurrent units (for RNN and its variants). In the input layer, the number of neurons should match the dimension of the input voiceprint feature vector. The number of neurons in the hidden layer can be adjusted through experiments, generally between dozens and hundreds. The activation function can be selected from ReLU, tanh, etc. The output layer uses the softmax function for multi-classification (assuming that multiple speakers are recognized) according to the recognition task, and outputs the probability distribution of each speaker. Collect a large number of speech samples from different speakers and accurately label them to ensure that each sample corresponds to the correct speaker identity label. The number of samples for each speaker should be as balanced as possible, and cover different speech content, speaking speed, intonation, etc., to improve the generalization ability of the network. For each speech sample, extract the voiceprint feature vector. Then, normalize the feature vector and scale its numerical range to an appropriate interval. Use normal distribution or evenly distributed To initialize the weight matrix The bias term is usually initialized to 0 or a small constant such as 0.1.
[0046] Specifically, the preprocessed voiceprint feature vector is used as the network input and passed through each layer of the network in sequence. In the convolutional layer, the convolution kernel is convolved with the input feature map to extract local features, and then the output feature map is processed through the activation function. In the recurrent layer, based on the sequential characteristics of the speech signal, the features of each time step are processed in sequence, the hidden state is updated, and the final hidden state is output to the next layer. The fully connected layer linearly transforms the output of the previous layer and processes the activation function, gradually mapping the features to a higher level of abstract representation, and finally obtaining the probability prediction value of each speaker at the output layer.
[0047] Specifically, use the cross entropy loss function formula: , substitute the network's output prediction value and the corresponding true label into the loss function to calculate the loss value of the current batch of samples. Based on the calculated loss value, the back propagation algorithm is used to calculate the gradient of each weight parameter in the network. Using the calculated gradient value, the adaptive moment estimation optimization algorithm is used to update the network's weight parameters. In each iteration, the weight update formula is ,By repeating the process of forward propagation, calculating the loss function, ,backward propagation, and weight updating, the network gradually learns the model ,parameters that can accurately identify the voice prints of different speakers.
[0048] Specifically, the forward propagation, loss calculation, backpropagation, and weight update processes described above are encapsulated in a training loop. The model is trained iteratively on the entire training set until the network's performance (such as accuracy, loss, and other metrics) no longer improves or reaches a preset training stop condition. Within each iteration, the training data is typically divided into multiple small batches. During training, several hyperparameters need to be adjusted to optimize network performance. These hyperparameters include the learning rate, batch size, various parameters of the network architecture (such as the number of layers and neurons), and the number of training iterations.
[0049] Specifically, during training, the model is regularly evaluated using the test set to monitor performance and avoid overfitting. By observing how model performance metrics on the test set change with each training iteration, we can determine whether the model is overfitting (if metrics on the test set stop improving or even decrease, while metrics on the training set continue to rise, overfitting may occur), and take timely measures, such as stopping training early or adding regularization terms. Selecting the optimal performing model through test set evaluation can significantly improve the accuracy and robustness of voiceprint recognition.
[0050] Optionally, the voiceprint recognition model includes a convolutional layer, a recurrent layer connected to the convolutional layer, a fully connected layer connected to the recurrent layer, and an output layer connected to the fully connected layer; the step of inputting the voiceprint feature into a pre-trained voiceprint recognition model and determining the voiceprint recognition result based on the model output result comprises: inputting the voiceprint feature into a pre-trained voiceprint recognition model, extracting local features in the voiceprint feature through the convolutional layer, and generating a feature map; determining a feature sequence corresponding to the feature map, processing the features of each time step in turn according to the sequence characteristics of the feature sequence through the recurrent layer, updating the hidden state, and outputting the updated hidden state to the fully connected layer; performing linear transformation and activation function processing on the updated hidden state through the fully connected layer, mapping the features of the updated hidden state to an identity prediction value for matching the speaker identity corresponding to the voiceprint feature; outputting the identity prediction value through the output layer, and determining the voiceprint recognition result based on the identity prediction value.
[0051] Specifically, the extracted voiceprint features are input into the model. The convolutional layer extracts local patterns in the voiceprint features and generates a feature map. The feature map is converted into a feature sequence. The recurrent layer processes the features at each time step in turn, updating the hidden state. The fully connected layer applies a linear transformation and an activation function to the hidden state to obtain an identity prediction value. The output layer outputs the identity prediction value and selects the speaker with the highest probability as the voiceprint recognition result.
[0052] S130. When the voiceprint recognition result indicates that there is a matching participant identity, the target voice signal is parsed by the signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction.
[0053] Among them, the target task execution instruction can be understood as a specific operation command generated by input (such as voice instructions, text instructions, etc.).
[0054] Specifically, the system first confirms the voiceprint recognition result, confirming that the target voice signal successfully matches the voiceprint characteristics of a participant in the database. Once a matching participant's identity is confirmed, the system automatically activates the signaling processing module. The target voice signal is input into the signaling processing module for analysis. Using natural language processing technology, the system identifies keywords and phrases in the voice signal. These keywords are often related to the target task execution instructions. Semantic understanding of these keywords and phrases is performed to determine their meaning and intent. Based on this semantic understanding, specific target task execution instructions are generated.
[0055] Optionally, the method further includes:
[0056] When the target working mode is pure voice mode, the initial voice signal input by the participant is collected by the voice acquisition module, and the initial voice signal is converted into voice text data; the voice text data is parsed by the signaling processing module to obtain the target task execution instruction, and the target task is executed based on the target task execution instruction.
[0057] Among them, the pure voice mode can be understood as a mode that aims to achieve efficient transmission of normal conference sounds and ensure that participants can receive various voice information clearly and smoothly.
[0058] Specifically, in pure voice mode, the collected initial voice signal undergoes a series of voice preprocessing operations, including the use of low-pass, high-pass, and bandpass filtering techniques to remove unwanted frequency components. The signal is then amplified to optimize its quality and strength, laying a solid foundation for subsequent processing. The speech recognition module then converts the preprocessed initial voice signal into speech text data. The signaling processing module then conducts in-depth analysis of this text information, extracting key intent and instructions to drive the corresponding task execution process.
[0059] Optionally, the audio output module efficiently receives additional result data from the task execution process and converts it into analog audio signals using a high-precision digital-to-analog converter (DAC). Ultimately, the converted analog audio signals are perfectly presented through carefully adapted third-party audio playback devices, such as high-quality speakers or professional headphones, creating a smooth, accurate, and immersive pure voice interaction experience for users, fully meeting their diverse needs and expectations in pure voice mode.
[0060] For example, when there is important information that needs to be announced or the speech of a participant needs to be heard by other participants in a meeting, the system converts the text information into a natural and fluent voice signal through speech synthesis technology.
[0061] A high-precision digital-to-analog converter converts digital voice signals into analog audio signals, which are then appropriately amplified. Finally, the processed audio signals are played back through well-adapted third-party audio playback equipment (such as the conference room's speaker system), enabling efficient and accurate interaction between signaling and users.
[0062] Optionally, the method further includes: when the target working mode is a pure signaling mode,
[0063] The light control module controls the voice control indicator light to remain in a constantly on state; when the target working mode is the voice-only mode, the indicator light control module controls the voice control indicator light to remain in an off state.
[0064] Specifically, in pure signaling mode, the voice control indicator light should remain on to clearly indicate the current working status of the system. In this way, users or operators can intuitively see that the system is in pure signaling mode.
[0065] In an embodiment of the present invention, a "voice control" indicator light is added so that operators can more intuitively distinguish between pure voice mode and pure signaling mode. In pure voice mode, the voice control indicator light should remain off. For example, assume there is an intelligent conference system that supports two working modes: pure signaling mode and pure voice mode. When the system switches to pure signaling mode, the voice control indicator light will light up, indicating that the system is receiving and processing instructions through signals; and when the system switches to pure voice mode, the indicator light will go out, indicating that the system is receiving and processing instructions through voice. In this way, users can intuitively understand the current working mode of the system through the status of the indicator light.
[0066] The technical solution of the embodiment of the present invention involves, when the target operating mode is pure signaling mode, acquiring an initial voice signal input by a participant through a voice acquisition module, and performing deep noise reduction processing on the initial voice signal to obtain a target voice signal. The noise-reduced target voice signal is clearer, providing a high-quality data foundation for subsequent voiceprint recognition. The voiceprint recognition module then extracts voiceprint features from the target voice signal, inputs these features into a pre-trained voiceprint recognition model, and determines the voiceprint recognition result based on the model output. This more comprehensively captures the unique characteristics of the voice signal. Based on high-quality data, the pre-trained voiceprint recognition model can more accurately match participant identities, reducing the false recognition rate. Finally, if the voiceprint recognition result indicates a matching participant identity, the signaling processing module analyzes the target voice signal to obtain a target task execution instruction. The target task is then executed based on the target task execution instruction. This solves the problems of low voice control efficiency and low recognition accuracy, thereby improving meeting efficiency and achieving the beneficial effects of intelligent management.
[0067] Example 2
[0068] Figure 2 This is a flow chart of a method for voice control of a conference provided in the second embodiment of the present invention. This embodiment is a further optimization of the above embodiment. Figure 2 As shown, the method includes:
[0069] S210 : Detecting an input level value, and if the level value is not greater than a preset first level value threshold, determining that the target operating mode is a pure signaling mode.
[0070] The preset first level value threshold may be understood as a preset low level threshold, which may be preset based on experience and is not limited in this embodiment.
[0071] Specifically, if the input level is low, the system will quickly select "Signaling Only Mode." During system operation, the input level is accurately monitored in real time. Once a change in the input level is detected, the system immediately enters the mode detection state. If the input level is low, the system will quickly select "Signaling Only Mode," and the "Voice Control" indicator will remain on, providing the user with clear mode indication.
[0072] S220: When the level value is greater than a preset second level value threshold, determine that the target working mode is a pure voice mode.
[0073] The preset second level value threshold may be understood as a preset high level threshold, which may be preset based on experience and is not limited in this embodiment.
[0074] Specifically, the system accurately monitors the input level in real time during operation. Once a change in the input level is detected, it immediately enters mode detection. When the input level reaches a high level, the system automatically switches to "Voice-Only Mode." The "Voice Control" indicator remains off, intuitively indicating the current operating mode. If no level change is detected during the detection process, the system automatically enters a loop detection process, continuously monitoring the input level until a change is detected. This ensures that the system can respond promptly and accurately to different operating mode requirements, providing users with a stable and efficient user experience.
[0075] S230: When the level value is greater than a preset second level value threshold, determine that the target working mode is a pure voice mode.
[0076] S240. Extract voiceprint features from the target speech signal through a voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine a voiceprint recognition result based on the model output result.
[0077] S250: When the voiceprint recognition result indicates that there is a matching participant identity, the target voice signal is parsed by the signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction.
[0078] Optionally, a shielding cover can be added to the device's motherboard, with a switch designed to be adjusted up and down. When the switch is flipped down, the device automatically switches to pure signaling mode. In this mode, all information exchanged between the device and external devices will be stably and efficiently transmitted through the serial port to ensure the fast and accurate transmission of signaling data such as control commands, meeting the needs of application scenarios such as remote control and system configuration that require high accuracy and timeliness. When the switch is flipped up, the device enters pure voice mode. Similarly, all types of voice information exchanged between the device and external devices will also be transmitted through the serial port, ensuring smooth transmission of voice data and providing a high-quality audio transmission channel for applications such as voice calls and voice broadcasts, thereby achieving stable and reliable information exchange between the device and external devices in different working modes.
[0079] The technical solution of the embodiment of the present invention detects the input level value. If the level value is not greater than a preset first level threshold, the target operating mode is determined to be the pure signaling mode. If the level value is greater than a preset second level threshold, the target operating mode is determined to be the pure voice mode. This ensures that the system can respond to different operating mode requirements in a timely and accurate manner, providing users with a stable and efficient user experience.
[0080] Example 3
[0081] Figure 3 This is a structural diagram of a voice control device for a conference provided by the third embodiment of the present invention. Figure 3 As shown, the device includes: a voice signal acquisition module 310, a voiceprint recognition module 320 and a task execution module 330.
[0082] Among them, the voice signal acquisition module 310 is used to collect the initial voice signal input by the participant through the voice acquisition module when the target working mode is the pure signaling mode, and perform deep noise reduction processing on the initial voice signal to obtain the target voice signal; the voiceprint recognition module 320 is used to extract the voiceprint features in the target voice signal through the voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine the voiceprint recognition result based on the model output result; the task execution module 330 is used to parse the target voice signal through the signaling processing module when the voiceprint recognition result shows that there is a matching participant identity, so as to obtain the target task execution instruction, and execute the target task based on the target task execution instruction.
[0083] The technical solution of the embodiment of the present invention involves, when the target operating mode is pure signaling mode, acquiring an initial voice signal input by a participant through a voice acquisition module, and performing deep noise reduction processing on the initial voice signal to obtain a target voice signal. The noise-reduced target voice signal is clearer, providing a high-quality data foundation for subsequent voiceprint recognition. The voiceprint recognition module then extracts voiceprint features from the target voice signal, inputs these features into a pre-trained voiceprint recognition model, and determines the voiceprint recognition result based on the model output. This more comprehensively captures the unique characteristics of the voice signal. Based on high-quality data, the pre-trained voiceprint recognition model can more accurately match participant identities, reducing the false recognition rate. Finally, if the voiceprint recognition result indicates a matching participant identity, the signaling processing module analyzes the target voice signal to obtain a target task execution instruction. The target task is then executed based on the target task execution instruction. This solves the problems of low voice control efficiency and low recognition accuracy, thereby improving meeting efficiency and achieving the beneficial effects of intelligent management.
[0084] Optionally, the voiceprint recognition module includes:
[0085] a feature extraction unit, configured to preprocess the target speech signal to obtain a target digital signal, and extract Mel-frequency cepstral coefficients and linear prediction cepstral coefficients of the target digital signal; wherein the preprocessing includes at least two of signal conversion processing, enhancement processing, windowing processing, and framing processing;
[0086] A voiceprint feature extraction unit is used to determine the voiceprint features in the target speech signal based on a feature fusion method, a pattern recognition algorithm, the Mel-frequency cepstral coefficients and the linear prediction cepstral coefficients.
[0087] Optionally, the feature extraction unit includes:
[0088] a spectrum analysis subunit, configured to perform spectrum analysis on the target digital signal to obtain a power spectrum corresponding to the target digital signal, and determine a spectrum based on the power spectrum;
[0089] The mel filter subunit is configured to perform mel filtering on the spectrum, perform discrete cosine transform based on the filtered spectrum to obtain the mel-frequency cepstral coefficients, and extract the mel-frequency cepstral coefficients.
[0090] Optionally, the feature extraction unit includes:
[0091] a linear prediction subunit, configured to perform a linear prediction analysis on the target digital signal to obtain a linear prediction coefficient corresponding to the target digital signal;
[0092] The linear conversion subunit is used to convert the linear prediction coefficients into linear prediction cepstral coefficients and extract the linear prediction cepstral coefficients.
[0093] Optionally, the device further includes:
[0094] a sample set construction module, configured to construct a sample set before inputting the voiceprint features into a pre-trained voiceprint recognition model and determining the voiceprint recognition result based on the model output, and to divide the sample set into a training set and a test set based on a preset ratio; wherein the sample set includes a preset number of sample voiceprint features and identity tags corresponding to the sample voiceprint features;
[0095] A model training module is used to perform a preset number of iterative training on a pre-established deep neural network based on the training set. In each iteration, the loss is calculated based on the probability prediction value output by the model and the waiting identity label, and the weight of the network is adjusted through the back propagation algorithm;
[0096] The model determination module is used to test each trained model based on the test set to obtain multiple model performance indicators, and determine the voiceprint recognition model based on the multiple model performance indicators.
[0097] Optionally, the voiceprint recognition model includes a convolutional layer, a recurrent layer connected to the convolutional layer, a fully connected layer connected to the recurrent layer, and an output layer connected to the fully connected layer; accordingly, the voiceprint recognition module includes:
[0098] A convolution unit, configured to input the voiceprint features into a pre-trained voiceprint recognition model, extract local features from the voiceprint features through the convolution layer, and generate a feature map;
[0099] a hidden update unit, configured to determine a feature sequence corresponding to the feature map, process the features of each time step in sequence according to the sequence characteristics of the feature sequence through the recurrent layer, update the hidden state, and output the updated hidden state to the fully connected layer;
[0100] a prediction value determination unit, configured to perform linear transformation and activation function processing on the updated hidden state through the fully connected layer, and map the features of the updated hidden state to an identity prediction value for matching the speaker identity corresponding to the voiceprint feature;
[0101] The result output unit is used to output the identity prediction value through the output layer, and determine the voiceprint recognition result based on the identity prediction value.
[0102] Optionally, the device further includes:
[0103] a first mode determination module configured to detect a level value of an initial voice signal input by a participant before the voice acquisition module acquires the initial voice signal, and determine that the target operating mode is the pure signaling mode if the level value is not greater than a preset first level value threshold;
[0104] The second mode determination module is configured to determine that the target working mode is a pure voice mode when the level value is greater than a preset second level value threshold.
[0105] Optionally, the device further includes:
[0106] A text conversion module is used to, when the target working mode is the pure voice mode, collect the initial voice signal input by the participant by acquiring the voice collection module, and convert the initial voice signal into voice text data;
[0107] The text parsing module is used to parse the voice text data through the signaling processing module to obtain a target task execution instruction, and execute the target task based on the target task execution instruction.
[0108] Optionally, the device further includes:
[0109] The first indicator light control module is used to indicate the
[0110] The light control module controls the voice control indicator light to remain on;
[0111] The second indicator light control module is used to indicate when the target working mode is pure voice mode.
[0112] The light control module controls the voice control indicator light to remain off.
[0113] The voice control device for a conference provided by an embodiment of the present invention can execute the voice control method for a conference provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0114] Example 4
[0115] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0116] like Figure 4 As shown, electronic device 10 includes at least one processor 11 and memory, such as read-only memory (ROM) 12 and random access memory (RAM) 13, communicatively connected to at least one processor 11. The memory stores computer programs executable by the at least one processor. Processor 11 can perform various appropriate actions and processes based on the computer programs stored in ROM 12 or loaded from storage unit 18 into RAM 13. RAM 13 can also store various programs and data required for the operation of electronic device 10. Processor 11, ROM 12, and RAM 13 are interconnected via bus 14. An input / output (I / O) interface 15 is also connected to bus 14.
[0117] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0118] Processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any other suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the voice control method for a conference call.
[0119] In some embodiments, the method for voice control of a conference call can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for voice control of a conference call described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method for voice control of a conference call in any other suitable manner (e.g., via firmware).
[0120] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0121] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0122] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or apparatus. A computer-readable storage medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0123] To provide interaction with a service acquirer, the systems and techniques described herein can be implemented on an electronic device that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the service acquirer; and a keyboard and pointing device (e.g., a mouse or trackball), through which the service acquirer can provide input to the electronic device. Other types of devices can also be used to provide interaction with the service acquirer; for example, feedback provided to the service acquirer can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the service acquirer can be received in any form (including acoustic input, voice input, or tactile input).
[0124] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a service acquirer computer having a graphical service acquirer interface or a web browser through which a service acquirer can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0125] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0126] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0127] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for voice control of a conference, characterized in that: include: When the target working mode is pure signaling mode, the initial voice signal input by the participant is collected by acquiring the voice acquisition module, and the initial voice signal is subjected to deep noise reduction processing to obtain the target voice signal; wherein, the pure signaling mode is a working mode for transmitting various control instructions; Extracting voiceprint features from the target speech signal through a voiceprint recognition module, inputting the voiceprint features into a pre-trained voiceprint recognition model, and determining a voiceprint recognition result based on the model output; If the voiceprint recognition result indicates that there is a matching participant identity, the target voice signal is parsed by the signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction; When the target working mode is pure signaling mode, the voice control indicator light is controlled by the indicator light control module to remain on; When the target working mode is the pure voice mode, the voice control indicator light is controlled by the indicator light control module to remain off; wherein the pure voice mode is a working mode for receiving various types of voice information; The method further includes: controlling the device to switch to the target working mode through a switch on a shielding cover of a mainboard of the device.
2. The method according to claim 1, characterized in that Extracting voiceprint features from the target voice signal through a voiceprint recognition module includes: Preprocessing the target speech signal to obtain a target digital signal, and extracting Mel-frequency cepstral coefficients and linear prediction cepstral coefficients of the target digital signal; the preprocessing includes at least two of signal conversion processing, enhancement processing, windowing processing, and framing processing; The voiceprint features in the target speech signal are determined based on a feature fusion method, a pattern recognition algorithm, the Mel-frequency cepstral coefficients and the linear prediction cepstral coefficients.
3. The method according to claim 2, characterized in that The extracting of Mel-frequency cepstral coefficients of the target digital signal includes: performing spectrum analysis on the target digital signal to obtain a power spectrum corresponding to the target digital signal, and determining a frequency spectrum based on the power spectrum; Mel filtering is performed on the spectrum, discrete cosine transform is performed based on the filtered spectrum to obtain the Mel-frequency cepstral coefficients, and the Mel-frequency cepstral coefficients are extracted.
4. The method according to claim 2, characterized in that The extracting of linear prediction cepstral coefficients of the target digital signal includes: Performing a linear prediction analysis on the target digital signal to obtain a linear prediction coefficient corresponding to the target digital signal; The linear prediction coefficients are converted into linear prediction cepstral coefficients, and the linear prediction cepstral coefficients are extracted.
5. The method according to claim 1, characterized in that Before inputting the voiceprint feature into a pre-trained voiceprint recognition model and determining the voiceprint recognition result based on the model output result, the method further includes: Constructing a sample set, and dividing the sample set into a training set and a test set based on a preset ratio; wherein the sample set includes a preset number of sample voiceprint features and identity labels corresponding to the sample voiceprint features; Performing a preset number of iterative training on a pre-established deep neural network based on the training set, wherein in each iteration, the loss is calculated based on the probability prediction value output by the model and the waiting identity label, and the weight of the network is adjusted through a back-propagation algorithm; Each trained model is tested based on the test set to obtain multiple model performance indicators, and the voiceprint recognition model is determined based on the multiple model performance indicators.
6. The method according to claim 1, characterized in that The voiceprint recognition model includes a convolutional layer, a recurrent layer connected to the convolutional layer, a fully connected layer connected to the recurrent layer, and an output layer connected to the fully connected layer; The step of inputting the voiceprint feature into a pre-trained voiceprint recognition model and determining the voiceprint recognition result based on the model output result includes: Inputting the voiceprint features into a pre-trained voiceprint recognition model, extracting local features from the voiceprint features through the convolutional layer, and generating a feature map; Determine a feature sequence corresponding to the feature map, process the features of each time step in sequence according to the sequence characteristics of the feature sequence through the recurrent layer, update the hidden state, and output the updated hidden state to the fully connected layer; Performing linear transformation and activation function processing on the updated hidden state through the fully connected layer, mapping the features of the updated hidden state to the identity prediction value of the speaker identity matching corresponding to the voiceprint feature; The identity prediction value is output through the output layer, and the voiceprint recognition result is determined based on the identity prediction value.
7. The method according to claim 1, characterized in that Before acquiring the initial voice signal input by the participant through the voice acquisition module, the following steps are also included: detecting an input level value, and determining that the target operating mode is a pure signaling mode if the level value is not greater than a preset first level value threshold; When the level value is greater than a preset second level value threshold, it is determined that the target working mode is a pure voice mode.
8. The method according to claim 1, characterized in that Also includes: When the target working mode is the pure voice mode, the initial voice signal input by the participant is collected by the voice collection module, and the initial voice signal is converted into voice text data; The voice text data is parsed by a signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction.
9. A voice control device for a conference, characterized in that: include: A voice signal acquisition module is configured to acquire an initial voice signal input by a participant through a voice acquisition module and perform deep noise reduction processing on the initial voice signal to obtain a target voice signal when the target operating mode is a pure signaling mode; wherein the pure signaling mode is an operating mode for transmitting various control instructions; A voiceprint recognition module is configured to extract voiceprint features from the target speech signal through the voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine a voiceprint recognition result based on an output result of the model; A task execution module is configured to, when the voiceprint recognition result indicates that a matching participant identity exists, parse the target voice signal through the signaling processing module to obtain a target task execution instruction, and execute the target task based on the target task execution instruction; The first indicator light control module is used to indicate the The light control module controls the voice control indicator light to remain on; The second indicator light control module is used to indicate when the target working mode is pure voice mode. The light control module controls the voice control indicator light to remain off; wherein, the pure voice mode is a working mode for receiving various voice messages; The switch control module is used to control the device to switch the target working mode through the switch on the shielding cover of the device mainboard.
Citation Information
Patent Citations
AI intelligent conference system based on voice and semantics and implementation method thereof
CN109474763A
Voice control method, device and equipment, medium and intelligent voice acquisition system
CN114512127A
Identity authentication method and device based on voiceprint recognition, computer equipment and storage medium
CN117668801A