Voice control method and device for conference

By adopting deep noise reduction processing and voiceprint recognition technology in the conference system, the recognition accuracy and efficiency of traditional voice control methods under environmental noise are solved, and more efficient voice control and conference management are achieved.

CN120108402AActive Publication Date: 2025-06-06BEIJING HUAJIAN YUNDING TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510578459.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

Traditional voice control methods are difficult to accurately recognize voice commands under the influence of environmental noise, and there is a problem of low control efficiency.

Method used

In pure signaling mode, the initial voice signal is obtained through the voice acquisition module and deep noise reduction processing is performed to obtain a clear target voice signal. Then, the voiceprint features are extracted through the voiceprint recognition module and input them into the pre-trained voiceprint recognition model, determine the identity of the participant, and parse the voice signal to perform the task.

Benefits of technology

Through deep noise reduction processing and voiceprint recognition technology, the accuracy and efficiency of voice control are improved, the misrecognition rate is reduced, and more efficient meeting management is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108402A_ABST
    Figure CN120108402A_ABST
Patent Text Reader

Abstract

The invention discloses a voice control method and device for a conference. The method comprises the following steps: when a target working mode is a pure signaling mode, acquiring an initial voice signal input by a participant through a voice acquisition module, and performing deep noise reduction processing on the initial voice signal to obtain a target voice signal; extracting voiceprint features in the target voice signal through a voiceprint recognition module, inputting the voiceprint features into a pre-trained voiceprint recognition model, and determining a voiceprint recognition result based on a model output result; and under the condition that the voiceprint recognition result is that the matched participant identity exists, analyzing the target voice signal through a signaling processing module to obtain a target task execution instruction, and executing a target task based on the target task execution instruction. Conference efficiency is improved, and intelligent management is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent conference control, and in particular to a method and device for voice control of a conference. Background Art

[0002] With the rapid development of information technology, the intelligence of conference systems has become a key trend to improve conference efficiency and optimize conference experience. Traditional conference systems often rely on manual operations, such as manually controlling conference equipment and manually recording conference content. These methods are not only inefficient but also prone to errors.

[0003] As one of the core technologies for realizing the intelligent conference system, speech signal processing technology has made significant progress in recent years, providing strong technical support for the intelligent conference system. Traditional voice control methods have the problem that environmental noise cannot be filtered out, resulting in inaccurate recognition of voice commands, and there is a certain delay, resulting in low control efficiency. Summary of the invention

[0004] The present invention provides a method and device for voice control of a conference, so as to solve the problems of low voice control efficiency and low recognition accuracy.

[0005] According to one aspect of the present invention, a method for voice control of a conference is provided, the method comprising:

[0006] When the target working mode is the pure signaling mode, the initial voice signal input by the participant is collected by acquiring the voice collection module, and the initial voice signal is subjected to deep noise reduction processing to obtain the target voice signal;

[0007] The voiceprint features in the target voice signal are extracted through the voiceprint recognition module, and the voiceprint features are input into a pre-trained voiceprint recognition model, and the voiceprint recognition result is determined based on the model output result; when the voiceprint recognition result shows that there is a matching participant identity, the target voice signal is parsed through the signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction.

[0008] According to another aspect of the present invention, a voice control method and device is provided, the device comprising:

[0009] A voice signal acquisition module is used to acquire an initial voice signal input by a participant through a voice acquisition module when the target working mode is a pure signaling mode, and perform deep noise reduction processing on the initial voice signal to obtain a target voice signal;

[0010] A voiceprint recognition module is used to extract voiceprint features from the target voice signal through the voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine the voiceprint recognition result based on the model output result; a task execution module is used to parse the target voice signal through the signaling processing module when the voiceprint recognition result shows that there is a matching participant identity, so as to obtain a target task execution instruction, and execute the target task based on the target task execution instruction.

[0011] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0012] at least one processor; and

[0013] a memory communicatively connected to the at least one processor; wherein,

[0014] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the voice control method for a conference described in any embodiment of the present invention.

[0015] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the voice control method for a conference described in any embodiment of the present invention when executed.

[0016] The technical solution of the embodiment of the present invention is to obtain the initial voice signal input by the participant through the voice acquisition module when the target working mode is the pure signaling mode, and perform deep noise reduction processing on the initial voice signal to obtain the target voice signal; the target voice signal after noise reduction is clearer, providing a high-quality data basis for subsequent voiceprint recognition. Then, the voiceprint features in the target voice signal are extracted by the voiceprint recognition module, and the voiceprint features are input into the pre-trained voiceprint recognition model, and the voiceprint recognition result is determined based on the model output result; it can capture the unique features of the voice signal more comprehensively. The pre-trained voiceprint recognition model can more accurately match the identity of the participant on the basis of high-quality data and reduce the misrecognition rate. Finally, when the voiceprint recognition result shows that there is a matching participant identity, the target voice signal is parsed by the signaling processing module to obtain the target task execution instruction, and the target task is executed based on the target task execution instruction, which solves the problems of low voice control efficiency and low recognition accuracy, and achieves the beneficial effects of improving meeting efficiency and realizing intelligent management.

[0017] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 is a flow chart of a voice control method for a conference provided according to Embodiment 1 of the present invention;

[0020] Figure 2 is a flow chart of a voice control method for a conference provided according to Embodiment 2 of the present invention;

[0021] Figure 3 is a structural diagram of a voice control device for a conference provided according to Embodiment 3 of the present invention;

[0022] Figure 4 The present invention is a schematic diagram of the structure of an electronic device for implementing the voice control method for a conference according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] Embodiment 1

[0026] Figure 1 A flowchart of a conference voice control method is provided for the first embodiment of the present invention. This embodiment is applicable to the conference voice control situation. The method can be executed by a conference voice control device. The conference voice control device can be implemented in the form of hardware and / or software. The conference voice control device can be configured in an electronic device. Figure 1 As shown, the method includes:

[0027] S110. When the target working mode is the pure signaling mode, an initial voice signal input by a participant is collected by acquiring a voice collection module, and deep noise reduction processing is performed on the initial voice signal to obtain a target voice signal.

[0028] The initial voice signal can be understood as pulse code modulation signal data. The target voice signal can be understood as the voice signal data after denoising. The pure signaling mode can be understood as a working mode that focuses on the accurate transmission of various control instructions to ensure the stable operation and flexible control of the system.

[0029] Specifically, in the pure signaling mode, the voice acquisition module will first obtain PCM (pulse code modulation) data, which carries key voice information. To improve voice quality, the collected PCM data is subjected to deep noise reduction processing by introducing automatic gain control, automatic echo cancellation and automatic noise suppression, effectively filtering out interference factors such as environmental noise and echo, so that subsequent processing can be carried out based on clearer and purer voice signals.

[0030] S120. Extracting voiceprint features from the target speech signal through a voiceprint recognition module, inputting the voiceprint features into a pre-trained voiceprint recognition model, and determining a voiceprint recognition result based on an output result of the model.

[0031] The voiceprint feature can be understood as a set of speech features that characterize and identify the speaker. The voiceprint recognition result includes the result of identifying the speaker, including at least one of the speaker's identity, the matching score with each voiceprint template in the database, and the matching status (success or failure).

[0032] Specifically, the target speech signal after noise reduction processing is accurately imported into the voiceprint recognition module. In this module, two advanced feature extraction technologies, Mel frequency cepstral coefficients and linear prediction cepstral coefficients, are used to extract voiceprint features from the processed speech signal. The extracted voiceprint features are input into the pre-trained voiceprint recognition model for matching operations, thereby completing the voiceprint recognition task efficiently and accurately, providing a solid guarantee for the security and accuracy of the system.

[0033] Optionally, the voiceprint features in the target voice signal are extracted through a voiceprint recognition module, including: preprocessing the target voice signal to obtain a target digital signal, and extracting Mel-frequency cepstral coefficients and linear prediction cepstral coefficients of the target digital signal; the preprocessing includes at least two of signal conversion processing, enhancement processing, windowing processing and frame processing; and determining the voiceprint features in the target voice signal based on a feature fusion method, a pattern recognition algorithm, the Mel-frequency cepstral coefficients and the linear prediction cepstral coefficients.

[0034] Specifically, signal conversion processing converts the original speech signal into a digital signal suitable for subsequent processing. The analog speech signal is sampled and quantized through an analog-to-digital converter (ADC) to obtain a digital signal. Enhancement processing uses a noise reduction algorithm (such as spectral subtraction, Wiener filtering, deep noise reduction algorithm, etc.) to reduce the noise of the speech signal. Windowing processing multiplies the speech signal by a function such as a Hamming window to make the signal smoother in the time domain. Framing processing divides the speech signal into frames according to a fixed frame length and frame shift.

[0035] Specifically, voiceprint features that can characterize the identity of the speaker are extracted from the target digital signal. Common voiceprint features include Mel frequency cepstral coefficients and linear prediction cepstral coefficients. Features such as Mel frequency cepstral coefficients and linear prediction cepstral coefficients can be fused by serial fusion, parallel fusion or weighted fusion. The fused voiceprint features are classified and identified to determine the identity of the speaker. Pattern recognition algorithms can include Gaussian mixture models, deep neural networks and support vector machines.

[0036] Optionally, extracting Mel-frequency cepstrum coefficients of the target digital signal includes:

[0037] Performing spectral analysis on the target digital signal to obtain a power spectrum corresponding to the target digital signal, and determining a spectrum based on the power spectrum; performing Mel filtering on the spectrum, performing discrete cosine transform based on the filtered spectrum to obtain the Mel-frequency cepstrum coefficients, and extracting the Mel-frequency cepstrum coefficients.

[0038] Specifically, a fast Fourier transform is performed on the target digital signal to convert the time domain signal into a frequency domain signal. The complex spectrum of the signal is obtained, which contains amplitude and phase information. The amplitude of the complex spectrum is squared to obtain the power spectrum of the signal. The power spectrum reflects the energy distribution of the signal at different frequencies. The power spectrum is logarithmized to compress the dynamic range and enhance the visibility of low-amplitude components. A set of bandpass filters are designed on the Mel frequency scale, each of which covers a certain frequency range. The power spectrum or logarithmic power spectrum is passed through the Mel filter bank to obtain the energy at the output of each filter, and then a Mel frequency energy spectrum is obtained, in which each element corresponds to the output energy of a Mel filter. The Mel frequency energy spectrum is converted into cepstrum coefficients by discrete cosine transform to remove the correlation between filters and obtain a more compact feature representation.

[0039] For example, using the first-order finite difference equation , enhance the high frequency part, make the signal spectrum flatter, compensate for the natural attenuation of the target digital signal in the high frequency part, highlight the high frequency information in the speech signal, and facilitate the subsequent feature extraction; use the function on the original speech signal , perform frame processing to decompose the target digital signal into a series of short-term stable signals, which is convenient for subsequent feature extraction on each frame; use the Hamming window function Perform windowing on each frame of signal to reduce spectrum leakage caused by frame truncation and make the energy of each frame of signal in the frequency domain more concentrated; perform fast Fourier transform on each frame of signal after windowing to obtain the spectrum , convert the time domain signal into a frequency domain signal to obtain the spectrum information of the speech signal; construct a set of Mel filters with a center frequency The conversion relationship between (unit: Hz) and Mel frequency m (unit: Mel) is For example, suppose P Mel filters are constructed, and for each filter (p is between 20-40), calculate the spectrum after Mel filtering , simulating the human ear's perception of frequency. Taking the logarithm gives , converting the multiplication operation of the spectrum into an addition operation, compressing the dynamic range, reducing the error of numerical calculation, and simulating the human ear's perception of sound intensity to a certain extent; taking the logarithm of the Mel spectrum Perform DCT to obtain the Mel frequency cepstrum coefficients (C value is between 12-16).

[0040] Optionally, extracting the linear prediction cepstral coefficients of the target digital signal includes: performing linear prediction analysis on the target digital signal to obtain linear prediction coefficients corresponding to the target digital signal; converting the linear prediction coefficients into linear prediction cepstral coefficients, and extracting the linear prediction cepstral coefficients.

[0041] Specifically, the target digital signal is divided into short time frames, each frame usually contains a preset millisecond of voice data, and the autocorrelation function of each frame signal is calculated to estimate the periodicity of the signal. The linear prediction equation is solved using the autocorrelation coefficient, and the linear prediction coefficient is converted into the cepstral domain through a recursive formula to obtain the linear prediction cepstral coefficients, which have better robustness and discrimination ability in speech recognition and voiceprint recognition.

[0042] Exemplarily, the function is used on the target digital signal , perform frame processing to decompose the speech signal into a series of short-term stable signals, which is convenient for subsequent feature extraction on each frame;

[0043] Use the Hamming window function Perform windowing on each frame of signal. With window function Multiply to get the windowed frame ; Use the autocorrelation function (p is the linear prediction order, ranging from 10-16); linear prediction equation ,in is the target digital signal is the linear prediction coefficient, p is the prediction order, is the prediction error. When solving by the autocorrelation method, the autocorrelation matrix R and vector r are constructed based on the calculated autocorrelation function, and then the equation is solved Get the linear prediction coefficient According to the linear prediction coefficient , the prediction error energy can be calculated , and thus obtain the gain ; According to the linear prediction coefficient Calculate the linear prediction cepstral coefficients. By recursive formula , (C is the order of the linear prediction cepstral coefficients, ranging from 12 to 16).

[0044] Optionally, before the voiceprint features are input into a pre-trained voiceprint recognition model and the voiceprint recognition results are determined based on the model output results, the method further includes: constructing a sample set, and dividing the sample set into a training set and a test set based on a preset ratio; wherein the sample set includes a preset number of sample voiceprint features and identity labels corresponding to the sample voiceprint features; performing iterative training for a preset number of times on a pre-established deep neural network based on the training set, and in each iteration, calculating the loss based on the probability prediction value output by the model and the waiting identity label, and adjusting the weight of the network through a back propagation algorithm; and testing each trained model based on the test set to obtain multiple model performance indicators, and determining the voiceprint recognition model based on the multiple model performance indicators.

[0045] Specifically, select a suitable network architecture according to the task requirements and data characteristics. Determine parameters such as the number of neurons in each layer, activation function, convolution kernel size and stride (for CNN), and number of recurrent units (for RNN and its variants). In the input layer, the number of neurons should match the dimension of the input voiceprint feature vector. The number of neurons in the hidden layer can be adjusted experimentally, generally between dozens and hundreds. The activation function can be selected from ReLU, tanh, etc. The output layer uses the softmax function for multi-classification (assuming that multiple speakers are recognized) according to the recognition task, and outputs the probability distribution of each speaker. Collect a large number of speech samples from different speakers and accurately annotate them to ensure that each sample corresponds to the correct speaker identity label. The number of samples for each speaker should be as balanced as possible, and cover different speech content, speaking speed, intonation, etc., to improve the generalization ability of the network. For each speech sample, extract the voiceprint feature vector. Then, normalize the feature vector and scale its numerical range to an appropriate interval. Use normal distribution or evenly distributed To initialize the weight matrix The bias term is usually initialized to 0 or a small constant such as 0.1.

[0046] Specifically, the preprocessed voiceprint feature vector is used as the input of the network and passes through each layer of the network in turn. In the convolution layer, the convolution kernel is used to perform a convolution operation with the input feature map to extract local features, and then the output feature map is obtained through activation function processing; in the recurrent layer, according to the sequence characteristics of the speech signal, the features of each time step are processed in turn, the hidden state is updated, and the final hidden state is output to the next layer; the fully connected layer performs a linear transformation and activation function processing on the output of the previous layer, gradually maps the features to a higher level of abstract representation, and finally obtains the probability prediction value of each speaker in the output layer.

[0047] Specifically, use the cross entropy loss function formula: , substitute the network's output prediction value and the corresponding true label into the loss function to calculate the loss value of the current batch of samples. Based on the calculated loss value, use the back propagation algorithm to calculate the gradient of each weight parameter in the network. Use the calculated gradient value to update the network's weight parameters according to the adaptive moment estimation optimization algorithm. In each iteration, the weight update formula is: ,By repeating the process of forward propagation, calculating the loss function, back propagation and weight updating, the network gradually learns the model parameters that can accurately identify the voiceprints of different speakers.

[0048] Specifically, the above forward propagation, loss calculation, back propagation and weight update processes are encapsulated in a training loop, and the model is iteratively trained multiple times based on the entire training set until the performance of the network (such as accuracy, loss value and other indicators) no longer improves or reaches the preset training stop condition. In each iteration, the training data is usually divided into multiple small batches. During the training process, some hyperparameters need to be adjusted to optimize the performance of the network. These hyperparameters include learning rate, batch size, various parameters in the network architecture (such as the number of layers, number of neurons, etc.), number of training iterations, etc.

[0049] Specifically, during the training process, the model is regularly evaluated using the test set to monitor the model's performance and avoid overfitting. By observing how the model performance indicators on the test set change with training iterations, it is possible to determine whether the model is overfitting (if the indicators on the test set no longer improve or even decrease, while the indicators on the training set continue to rise, overfitting may occur), and timely measures can be taken, such as stopping training early, adding regularization terms, etc. By evaluating the test set and selecting the model with the best performance, the accuracy and robustness of voiceprint recognition can be significantly improved.

[0050] Optionally, the voiceprint recognition model includes a convolutional layer, a recurrent layer connected to the convolutional layer, a fully connected layer connected to the recurrent layer, and an output layer connected to the fully connected layer; the step of inputting the voiceprint feature into a pre-trained voiceprint recognition model and determining the voiceprint recognition result based on the model output result comprises: inputting the voiceprint feature into a pre-trained voiceprint recognition model, extracting local features in the voiceprint feature through the convolutional layer, and generating a feature map; determining a feature sequence corresponding to the feature map, processing the features of each time step in sequence according to the sequence characteristics of the feature sequence through the recurrent layer, updating the hidden state, and outputting the updated hidden state to the fully connected layer; performing linear transformation and activation function processing on the updated hidden state through the fully connected layer, mapping the features of the updated hidden state to an identity prediction value corresponding to the voiceprint feature and matching the speaker identity; outputting the identity prediction value through the output layer, and determining the voiceprint recognition result based on the identity prediction value.

[0051] Specifically, the extracted voiceprint features are input into the model. The convolution layer extracts local patterns in the voiceprint features and generates feature maps. The feature maps are converted into feature sequences, and the recurrent layer processes the features of each time step in turn and updates the hidden state. The fully connected layer performs linear transformation and activation function processing on the hidden state to obtain the identity prediction value. The output layer outputs the identity prediction value and selects the speaker identity with the highest probability as the voiceprint recognition result.

[0052] S130. When the voiceprint recognition result shows that there is a matching participant identity, the target voice signal is parsed by a signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction.

[0053] Among them, the target task execution instruction can be understood as a specific operation command generated by input (such as voice instructions, text instructions, etc.).

[0054] Specifically, first confirm the voiceprint recognition result, that is, the target voice signal successfully matches the voiceprint feature of a participant in the database. After confirming the existence of a matching participant identity, the system automatically activates the signaling processing module. The target voice signal is input into the signaling processing module for parsing. Natural language processing technology is used to identify keywords and phrases in the voice signal. These keywords are usually related to the target task execution instructions. The identified keywords and phrases are semantically understood to determine their meaning and intention. Based on the semantic understanding results, specific target task execution instructions are generated.

[0055] Optionally, the method further includes:

[0056] When the target working mode is the pure voice mode, the initial voice signal input by the participant is collected by the voice acquisition module, and the initial voice signal is converted into voice text data; the voice text data is parsed by the signaling processing module to obtain the target task execution instruction, and the target task is executed based on the target task execution instruction.

[0057] Among them, the pure voice mode can be understood as a mode that aims to achieve efficient transmission of normal conference sounds and ensure that participants can receive various voice information clearly and smoothly.

[0058] Specifically, in pure voice mode, a series of voice preprocessing operations are first performed on the collected initial voice signal, including the use of low-pass, high-pass and band-pass filtering techniques to remove unnecessary frequency components, and the signal is appropriately amplified to optimize its quality and strength, thereby laying a good foundation for subsequent processing procedures. Subsequently, the voice recognition module converts the preprocessed initial voice signal into voice text data, and then the signaling processing module conducts in-depth analysis of these text information, extracting the key intentions and instructions to drive the corresponding task execution process.

[0059] Optionally, other result data from the task execution process is efficiently received through the audio output module and converted into analog audio signals. This conversion process is achieved through a high-precision digital-to-analog converter (DAC). Finally, with the help of carefully adapted third-party audio playback devices, such as high-quality speakers or professional headphones, the converted analog audio signals are perfectly presented, creating a smooth, accurate and immersive pure voice interaction experience for users, fully meeting the diverse needs and expectations of users in pure voice mode.

[0060] For example, when there is important information to be announced in a meeting or the speech of a participant needs to be heard by other participants, the system converts the text information into a natural and fluent voice signal through speech synthesis technology.

[0061] Use high-precision digital-to-analog converters to convert digital voice signals into analog audio signals, and perform appropriate audio amplification on the analog audio signals. Finally, the processed audio signals are played through well-adapted third-party audio playback equipment (such as the speaker system in the conference room), achieving efficient and accurate interaction between signaling and users.

[0062] Optionally, the method further includes: when the target working mode is a pure signaling mode, by indicating

[0063] The light control module controls the voice control indicator light to remain in a normally on state; when the target working mode is the pure voice mode, the voice control indicator light is controlled by the indicator light control module to remain in an off state.

[0064] Specifically, in the pure signaling mode, in order to clearly indicate the current working status of the system, the voice control indicator light should remain on, so that the user or operator can intuitively see that the system is in the pure signaling working mode.

[0065] In an embodiment of the present invention, a "voice control" indicator light is added so that the operator can more intuitively distinguish between pure voice mode and pure signaling mode. In pure voice mode, the voice control indicator light should remain off. For example, assume that there is an intelligent conference system that supports two working modes: pure signaling mode and pure voice mode. When the system switches to pure signaling mode, the voice control indicator light will light up, indicating that the system is receiving and processing instructions through signals; and when the system switches to pure voice mode, the indicator light will go out, indicating that the system is receiving and processing instructions through voice. In this way, the user can intuitively understand the current working mode of the system through the status of the indicator light.

[0066] The technical solution of the embodiment of the present invention is to obtain the initial voice signal input by the participant through the voice acquisition module when the target working mode is the pure signaling mode, and perform deep noise reduction processing on the initial voice signal to obtain the target voice signal; the target voice signal after noise reduction is clearer, providing a high-quality data basis for subsequent voiceprint recognition. Then, the voiceprint features in the target voice signal are extracted by the voiceprint recognition module, and the voiceprint features are input into the pre-trained voiceprint recognition model, and the voiceprint recognition result is determined based on the model output result; it can capture the unique features of the voice signal more comprehensively. The pre-trained voiceprint recognition model can more accurately match the identity of the participant on the basis of high-quality data and reduce the misrecognition rate. Finally, when the voiceprint recognition result shows that there is a matching participant identity, the target voice signal is parsed by the signaling processing module to obtain the target task execution instruction, and the target task is executed based on the target task execution instruction, which solves the problems of low voice control efficiency and low recognition accuracy, and achieves the beneficial effects of improving meeting efficiency and realizing intelligent management.

[0067] Embodiment 2

[0068] Figure 2 This is a flow chart of a method for voice control of a conference provided in Embodiment 2 of the present invention. This embodiment is a further optimization of the above embodiment. Figure 2 As shown, the method includes:

[0069] S210 , detecting an input level value, and determining that the target operating mode is a pure signaling mode when the level value is not greater than a preset first level value threshold.

[0070] The preset first level value threshold may be understood as a preset low level threshold, which may be preset based on experience, and is not limited in this embodiment.

[0071] Specifically, if the input level value is low, the system will quickly select the "pure signaling mode". During the operation of the system, the input level value will be accurately monitored in real time. Once the input level value is detected to have changed, it will immediately enter the mode detection state. If the input level value is low, the system will quickly select the "pure signaling mode", and accordingly, the "voice control" indicator light will remain on, providing users with clear mode indication.

[0072] S220: When the level value is greater than a preset second level value threshold, determine that the target working mode is a pure voice mode.

[0073] The preset second level value threshold may be understood as a preset high level threshold, which may be preset based on experience, and is not limited in this embodiment.

[0074] Specifically, during the operation of the system, the input level value will be accurately monitored in real time. Once the input level value is detected to have changed, it will immediately enter the mode detection state. When the input level value is high, the system will automatically switch to the "pure voice mode". At this time, the "voice control" indicator light remains off, indicating the current working mode to the user in an intuitive way. If no change in the level value is detected during the detection process, the system will automatically enter a loop detection process and continue to monitor the input level until the level change is successfully captured, thereby ensuring that the system can respond to different working mode requirements in a timely and accurate manner, providing users with a stable and efficient user experience.

[0075] S230: When the level value is greater than a preset second level value threshold, determine that the target working mode is a pure voice mode.

[0076] S240. Extract voiceprint features from the target speech signal through a voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine a voiceprint recognition result based on an output result of the model.

[0077] S250: When the voiceprint recognition result shows that there is a matching participant identity, the target voice signal is parsed by a signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction.

[0078] Optionally, a shielding cover can be added to the mainboard of the device, and a switch that can be adjusted up and down can be designed on the shielding cover. When the switch is turned down, the device will automatically switch to pure signaling mode. In this mode, all information exchanged between the device and external devices will be transmitted stably and efficiently through the serial port to ensure the fast and accurate transmission of signaling data such as control instructions, and meet the application scenarios such as remote control and system configuration that require high accuracy and timeliness. When the switch is turned up, the device will enter pure voice mode. Similarly, all kinds of voice information exchanged between the device and external devices will also be transmitted through the serial port to ensure smooth transmission of voice data and provide high-quality audio transmission channels for applications such as voice calls and voice broadcasts, thereby achieving stable and reliable information exchange between the device and external devices in different working modes.

[0079] The technical solution of the embodiment of the present invention determines that the target working mode is a pure signaling mode by detecting the input level value, when the level value is not greater than a preset first level value threshold. When the level value is greater than a preset second level value threshold, the target working mode is determined to be a pure voice mode. This ensures that the system can respond to different working mode requirements in a timely and accurate manner, and provides users with a stable and efficient use experience.

[0080] Embodiment 3

[0081] Figure 3 This is a schematic diagram of the structure of a voice control device for a conference provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a voice signal acquisition module 310, a voiceprint recognition module 320 and a task execution module 330.

[0082] Among them, the voice signal acquisition module 310 is used to, when the target working mode is the pure signaling mode, acquire the initial voice signal input by the participant by acquiring the voice acquisition module, and perform deep noise reduction processing on the initial voice signal to obtain the target voice signal; the voiceprint recognition module 320 is used to extract the voiceprint features in the target voice signal through the voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine the voiceprint recognition result based on the model output result; the task execution module 330 is used to parse the target voice signal through the signaling processing module when the voiceprint recognition result shows that there is a matching participant identity, so as to obtain the target task execution instruction, and execute the target task based on the target task execution instruction.

[0083] The technical solution of the embodiment of the present invention is to obtain the initial voice signal input by the participant through the voice acquisition module when the target working mode is the pure signaling mode, and perform deep noise reduction processing on the initial voice signal to obtain the target voice signal; the target voice signal after noise reduction is clearer, providing a high-quality data basis for subsequent voiceprint recognition. Then, the voiceprint features in the target voice signal are extracted by the voiceprint recognition module, and the voiceprint features are input into the pre-trained voiceprint recognition model, and the voiceprint recognition result is determined based on the model output result; it can capture the unique features of the voice signal more comprehensively. The pre-trained voiceprint recognition model can more accurately match the identity of the participant on the basis of high-quality data and reduce the misrecognition rate. Finally, when the voiceprint recognition result shows that there is a matching participant identity, the target voice signal is parsed by the signaling processing module to obtain the target task execution instruction, and the target task is executed based on the target task execution instruction, which solves the problems of low voice control efficiency and low recognition accuracy, and achieves the beneficial effects of improving meeting efficiency and realizing intelligent management.

[0084] Optionally, the voiceprint recognition module includes:

[0085] A feature extraction unit, used for preprocessing the target speech signal to obtain a target digital signal, and extracting Mel frequency cepstral coefficients and linear prediction cepstral coefficients of the target digital signal; the preprocessing includes at least two of signal conversion processing, enhancement processing, windowing processing and frame processing;

[0086] A voiceprint feature extraction unit is used to determine the voiceprint features in the target speech signal based on a feature fusion method, a pattern recognition algorithm, the Mel-frequency cepstral coefficients and the linear prediction cepstral coefficients.

[0087] Optionally, the feature extraction unit includes:

[0088] A spectrum analysis subunit, configured to perform spectrum analysis on the target digital signal to obtain a power spectrum corresponding to the target digital signal, and determine a spectrum based on the power spectrum;

[0089] The mel filter subunit is used to perform mel filtering on the frequency spectrum, perform discrete cosine transform based on the filtered frequency spectrum to obtain the mel-frequency cepstral coefficients, and extract the mel-frequency cepstral coefficients.

[0090] Optionally, the feature extraction unit includes:

[0091] A linear prediction subunit, configured to perform a linear prediction analysis on the target digital signal to obtain a linear prediction coefficient corresponding to the target digital signal;

[0092] The linear conversion subunit is used to convert the linear prediction coefficients into linear prediction cepstral coefficients and extract the linear prediction cepstral coefficients.

[0093] Optionally, the device further comprises:

[0094] A sample set construction module, used for constructing a sample set before inputting the voiceprint features into a pre-trained voiceprint recognition model and determining the voiceprint recognition result based on the model output result, and dividing the sample set into a training set and a test set based on a preset ratio; wherein the sample set includes a preset number of sample voiceprint features and identity tags corresponding to the sample voiceprint features;

[0095] A model training module, used to perform a preset number of iterative training on the pre-established deep neural network based on the training set, and in each iteration, the loss is calculated according to the probability prediction value output by the model and the waiting identity label, and the weight of the network is adjusted by the back propagation algorithm;

[0096] The model determination module is used to test each trained model based on the test set to obtain multiple model performance indicators, and determine the voiceprint recognition model based on the multiple model performance indicators.

[0097] Optionally, the voiceprint recognition model includes a convolutional layer, a recurrent layer connected to the convolutional layer, a fully connected layer connected to the recurrent layer, and an output layer connected to the fully connected layer; accordingly, the voiceprint recognition module includes:

[0098] A convolution unit, used for inputting the voiceprint features into a pre-trained voiceprint recognition model, extracting local features from the voiceprint features through the convolution layer, and generating a feature map;

[0099] A hidden update unit, used to determine a feature sequence corresponding to the feature map, process the features of each time step in sequence according to the sequence characteristics of the feature sequence through the recurrent layer, update the hidden state, and output the updated hidden state to the fully connected layer;

[0100] A prediction value determination unit, configured to perform linear transformation and activation function processing on the updated hidden state through the fully connected layer, and map the features of the updated hidden state to an identity prediction value that matches the speaker identity corresponding to the voiceprint feature;

[0101] The result output unit is used to output the identity prediction value through the output layer, and determine the voiceprint recognition result based on the identity prediction value.

[0102] Optionally, the device further comprises:

[0103] A first mode determination module, configured to detect a level value of the input before acquiring an initial voice signal input by a participant through a voice acquisition module, and determine that the target working mode is a pure signaling mode when the level value is not greater than a preset first level value threshold;

[0104] The second mode determination module is used to determine that the target working mode is a pure voice mode when the level value is greater than a preset second level value threshold.

[0105] Optionally, the device further comprises:

[0106] A text conversion module, used for, when the target working mode is the pure voice mode, collecting the initial voice signal input by the participant by acquiring the voice collection module, and converting the initial voice signal into voice text data;

[0107] The text parsing module is used to parse the voice text data through the signaling processing module to obtain a target task execution instruction, and execute the target task based on the target task execution instruction.

[0108] Optionally, the device further comprises:

[0109] The first indicator light control module is used to indicate when the target working mode is the pure signaling mode.

[0110] The light control module controls the voice control indicator light to remain on;

[0111] The second indicator light control module is used to indicate when the target working mode is the pure voice mode.

[0112] The light control module controls the voice control indicator light to remain off.

[0113] The voice control device for a conference provided in the embodiment of the present invention can execute the voice control method for a conference provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0114] Embodiment 4

[0115] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0116] like Figure 4 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0117] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0118] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as voice control of a method conference.

[0119] In some embodiments, the voice control of the method conference can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the voice control of the method conference described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the voice control of the method conference by any other suitable means (for example, by means of firmware).

[0120] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0122] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide interaction with a service acquirer, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the service acquirer; and a keyboard and a pointing device (e.g., a mouse or trackball), through which the service acquirer can provide input to the electronic device. Other types of devices may also be used to provide interaction with the service acquirer; for example, the feedback provided to the service acquirer may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the service acquirer may be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a service acquirer computer having a graphical service acquirer interface or a web browser through which the service acquirer can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0125] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0126] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0127] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for voice control of a conference, characterized in that: include: When the target working mode is the pure signaling mode, the initial voice signal input by the participant is collected by acquiring the voice collection module, and the initial voice signal is subjected to deep noise reduction processing to obtain the target voice signal; Extracting voiceprint features from the target voice signal through a voiceprint recognition module, inputting the voiceprint features into a pre-trained voiceprint recognition model, and determining a voiceprint recognition result based on an output result of the model; When the voiceprint recognition result shows that there is a matching participant identity, the target voice signal is parsed by the signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction.

2. The method according to claim 1, characterized in that Extracting the voiceprint features in the target voice signal through a voiceprint recognition module includes: Preprocessing the target speech signal to obtain a target digital signal, and extracting Mel-frequency cepstral coefficients and linear prediction cepstral coefficients of the target digital signal; the preprocessing includes at least two of signal conversion processing, enhancement processing, windowing processing and frame processing; The voiceprint features in the target speech signal are determined based on a feature fusion method, a pattern recognition algorithm, the Mel-frequency cepstral coefficients and the linear prediction cepstral coefficients.

3. The method according to claim 2, characterized in that The step of extracting Mel-frequency cepstrum coefficients of the target digital signal comprises: Performing spectrum analysis on the target digital signal to obtain a power spectrum corresponding to the target digital signal, and determining a spectrum based on the power spectrum; Mel filtering is performed on the frequency spectrum, and discrete cosine transform is performed based on the filtered frequency spectrum to obtain the Mel-frequency cepstral coefficients, and the Mel-frequency cepstral coefficients are extracted.

4. The method according to claim 2, characterized in that: The step of extracting linear prediction cepstral coefficients of the target digital signal comprises: Performing a linear prediction analysis on the target digital signal to obtain a linear prediction coefficient corresponding to the target digital signal; The linear prediction coefficients are converted into linear prediction cepstral coefficients, and the linear prediction cepstral coefficients are extracted.

5. The method according to claim 1, characterized in that: Before inputting the voiceprint feature into a pre-trained voiceprint recognition model and determining the voiceprint recognition result based on the model output result, the method further includes: Constructing a sample set, and dividing the sample set into a training set and a test set based on a preset ratio; wherein the sample set includes a preset number of sample voiceprint features and identity tags corresponding to the sample voiceprint features; Performing a preset number of iterative training on the pre-established deep neural network based on the training set, in each iteration, calculating the loss according to the probability prediction value output by the model and the waiting identity label, and adjusting the weight of the network through the back propagation algorithm; Each trained model is tested based on the test set to obtain a plurality of model performance indicators, and the voiceprint recognition model is determined based on the plurality of model performance indicators.

6. The method according to claim 1, characterized in that The voiceprint recognition model includes a convolutional layer, a circulation layer connected to the convolutional layer, a fully connected layer connected to the circulation layer, and an output layer connected to the fully connected layer; The step of inputting the voiceprint feature into a pre-trained voiceprint recognition model and determining the voiceprint recognition result based on the model output result includes: Inputting the voiceprint feature into a pre-trained voiceprint recognition model, extracting local features in the voiceprint feature through the convolution layer, and generating a feature map; Determine a feature sequence corresponding to the feature graph, process the features of each time step in sequence according to the sequence characteristics of the feature sequence through the recurrent layer, update the hidden state, and output the updated hidden state to the fully connected layer; Performing linear transformation and activation function processing on the updated hidden state through the fully connected layer, mapping the features of the updated hidden state to the identity prediction value of the speaker identity matching corresponding to the voiceprint feature; The identity prediction value is output through the output layer, and the voiceprint recognition result is determined based on the identity prediction value.

7. The method according to claim 1, characterized in that Before acquiring the initial voice signal input by the participant through the voice acquisition module, it also includes: detecting an input level value, and determining that the target working mode is a pure signaling mode when the level value is not greater than a preset first level value threshold; When the level value is greater than a preset second level value threshold, it is determined that the target working mode is a pure voice mode.

8. The method according to claim 1, characterized in that Also includes: When the target working mode is the pure voice mode, the initial voice signal input by the participant is collected by acquiring the voice collection module, and the initial voice signal is converted into voice text data; The voice text data is parsed by a signaling processing module to obtain a target task execution instruction, and the target task is executed based on the target task execution instruction.

9. The method according to claim 1, characterized in that: Also includes: When the target working mode is the pure signaling mode, the voice control indicator light is controlled by the indicator light control module to remain on; When the target working mode is the pure voice mode, the voice control indicator light is controlled by the indicator light control module to remain off.

10. A voice control device for a conference, characterized in that: include: A voice signal acquisition module is used to acquire an initial voice signal input by a participant through a voice acquisition module when the target working mode is a pure signaling mode, and perform deep noise reduction processing on the initial voice signal to obtain a target voice signal; A voiceprint recognition module is used to extract voiceprint features from the target voice signal through the voiceprint recognition module, input the voiceprint features into a pre-trained voiceprint recognition model, and determine a voiceprint recognition result based on an output result of the model; The task execution module is used to parse the target voice signal through the signaling processing module to obtain a target task execution instruction when the voiceprint recognition result shows that there is a matching participant identity, and execute the target task based on the target task execution instruction.

Citation Information

Patent Citations

  • AI intelligent conference system based on voice and semantics and implementation method thereof

    CN109474763A

  • Conference participant voiceprint recognition method and device, electronic equipment and storage medium

    CN113643708A

  • Voice control method, device and equipment, medium and intelligent voice acquisition system

    CN114512127A

  • Personnel voiceprint recognition, authentication, noise reduction and voice enhancement method, system and device for power dispatching system

    CN116312561A

  • Identity authentication method and device based on voiceprint recognition, computer equipment and storage medium

    CN117668801A