A laboratory voice monitoring system based on speech recognition and natural language processing

By using speech recognition and natural language processing technology in the laboratory voice monitoring system, the single-person voice information of multi-person conversations in the laboratory is separated and identified, and the risk level is obtained by combining the risk information database to obtain risk levels, the problem of poor recognition effect of the existing system in multi-person conversation scenarios is solved, and more efficient laboratory safety monitoring is achieved.

CN114927124BActive Publication Date: 2025-06-20SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210208601.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-04
Publication Date
2025-06-20
Estimated Expiration
2042-03-04

AI Technical Summary

Technical Problem

When the existing laboratory voice monitoring system handles multi-person conversation scenarios, it is difficult to effectively extract and identify conversation information, and it is impossible to effectively monitor the laboratory situation.

Method used

The laboratory voice monitoring system based on speech recognition and natural language processing is adopted to obtain laboratory audio information through the speech acquisition module, and the single-person voice information is separated by spatial clustering algorithm and speaker clustering segmentation algorithm, and features are extracted through natural language processing, and risk levels are obtained in combination with the risk information database.

Benefits of technology

It realizes effective identification and monitoring of multi-person dialogue scenarios in the laboratory, improves the accuracy of speech recognition, can accurately identify and output risk levels, and enhances the laboratory's safety monitoring capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114927124B_ABST
    Figure CN114927124B_ABST
Patent Text Reader

Abstract

The present invention relates to a laboratory voice monitoring system based on speech recognition and natural language processing, which includes a voice acquisition module, a communication module, and a server management module. The voice acquisition module is connected to the server management module through the communication module: the voice acquisition module is used to collect audio information in the laboratory and optimize and separate the voice information based on an auxiliary function; the communication module is used to transmit the voice information to the server management module; the server management module is used to separate the voice information into multiple single-person voice information according to a spatial clustering algorithm, and perform annotation on each single-person voice information in the time dimension according to a speaker clustering segmentation algorithm. After extracting features from the annotated voice information, the features are combined with a risk information database to obtain the risk level corresponding to the voice information, and the risk level, voice information, and annotation information are output. Compared with the prior art, the present invention has the advantages of being able to separate the voices of speakers in the laboratory and having high recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of speech recognition and natural language processing, and in particular to a laboratory voice monitoring system based on speech recognition and natural language processing. Background Art

[0002] Laboratories are important places in the process of scientific research and teaching, and experimental teaching is an important link in cultivating the comprehensive abilities of talents. With the rapid development of the economy and the progress of science, various laboratories have been significantly improved both in terms of quantity and equipment. In order to give full play to the functionality of laboratories and ensure their safety, the remote integrated monitoring technology of laboratories plays an important role. In recent years, the intelligent monitoring technology of laboratories has been continuously developed, and the requirements for monitoring have become higher and higher.

[0003] Most of the current laboratory integrated monitoring system solutions are constructed based on Internet of Things technology. The monitoring methods mainly include security (video, access control, etc.), gas (conventional gas, toxic and harmful gas, etc.), environment (temperature and humidity, air conditioning, etc.), laboratory equipment / systems, power (UPS, power lights, etc.), fire monitoring, etc. With the explosive growth of deep learning technology, speech recognition and natural language processing technologies have also developed rapidly, and the accuracy has been greatly improved. In large laboratories shared by multiple people, voice, as an important resource, can be used to monitor the safety of the laboratory. By real-time monitoring and recording the conversation situations of experimental personnel, the situation of the laboratory can be reflected from the audio perspective. However, when the existing laboratory voice monitoring system processes laboratory scenarios, due to the often complex situation of multiple people talking in the laboratory, the effect of extracting and recognizing conversation information is poor, and the monitoring of the laboratory situation cannot be achieved. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a laboratory voice monitoring system based on speech recognition and natural language processing.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] A laboratory voice monitoring system based on speech recognition and natural language processing, including a voice collection module, a communication module, and a server management module. The voice collection module is connected to the server management module through the communication module:

[0007] The voice collection module is used to collect audio information in the laboratory and optimize and separate the voice information based on an auxiliary function;

[0008] The communication module is used to transmit the voice information to the server management module;

[0009] The server management module is used to separate the voice information into multiple single-person voice information according to the spatial clustering algorithm, annotate each single-person voice information in the time dimension according to the speaker clustering and segmentation algorithm, extract features from the annotated voice information, combine the features with the risk information database to obtain the risk level corresponding to the voice information, and output the risk level, voice information and annotation information.

[0010] Further, the voice acquisition module includes a microphone array sub-module and a separation sub-module:

[0011] The microphone array sub-module includes multiple omnidirectional microphones arranged in the laboratory, which are used to collect audio information containing spatial position information;

[0012] The separation sub-module separates the audio information into voice information and noise information according to the independent vector analysis algorithm optimized based on the auxiliary function, in combination with the spatial position information of the audio information.

[0013] Further, the server management module includes a signal reception queue sub-module, a multi-person speech recognition sub-module, a natural language processing sub-module and a query analysis sub-module:

[0014] The signal reception queue sub-module is used to receive voice information and transmit it to the multi-person speech recognition sub-module;

[0015] The multi-person speech recognition sub-module includes a pre-processor and a speech recognizer. The pre-processor is used to enhance the signal of the voice information, and the speech recognizer is used to separate the signal-enhanced voice information by using deep learning combined with the spatial clustering algorithm based on the Gaussian mixture model to obtain multiple single-person voice information; perform cutting on the signal turning points of the voice information through speaker segmentation, and finally perform annotation on the cut voice information through speaker clustering;

[0016] The natural language processing sub-module uses an acoustic model to convert the voice information passing through the multi-person speech recognition sub-module into text information, extracts the features of the text information by using the BERT model, and matches the features with the risk information database corresponding to the current laboratory to obtain the risk level;

[0017] The query analysis sub-module outputs the voice information, annotation information and risk level.

[0018] Further, the pre-processor uses a multi-objective learning algorithm based on the long short-term memory model to optimize the algorithm objectives according to the logarithmic power spectrum feature and the ideal ratio mask, and enhance the signal of the voice information.

[0019] Further, the acoustic model extracts the Mel cepstral features of the voice information and outputs the text after decoding the features.

[0020] Furthermore, before separating the speech signals, the speech recognizer eliminates the reverberation of the speech information based on the generalized weighted prediction error algorithm.

[0021] Furthermore, before separating the speech signals, the speech recognizer performs the following steps:

[0022] S1. Adopt a blind separation algorithm based on Independent Vector Analysis (IVA) to obtain multiple mutually independent speech information, and restore the amplitude to the range of the original observed signal through the back-projection algorithm;

[0023] S2. Use a noise estimation algorithm of the recursive average algorithm controlled by the minimum value to estimate the stationary noise on each individual speech information;

[0024] S3. Combine the stationary noise information of other speech information to estimate the non-stationary noise of the current speech information;

[0025] S4. Use the decision-directed algorithm to calculate the prior and posterior signal-to-noise ratios and perform multi-signal filtering processing;

[0026] S5. Synthesize all the processed speech information by means of linear superposition.

[0027] Furthermore, when the total number of clusters in the speaker clustering reaches the default maximum number of speakers, or the clustering value exceeds the set threshold, stop the speaker clustering and output the clustering result as the annotation information.

[0028] Furthermore, the server management module further includes an alarm sub-module. After obtaining the risk level, it judges whether the risk level exceeds the threshold. If so, it alarms the administrator.

[0029] Furthermore, the communication module adopts a half-duplex network composed of RS485 interfaces.

[0030] Compared with the prior art, the present invention has the following advantages:

[0031] 1. The present invention obtains the speech information in the laboratory through the speech acquisition module, transmits it to the server management module through the communication module. The server management module first separates the speakers of the speech information through the spatial clustering algorithm, and performs annotation according to the speaker clustering segmentation algorithm, then extracts the features related to the risk, combines with the set risk information database to obtain the risk level, and finally outputs the relevant information. The present invention separates the signals, can identify the speech information of different people in the laboratory, and then improves the overall accuracy of speech recognition.

[0032] 2. The present invention is provided with a natural language processing sub-module, which extracts the Mel cepstrum coefficient features of the voice information. This feature reflects the characteristics of the audio in a way that simulates the human ear, further ensuring the accuracy of the voice information during calculation. After converting this feature into text information, features are extracted through the BERT model and compared with the risk information database, enabling effective identification of potential risk information.

[0033] 3. The present invention adopts a multi-objective learning algorithm based on the long short-term memory model. The algorithm objective is optimized according to the logarithmic power spectrum feature and the ideal ratio mask, enhancing the signal of the voice information, achieving a lower computational complexity, and improving the generalization ability of the system.

[0034] 4. In addition to being applicable to laboratories, the present invention can be applied to other scenarios by changing the specific content of the risk information database, with a wide range of applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of the system of the present invention.

[0036] Figure 2 It is a schematic diagram of the blind source separation framework of the present invention.

[0037] Figure 3 It is a framework diagram for the multi-person speech recognition sub-module and the natural language processing sub-module of the present invention to realize text information recognition.

[0038] Figure 4 It is a flow chart for updating the risk information database of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0040] This embodiment provides a laboratory voice monitoring system based on speech recognition and natural language processing, as Figure 1 shown, including a voice collection module, a communication module, and a server management module. The voice collection module is connected to the server management module through the communication module:

[0041] The voice collection module is used to collect audio information in the laboratory and separate the voice information based on the optimization of the auxiliary function;

[0042] The communication module is used to transmit the voice information to the server management module;

[0043] The server management module is used to separate the voice information into multiple single-person voice information according to the spatial clustering algorithm, annotate each single-person voice information in the time dimension according to the speaker clustering and segmentation algorithm, extract features from the annotated voice information, combine the features with the risk information database to obtain the risk level corresponding to the voice information, and output the risk level, voice information and annotation information.

[0044] Among them, the voice acquisition module specifically includes a microphone array sub-module and a separation sub-module:

[0045] The microphone array sub-module includes multiple omnidirectional microphones set at different spatial positions in the laboratory, arranged in a certain shape rule to form a microphone array, which is used to collect audio information containing spatial position information. The microphone outputs a PDM signal converted from an analog signal.

[0046] The separation sub-module separates the audio according to the AIVA algorithm based on blind source separation (as Figure 2 shown). The AIVA algorithm uses an optimization method of auxiliary function approximation to replace the gradient descent process. It is an iterative algorithm used for maximum likelihood estimation or maximum a posteriori probability estimation of the parameters of a probability model containing hidden variables. Its objective function J(W) is expressed in the form of KL divergence:

[0047]

[0048] In the formula, J represents negative entropy, W represents the separation matrix, y t is the estimated source vector at time t, N is the number of channels, p is the distribution density function with frequency as the random variable, and q is the joint distribution density function.

[0049] Substituting into the KL divergence formula, the above formula can be simplified and expressed as:

[0050]

[0051] In the formula, E t is the expected value at time t, and det(·) is the determinant of the square matrix.

[0052] By using the optimization method of the auxiliary function, the form of the auxiliary function G(W,V) of the objective function is determined as follows:

[0053]

[0054] In the formula, V is the weighted variance matrix, ω is the angular frequency, τ is the window function, k is the weight value for updating each point of the separation matrix, and K is the number of frequency points.

[0055] Finally, the audio information is separated into voice information and noise information.

[0056] After obtaining the voice information, according to its different spatial positions, it is input into the server management module through multiple channels of the communication module. The communication module uses a half-duplex network composed of RS485 interfaces to communicate with the server management module.

[0057] The server management module includes a signal reception queue sub-module, a multi-person speech recognition sub-module, a natural language processing sub-module, a query analysis sub-module, an alarm sub-module, and an auxiliary sub-module:

[0058] The signal reception queue sub-module is used to receive voice information and transmit it to the multi-person speech recognition sub-module;

[0059] The multi-person speech recognition sub-module includes a pre-processor and a speech recognizer. The pre-processor is used to enhance the signal of the voice information, and the speech recognizer is used to separate the voice information after signal enhancement by using deep learning combined with a spatial clustering algorithm based on a Gaussian mixture model to obtain multiple single-person voice information; after speaker segmentation, the signal turning points of the voice information are cut, and finally the cut voice information is labeled by speaker clustering.

[0060] Among them, the steps executed by the pre-processor are specifically as follows:

[0061] Performing speech enhancement on the voice information, also known as speech noise reduction, refers to suppressing noise interference in the mixed observed speech signal. In the present invention, the existing model structure is improved based on a long short-term memory model to improve the learning ability and generalization ability of the model. At the same time, by designing different objective functions, multiple models are fused to find a balance between removing noise and maintaining the listening feeling. A multi-objective learning algorithm is introduced, and at the output end, the logarithmic power spectrum feature and the ideal ratio mask are used simultaneously to guide the optimization of the speech enhancement model. Its objective function T MTL The expression is as follows:

[0062]

[0063] In the formula, represents the logarithmic power spectrum (LPS) feature of the input mixed noisy speech, is the LPS feature of the reference noise-free speech, is the time-frequency mask, (m,n) is the time-frequency unit, m is the index of the time frame, n is the index of the frequency point, and λ is the weight for adjusting the ratio between the two learning objectives.

[0064] Introducing multi-task learning to establish an LSTM model, and obtaining approximate clean speech features, including LPS features and IRM features, by sharing weight parameters before output to obtain a smaller model size and achieve lower computational complexity. At the same time, adding multiple regularization terms to avoid overfitting and improve the robustness and generalization ability of the model.

[0065] The speech recognizer specifically performs the following steps:

[0066] Step A1: Use the generalized weighted prediction error algorithm to remove reverberation.

[0067] Step A2: Adopt the blind separation algorithm based on AIVA to obtain multiple independent speech information, and restore the amplitude to the range of the original observed signal through the back-projection algorithm.

[0068] Step A3: Use the noise estimation algorithm of the recursive average algorithm controlled by the minimum value to estimate the stationary noise on the current speech information after amplitude recovery, and then combine the stationary noise information of other speech information to calculate the non-stationary noise of the current speech information.

[0069] Step A4: According to the stationary noise and non-stationary noise, use the decision-directed algorithm to calculate the priori and posteriori signal-to-noise ratios and perform multi-channel filtering processing.

[0070] Step A5: Multiple processed single-channel speeches are synthesized into one speech by linear superposition.

[0071] Step A6: For the target speaker, superimpose the speeches of other speakers on the target person according to different signal-to-noise ratios for data clarification and enhancement, and the objective function T IM The expression is as follows:

[0072]

[0073] In the formula, x LPS (m,n) is the logarithmic power spectrum of the noisy speech.

[0074] Step A7: Use the standard generalized eigenvalue decomposition algorithm to process the data obtained in Step A6. Utilize deep learning combined with the spatial clustering algorithm based on the Gaussian mixture model to perform clustering without prior knowledge. After the algorithm converges, the GEVD beam uses the time-frequency mask of the speaker to generate the finally separated single-speaker speech information.

[0075] Step A8: Use the speaker segmentation algorithm to cut at the turning points of the audio, and divide the entire audio into multiple small segments, each of which contains the speech of only one speaker.

[0076] Step A9: Use the speaker clustering algorithm, preferably the agglomerative hierarchical clustering algorithm in this embodiment, to cluster all the segmented segments. The purpose of clustering is to re-divide the multiple scattered segments according to the speaker identity, so that each clustering cluster only contains one speaker. When the total number of clustering clusters of the speaker clustering reaches the default maximum number of speakers, or the clustering value exceeds the set threshold, stop the speaker clustering and output the clustering result as the annotation information.

[0077] The natural language processing sub-module uses an acoustic model to convert the speech information that has passed through the multi-person speech recognition sub-module into text information, and uses the BERT model to extract the features of the text information, and matches the features with the risk information database corresponding to the current laboratory to obtain the risk level.

[0078] Using Mel-frequency cepstral coefficients, the sound is converted into a matrix of A rows and B columns, called the observation sequence, where A is the acoustic feature dimension and B is the total number of frames. The decoding process recognizes the frames into states, then combines them into phonemes, and finally synthesizes the text information. The flow chart of the above process from inputting the original speech information to outputting the text information is as Figure 3 shown. The training stage can be summarized as a signal processing model for obtaining extractable features, and the testing stage can be summarized as converting audio information into text information.

[0079] The BERT model is used to extract the features of the text information, encode the sentences of the new task, and encode sentences of any length into fixed-length vectors. That is, the hidden layer states of one or more layers of the input sequence are concatenated as the representation of the sequence and the input of the downstream model algorithm. This method has no backpropagation process.

[0080] After extracting the text features, determine the risk information database corresponding to the current laboratory scenario, make a judgment according to the pre-established risk information database and risk level standard, and output the corresponding risk level. The risk information database is a database manually collected and maintained and listing relevant risk event information. The update method is as Figure 3 shown. Once new information is collected, the risk database needs to be updated. First, the risk information is input into the deep learning model for training. After obtaining the word vectors, the trained deep learning model is input to output the judgment result, and the qualified word vectors are expanded into the risk database. For example, "fire", "smoke", etc. are first-level laboratory accident risks.

[0081] After obtaining the risk level, the alarm sub-module determines whether the risk level exceeds the threshold. If so, it alarms the administrator by text message or email.

[0082] The query and analysis sub-module outputs the speech information, annotation information, and risk level.

[0083] The auxiliary sub-module mainly includes system auxiliary functions such as user permission management, risk threshold setting, dictionary setting, and log management.

[0084] The working process of this embodiment is as follows:

[0085] First, the microphone array sub-module collects audio information in the laboratory, including noise information and speech information of multiple people speaking. Then, the separation sub-module separates the noise information and speech information, retains only the speech information, and transmits it through the communication module. It is received by the signal receiving queue sub-module and transmitted to the multi-person speech recognition sub-module. First, preprocessing is performed. The speech information is enhanced by the pre-processor, and then through the speech recognizer, the speech information is segmented into the speech information of each individual speaker and labeled. Then, through the natural speech processing sub-module, the speech information is converted into text information. Combining with the risk information database and the bidirectional neural network, the risk level is obtained. Finally, the risk level, speech information, and annotation information are output through the query analysis sub-module, and an alarm is issued through the alarm sub-module. And the system is maintained and managed regularly through the auxiliary sub-module.

[0086] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A laboratory voice monitoring system based on speech recognition and natural language processing, characterized in that, It includes a voice acquisition module, a communication module, and a server management module. The voice acquisition module is connected to the server management module through the communication module: The voice acquisition module is used to collect audio information in the laboratory and optimize and separate the voice information based on an auxiliary function; The communication module is used to transmit the voice information to the server management module; The server management module is used to separate the voice information into multiple single-person voice information according to the spatial clustering algorithm, annotate each single-person voice information in the time dimension according to the speaker clustering segmentation algorithm, extract features from the annotated voice information, combine the features with the risk information database to obtain the risk level corresponding to the voice information, and output the risk level, voice information, and annotation information; The server management module includes a signal reception queue sub-module, a multi-person speech recognition sub-module, a natural language processing sub-module, and a query analysis sub-module: The signal reception queue sub-module is used to receive the voice information and transmit it to the multi-person speech recognition sub-module; The multi-person speech recognition sub-module includes a pre-processor and a speech recognizer. The pre-processor is used to enhance the signal of the voice information, and the speech recognizer is used to separate the signal-enhanced voice information into multiple single-person voice information by using deep learning combined with a spatial clustering algorithm based on a Gaussian mixture model; perform cutting on the signal turning points of the voice information through speaker segmentation, and finally perform annotation on the cut voice information through speaker clustering; The natural language processing sub-module uses an acoustic model to convert the voice information passing through the multi-person speech recognition sub-module into text information, extracts features of the text information by using a BERT model, and matches the features with the risk information database corresponding to the current laboratory to obtain the risk level; The query analysis sub-module outputs the voice information, annotation information, and risk level; Before separating the voice signal, the speech recognizer performs the following steps: Step A1: Use a generalized weighted prediction error algorithm to remove reverberation; Step A2: Adopt a blind separation algorithm based on Independent Vector Analysis (IVA), take multiple mutually independent voice information, and restore the amplitude to the range of the original observed signal through the back-projection algorithm; Step A3: Use a noise estimation algorithm of a recursive average algorithm controlled by a minimum value to estimate the stationary noise on each individual voice information, and estimate the non-stationary noise of the current voice information by using the stationary noise information of other voice information; Step A4: Use a decision-directed algorithm to calculate the priori and posteriori signal-to-noise ratios and perform multi-signal filtering processing; Step A5: Multiple processed single-channel voices are synthesized into one voice by linear superposition.

2. The laboratory voice monitoring system based on speech recognition and natural language processing according to claim 1, characterized in that, The voice acquisition module includes a microphone array sub-module and a separation sub-module: The microphone array sub-module includes multiple omnidirectional microphones arranged in the laboratory, which are used to collect audio information containing spatial position information; The separation sub-module separates the audio information into voice information and noise information according to the independent vector analysis algorithm optimized based on an auxiliary function, combined with the spatial position information of the audio information.

3. The laboratory voice monitoring system based on speech recognition and natural language processing according to claim 1, characterized in that, The preprocessor adopts a multi-objective learning algorithm based on the long short-term memory model, optimizes the algorithm objectives according to the logarithmic power spectrum features and the ideal ratio mask, and enhances the signal of the speech information.

4. The laboratory voice monitoring system based on speech recognition and natural language processing according to claim 1, characterized in that, The acoustic model extracts the mel cepstrum features of the speech information and outputs the text after decoding the features.

5. The laboratory voice monitoring system based on speech recognition and natural language processing according to claim 1, characterized in that, Before separating the speech signals, the speech recognizer eliminates the reverberation of the speech information based on the generalized weighted prediction error algorithm.

6. The laboratory voice monitoring system based on speech recognition and natural language processing according to claim 1, characterized in that, When the total number of clusters in the speaker clustering reaches the default maximum number of speakers, or the clustering value exceeds the set threshold, the speaker clustering is stopped, and the clustering result is output as the annotation information.

7. The laboratory voice monitoring system based on speech recognition and natural language processing according to claim 1, characterized in that, The server management module further includes an alarm sub-module. After obtaining the risk level, it judges whether the risk level exceeds the threshold. If so, it alarms the administrator.

8. The laboratory voice monitoring system based on speech recognition and natural language processing according to claim 1, characterized in that, The communication module adopts a half-duplex network composed of RS485 interfaces.

Citation Information

Patent Citations

  • Speaker marking method and system based on density peak value clustering and variational Bayes

    CN106971713A

  • Safety monitoring method for shop

    CN109300279A