Meeting minutes generation method, device, electronic device and storage medium
By extracting the spectral diagram of conference voice data and using the intelligent decoding engine to perform voice recognition and error correction operations, the problem of low speech recognition accuracy in complex conference scenarios is solved, and efficient and accurate conference minutes are achieved.
Patent Information
- Application Number
- CN202111358381.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-16
AI Technical Summary
The prior art is difficult to quickly and accurately identify user speech information in complex conference scenarios, resulting in low accuracy of the results of the conference minutes.
By extracting the spectral diagram of conference voice data, using the acoustic model and language model of the preset intelligent decoding engine, the first probability value between the signal characteristics and the phoneme template and the second probability value between the phoneme characteristics and the text template are determined, end-to-end speech recognition is performed, and the generated text data is corrected to generate meeting minutes.
It improves the efficiency and accuracy of speech recognition in complex scenarios, ensuring the accuracy and reliability of meeting minutes.
Smart Images

Figure CN114203180B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of conference record, and in particular to a method, device, electronic device and storage medium for generating conference minutes. Background Art
[0002] When holding a meeting, the traditional method is to manually record the content of the meeting process and manually organize it into meeting minutes, but the manual method is inefficient. In order to improve the recording efficiency, speech recognition technology is applied to meeting records to achieve intelligent generation of meeting minutes.
[0003] However, meeting records are real-time and complex. Different people say the same thing, or the same person says the same thing at different times, in different physiological and psychological states, which can be very different. In the continuous speech of users, there are no obvious boundaries between phonemes, syllables and words, and each pronunciation unit has a co-articulation phenomenon that is strongly affected by the context. The current speech recognition models are all targeted at specific users or specific scenarios, and it is difficult to quickly and accurately recognize user speech information in the complex scenarios of meeting records. Summary of the invention
[0004] The present application provides a method, device, electronic device and storage medium for generating meeting minutes, so as to solve the technical problem of low accuracy of the generated results of meeting minutes.
[0005] In order to solve the above technical problems, in a first aspect, an embodiment of the present application provides a method for generating meeting minutes, comprising:
[0006] Extract spectrogram of conference speech data;
[0007] Using the acoustic model of the preset intelligent decoding engine, according to the spectrogram, a first probability value between the signal feature of the conference speech data and the phoneme template is determined to obtain the phoneme feature corresponding to the signal feature;
[0008] Determine a second probability value between the phoneme feature and the text template using a language model of a preset intelligent decoding engine;
[0009] Using a decoder of a preset intelligent decoding engine, according to the first probability value and the second probability value, the conference voice data is decoded to obtain conference text data;
[0010] Perform error correction operations on meeting text data and generate meeting minutes.
[0011] This embodiment extracts the spectrogram of the conference voice data to process the voice data within a period of time, thereby achieving the purpose of continuous voice processing; uses the acoustic model of the preset intelligent decoding engine to determine the first probability value between the signal feature of the conference voice data and the phoneme template according to the spectrogram, and obtains the phoneme feature corresponding to the signal feature; and uses the language model of the preset intelligent decoding engine to determine the second probability value between the phoneme feature and the text template; and uses the decoder of the preset intelligent decoding engine to decode the conference voice data according to the first probability value and the second probability value to obtain conference text data, so as to achieve end-to-end voice recognition without directly extracting voice features, thereby improving the efficiency and accuracy of voice recognition in complex scenarios; finally, error correction operations are performed on the conference text data to generate meeting minutes, and further ensure the accuracy of the final result.
[0012] In one embodiment, extracting a spectrogram of conference speech data includes:
[0013] Divide the conference voice data into frames to obtain multiple frames of voice signals;
[0014] Performing windowing processing on each frame of speech signal, and performing fast Fourier transform on the windowed speech signal to obtain the frequency spectrum of each frame of speech signal;
[0015] Multiple spectra are superimposed to obtain a spectrogram.
[0016] This embodiment performs framing, windowing and fast Fourier transformation on the conference voice data to convert the conference voice data from a time domain signal to a frequency domain signal, so as to better characterize the signal characteristics of the conference voice data.
[0017] In one embodiment, the acoustic model of the preset intelligent decoding engine is used to determine the first probability value between the signal feature of the conference speech data and the phoneme template according to the spectrogram, and the phoneme feature corresponding to the signal feature is obtained, including:
[0018] Using the acoustic model, calculating a first probability value between a signal feature of the spectrogram and a phoneme template in a preset language library, wherein the preset language library includes phoneme templates of a small vocabulary;
[0019] The phoneme template having the largest first probability value with the signal feature is determined as the phoneme feature.
[0020] This embodiment performs probability calculation through phoneme templates of a preset language library so that speech recognition can support small vocabulary and dialect recognition and has language recognition diversification.
[0021] In one embodiment, the language model is a trigram model, and the second probability value between the phoneme feature and the text template is determined by using the language model of the preset intelligent decoding engine, including:
[0022] The trigram model is used to calculate a second probability value between the phoneme feature and a text template in a preset text library.
[0023] This embodiment calculates the second probability value through the ternary model, which can avoid the data sparsity problem and thus improve the accuracy of the calculation result.
[0024] In one embodiment, the conference voice data is decoded by using a decoder of a preset intelligent decoding engine according to the first probability value and the second probability value to obtain the conference text data, including:
[0025] The conference voice data is decoded using the decoding function of the decoder according to the first probability value and the second probability value to obtain the conference text data. The decoding function is:
[0026] w * = argmax w (logP(w|o)+λlogP LM (w)+γlen(w));
[0027] Where P(ω|o) is the first probability value, P LM (ω) is the second probability value, and γlen(ω) is the length of the conference voice data.
[0028] This embodiment directly outputs the language recognition result through the decoding function of the decoder, realizes an end-to-end language recognition process, and improves the language recognition efficiency and recognition accuracy.
[0029] In one embodiment, error correction operations are performed on conference text data to generate conference minutes, including:
[0030] Perform word segmentation on the conference text data to obtain candidate error results;
[0031] Perform word replacement operations on candidate error results to generate meeting minutes.
[0032] Through error correction operations, this embodiment can effectively recognize Mandarin with a slight accent, Cantonese, Sichuan dialects, and foreign languages such as English, and can automatically correct errors based on sentence meanings, automatically segment words and sentences, and add punctuation, making input faster and communication smoother.
[0033] In one embodiment, before extracting the spectrogram of the conference speech data, the method further includes:
[0034] Collect conference voice data;
[0035] Perform voiceprint recognition on the conference voice data to determine the speaker corresponding to the conference voice data.
[0036] This embodiment uses voiceprint recognition to accurately record the speech content of each speaker and ensure the accuracy of the recorded information.
[0037] In a second aspect, an embodiment of the present application provides a device for generating meeting minutes, including:
[0038] An extraction module, used to extract the spectrogram of conference speech data;
[0039] A first determination module is used to determine a first probability value between a signal feature of the conference speech data and a phoneme template according to a spectrogram by using an acoustic model of a preset intelligent decoding engine, and obtain a phoneme feature corresponding to the signal feature;
[0040] A second determination module, used to determine a second probability value between the phoneme feature and the text template by using a language model of a preset intelligent decoding engine;
[0041] A decoding module, used to decode the conference voice data using a decoder of a preset intelligent decoding engine according to the first probability value and the second probability value to obtain conference text data;
[0042] The error correction module is used to perform error correction operations on the conference text data and generate meeting minutes.
[0043] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method for generating meeting minutes as described in the first aspect is implemented.
[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method for generating meeting minutes as described in the first aspect.
[0045] It should be noted that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A flowchart of a method for generating meeting minutes provided in an embodiment of the present application;
[0047] Figure 2 A schematic diagram of the structure of a device for generating meeting minutes provided in an embodiment of the present application;
[0048] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0050] As documented in related technologies, conference records are real-time and complex. Different people say the same thing, or the same person says the same thing at different times, in different physiological and psychological states, which can be very different. In continuous speech, there are no obvious boundaries between phonemes, syllables, and words, and each pronunciation unit has a co-articulation phenomenon that is strongly affected by the context. However, current speech recognition models are all targeted at specific users or specific scenarios, and it is difficult to quickly and accurately recognize user speech information in complex scenarios such as conference records.
[0051] To this end, the embodiments of the present application provide a method, device, electronic device and storage medium for generating meeting minutes, which extracts the spectrogram of the conference voice data to process the voice data within a period of time, thereby achieving the purpose of continuous voice processing; utilizes the acoustic model of a preset intelligent decoding engine to determine the first probability value between the signal feature of the conference voice data and the phoneme template according to the spectrogram, and obtains the phoneme feature corresponding to the signal feature, and utilizes the language model of the preset intelligent decoding engine to determine the second probability value between the phoneme feature and the text template, and utilizes the decoder of the preset intelligent decoding engine to decode the conference voice data according to the first probability value and the second probability value to obtain conference text data, so as to realize end-to-end voice recognition without directly extracting voice features, thereby improving the efficiency and accuracy of voice recognition; finally, error correction operations are performed on the conference text data to generate meeting minutes, further ensuring the accuracy of the final result.
[0052] Please refer to Figure 1 , Figure 1 The flowchart of a method for generating meeting minutes provided in an embodiment of the present application is shown in FIG. The method for generating meeting minutes in an embodiment of the present application can be applied to electronic devices, including but not limited to smart phones, tablet computers, laptop computers, and personal digital assistants. Figure 1 As shown, the method for generating meeting minutes includes steps S101 to S105, which are described in detail as follows:
[0053] Step S101, extracting the spectrogram of conference speech data.
[0054] In this step, the spectrogram is formed by superimposing the frequency spectrograms within a period of time. The main steps of optionally extracting the spectrogram include framing, windowing and fast Fourier transforming the conference speech data.
[0055] Step S102, using the acoustic model of the preset intelligent decoding engine, according to the spectrogram, determine the first probability value between the signal feature of the conference voice data and the phoneme template, and obtain the phoneme feature corresponding to the signal feature.
[0056] In this step, the preset decoding engine includes an acoustic model, a language model and a decoder. The acoustic model is mainly used to calculate the likelihood (first probability value) between the speech signal feature and each pronunciation template (phoneme template).
[0057] Optionally, the acoustic model is used to calculate a first probability value between the signal feature of the spectrogram and a phoneme template in a preset language library, wherein the preset language library includes phoneme templates of a small vocabulary; and the phoneme template having the largest first probability value with the signal feature is determined as the phoneme feature.
[0058] In this embodiment, the training data is processed by a convolutional neural network, the main features are extracted by maximum pooling, and the acoustic model is obtained by training according to the CTC loss function. For example, a model is created for a new voice or dialect and for the application of a small vocabulary; sufficient voice data is collected, for example, the length of the voice data of a single person is at least 1 hour, and the length of the voice data of 200 people is at least 50 hours; the training data is processed by a convolutional neural network for training and optimization.
[0059] Step S103, using the language model of the preset intelligent decoding engine to determine a second probability value between the phoneme feature and the text template.
[0060] In this step, the language model can ensure the grammatical structure of the text, so that the recognized sentences are fluent. The language model is a probabilistic statistical method that uses a trained language model to give a probability to any text. The higher the probability, the more grammatically fluent it is. By training a language model and comparing the probabilities of two sentences on the same language model to determine the accuracy of the grammar and the fluency of the sentence, it can reduce labor costs.
[0061] Step S104: using the decoder of the preset intelligent decoding engine, decode the conference voice data according to the first probability value and the second probability value to obtain conference text data.
[0062] In this step, optionally, the conference voice data is decoded according to the first probability value and the second probability value using a decoding function of the decoder to obtain conference text data, and the decoding function is:
[0063] w * = argmax w(logP(w|o)+λlogP LM (w)+γlen(w));
[0064] Where P(ω|o) is the first probability value, P LM (ω) is the second probability value, γlen(ω) is the length of the conference speech data. λ is the weight of the language model. The larger the λ is, the more dependent on the language model it is. Traverse all possible word sequences to find the position with the highest probability and output the result.
[0065] Step S105, performing error correction operations on the conference text data to generate conference minutes.
[0066] In this step, the error correction operation includes identifying typos, spelling errors, grammatical errors and common format errors, and returning modification suggestions. After confirming the modification suggestions, the typos and other errors are corrected and transmitted to the meeting minutes document function to generate meeting minutes.
[0067] Optionally, the automatically generated meeting minutes document can be downloaded to improve the efficiency of meeting recording.
[0068] It should be noted that this embodiment can recognize the audio stream as text in real time and return the start and end time of each sentence. It is suitable for scenarios such as long sentence voice input, audio and video subtitles, and conferences. It supports WebSocket API, Android, iOS, and Linux SDK, and can be called on multiple operating systems and multiple device terminals. It is used for audio file transcription, recognizes batch uploaded audio files as text, supports Mandarin and Chinese with a slight accent, and supports English recognition. It is suitable for scenarios such as recording quality inspection, meeting content summary, and audio content analysis.
[0069] In one embodiment, Figure 1 Based on the illustrated embodiment, step S101 includes:
[0070] Framing the conference voice data to obtain multiple frames of voice signals;
[0071] Performing windowing processing on the speech signal of each frame, and performing fast Fourier transform on the speech signal after the windowing processing to obtain the frequency spectrum of the speech signal of each frame;
[0072] The spectrogram is obtained by superimposing a plurality of the frequency spectra.
[0073] In this embodiment, the conference voice data is a voice signal. The voice signal is divided into frames and then windowed when processing the voice signal. That is, the conference voice data in a frame is multiplied by a window function to obtain a new frame of data. Each time a section of data is taken, the data is fast Fourier transformed and analyzed, and then the next section of data is taken and analyzed again.
[0074] Since it is difficult to see the characteristics of speech signals in the time domain, this embodiment performs a fast Fourier transform on each frame of the signal processed by the window function to convert the time domain image into a spectrum image of each frame, and then superimposes the spectrum of each window to obtain a spectrogram.
[0075] As you can understand, Fourier transform is a method of analyzing signals. It can analyze the components of a signal and also use these components to synthesize a signal. Many waveforms can be used as components of a signal, such as sine waves, square waves, sawtooth waves, etc. Fourier transform uses sine waves as components of a signal.
[0076] Any periodic function can be represented by an infinite series of sine and cosine functions, which is called the Fourier series. If there is a periodic function with a complex waveform, then the method of finding the frequencies of the sine and cosine functions that can be used to form this periodic function is called the Fourier transform, and the method of representing this periodic function by superimposing the sine and cosine functions of these frequencies is called the inverse Fourier transform.
[0077] In one embodiment, Figure 1 Based on the illustrated embodiment, step S103 includes:
[0078] The trigram model is used to calculate a second probability value between the phoneme feature and a text template in a preset text library.
[0079] In this embodiment, the language model is the probability distribution of a sequence of words. Specifically, the language model determines a probability distribution P for a text of length m, which represents the possibility of the existence of this text. In practice, if the length of the text is long, the calculation of P(wi|w1, w2, ..., wi-1) will be very difficult. Therefore, this embodiment uses the model to simplify it into an n-gram model, in which only the first n words of the current word need to be calculated when estimating the conditional probability in the n-gram model. Traditional methods generally use the ratio of frequency counts to estimate the n-gram conditional probability. When n is large, there is a data sparsity problem, resulting in inaccurate estimation results. Therefore, this embodiment adopts a ternary model to be able to cope with probability calculations at the million-word level.
[0080] In one embodiment, Figure 1 Based on the embodiment shown, step S105 includes:
[0081] Performing a word segmentation operation on the conference text data to obtain a candidate error result;
[0082] A word replacement operation is performed on the candidate error results to generate the meeting minutes.
[0083] In the present embodiment, the error correction operation includes error detection and error correction; wherein the error detection part first segments words (i.e., segmentation) through the Jieba Chinese word segmenter. Since the sentence contains typos, the segmentation result often has segmentation errors, so that errors can be detected from both the character granularity and the word granularity, and the suspected error results of these two granularities are integrated to form a candidate set of suspected error positions (i.e., candidate error results). The error correction part is based on the candidate set of suspected error positions, by traversing all suspected error positions, and using sound-like words and shape-like words to replace the words in the wrong position, and then calculating the sentence perplexity through the language model, comparing and sorting all candidate set results, and obtaining the optimal corrected word. The present embodiment can greatly reduce the occurrence of typos and wrong words, and avoids the problem of missed inspections.
[0084] In one embodiment, Figure 1 Based on the embodiment shown, before step S101, the following step is further included:
[0085] Collecting the conference voice data;
[0086] Perform voiceprint recognition on the conference voice data to determine the speaker corresponding to the conference voice data.
[0087] In this embodiment, the intelligent recording function is used to record system sound, microphone sound or both at the same time, and supports saving audio resources, voice dubbing, recording meeting minutes or transcribing audio formats. The data format of the conference voice data can include but is not limited to MP3, AAC, OGG, WMA, WAV or FLAC, so as to be suitable for uploading to major platforms and support audio sharing.
[0088] Through voiceprint recognition, the conference voice data is generated into a feature vector, and the voiceprint feature vector that has been previously entered into the feature vector database is compared to identify whether it is the same speaker. If it is not the same speaker, it is treated as a new speaker and the speech information is recorded; if it is the same speaker, the speech information is recorded as the same speaker. This embodiment can distinguish the voices of different speakers at the meeting, and can classify and record them well, making the meeting content more specific and rich, and avoiding unclear recorded information.
[0089] In order to execute the method for generating meeting minutes corresponding to the above method embodiment, to achieve the corresponding functions and technical effects. Figure 2 , Figure 2The structural block diagram of a device for generating conference minutes provided in an embodiment of the present application is shown. For the sake of convenience, only the parts related to the present embodiment are shown. The device for generating conference minutes provided in an embodiment of the present application includes:
[0090] Extraction module 201, used for extracting the spectrogram of conference speech data;
[0091] A first determination module 202 is used to determine a first probability value between a signal feature of the conference speech data and a phoneme template according to the spectrogram by using an acoustic model of a preset intelligent decoding engine, and obtain a phoneme feature corresponding to the signal feature;
[0092] A second determination module 203 is used to determine a second probability value between the phoneme feature and the text template by using the language model of the preset intelligent decoding engine;
[0093] A decoding module 204, configured to use a decoder of the preset intelligent decoding engine to decode the conference voice data according to the first probability value and the second probability value to obtain conference text data;
[0094] The error correction module 205 is used to perform error correction operations on the conference text data to generate meeting minutes.
[0095] In one embodiment, the extraction module 201 includes:
[0096] A framing unit, used for framing the conference voice data to obtain multiple frames of voice signals;
[0097] A windowing unit, used for performing windowing processing on the speech signal of each frame, and performing fast Fourier transform on the speech signal after the windowing processing to obtain the frequency spectrum of the speech signal of each frame;
[0098] The superposition unit is used to superimpose a plurality of the frequency spectra to obtain the spectrogram.
[0099] In one embodiment, the first determining module 202 includes:
[0100] A first calculation unit, configured to calculate, by using the acoustic model, a first probability value between a signal feature of the spectrogram and a phoneme template in a preset language library, wherein the preset language library includes phoneme templates of a small vocabulary;
[0101] The determination unit is used to determine the phoneme template with the largest first probability value with the signal feature as the phoneme feature.
[0102] In one embodiment, the second determining module 203 includes:
[0103] The second calculation unit is used to calculate a second probability value between the phoneme feature and a text template in a preset text library by using the trigram model.
[0104] In one embodiment, the decoding module 204 includes:
[0105] A decoding unit is used to decode the conference voice data according to the first probability value and the second probability value using a decoding function of the decoder to obtain conference text data, wherein the decoding function is:
[0106] w * = argmax w (logP(w|o)+λlogP LM (w)+γlen(w));
[0107] Where P(ω|o) is the first probability value, P LM (ω) is the second probability value.
[0108] In one embodiment, the error correction module 205 includes:
[0109] A word segmentation unit, used for performing a word segmentation operation on the conference text data to obtain a candidate error result;
[0110] A replacement unit is used to perform a word replacement operation on the candidate error results to generate the meeting minutes.
[0111] In one embodiment, the generating device further includes:
[0112] A collection module, used for collecting the conference voice data;
[0113] The third determination module is used to perform voiceprint recognition on the conference voice data to determine the speaker corresponding to the conference voice data.
[0114] The above-mentioned device for generating meeting minutes can implement the method for generating meeting minutes of the above-mentioned method embodiment. The optional items in the above-mentioned method embodiment are also applicable to this embodiment and will not be described in detail here. The rest of the contents of the embodiment of this application can refer to the contents of the above-mentioned method embodiment and will not be described in detail in this embodiment.
[0115] Figure 3 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment includes: at least one processor 30 ( Figure 3 Only one is shown in the figure) a processor, a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30, and when the processor 30 executes the computer program 32, the steps in any of the above method embodiments are implemented.
[0116] The electronic device 3 may be a computing device such as a smart phone, a tablet computer, a desktop computer, etc. The electronic device may include but is not limited to a processor 30 and a memory 31. Those skilled in the art will understand that Figure 3 It is only an example of the electronic device 3 and does not constitute a limitation on the electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0117] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0118] In some embodiments, the memory 31 may be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. In other embodiments, the memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 3. Further, the memory 31 may also include both an internal storage unit and an external storage device of the electronic device 3. The memory 31 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0119] In addition, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0120] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device implements the steps in the above-mentioned method embodiments when executing the computer program product.
[0121] In several embodiments provided in the present application, it is understood that each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
[0122] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can essentially be embodied in the form of a software product, or in other words, the part that contributes to the prior art or the part of the technical solution. The computer software product is stored in a storage medium, including several instructions for enabling a terminal device to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0123] The specific embodiments described above further describe the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of protection of the present application. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of protection of the present application.
Claims
1. A method for generating meeting minutes. It is characterized in that include: Extract spectrogram of conference speech data; Using the acoustic model of the preset intelligent decoding engine, according to the spectrogram, determine a first probability value between the signal feature of the conference speech data and the phoneme template, and obtain the phoneme feature corresponding to the signal feature; Determine a second probability value between the phoneme feature and the text template by using the language model of the preset intelligent decoding engine; Using the decoder of the preset intelligent decoding engine, the conference voice data is decoded according to the first probability value and the second probability value to obtain conference text data; Performing error correction operations on the conference text data to generate meeting minutes; The method of using the decoder of the preset intelligent decoding engine to decode the conference voice data according to the first probability value and the second probability value to obtain conference text data includes: The conference voice data is decoded by using a decoding function of the decoder according to the first probability value and the second probability value to obtain conference text data, wherein the decoding function is: oh * =argmax ω (logP(ω|o)+λlogP LM (ω)+γlen(ω)); Where P(ω|o) is the first probability value, λ is the weight of the language model, and P LM (ω) is the second probability value, and γlen(ω) is the length of the conference voice data.
2. The method for generating meeting minutes as claimed in claim 1, It is characterized in that The step of extracting a spectrogram of conference speech data comprises: Framing the conference voice data to obtain multiple frames of voice signals; Performing windowing processing on the speech signal of each frame, and performing fast Fourier transform on the speech signal after the windowing processing to obtain the frequency spectrum of the speech signal of each frame; The spectrogram is obtained by superimposing a plurality of the frequency spectra.
3. The method for generating meeting minutes as claimed in claim 1, It is characterized in that The method of using the acoustic model of the preset intelligent decoding engine to determine the first probability value between the signal feature of the conference voice data and the phoneme template according to the spectrogram, and obtaining the phoneme feature corresponding to the signal feature, includes: Using the acoustic model, calculating a first probability value between the signal feature of the spectrogram and a phoneme template in a preset language library, wherein the preset language library includes phoneme templates of a small vocabulary; The phoneme template having the largest first probability value with the signal feature is determined as the phoneme feature.
4. The method for generating meeting minutes as claimed in claim 1, It is characterized in that The language model is a ternary model, and the method of using the language model of the preset intelligent decoding engine to determine the second probability value between the phoneme feature and the text template includes: The trigram model is used to calculate a second probability value between the phoneme feature and a text template in a preset text library.
5. The method for generating meeting minutes as claimed in claim 1, It is characterized in that The error correction operation is performed on the conference text data to generate the conference minutes, including: Performing a word segmentation operation on the conference text data to obtain a candidate error result; A word replacement operation is performed on the candidate error results to generate the meeting minutes.
6. The method for generating meeting minutes as claimed in claim 1, It is characterized in that Before extracting the spectrogram of the conference speech data, the method further includes: Collecting the conference voice data; Perform voiceprint recognition on the conference voice data to determine the speaker corresponding to the conference voice data.
7. A device for generating meeting minutes, It is characterized in that include: An extraction module, used to extract the spectrogram of conference speech data; A first determination module is used to determine a first probability value between a signal feature of the conference speech data and a phoneme template according to the spectrogram by using an acoustic model of a preset intelligent decoding engine, and obtain a phoneme feature corresponding to the signal feature; A second determination module, configured to determine a second probability value between the phoneme feature and the text template by using the language model of the preset intelligent decoding engine; A decoding module, configured to use a decoder of the preset intelligent decoding engine to decode the conference voice data according to the first probability value and the second probability value to obtain conference text data; An error correction module, used to perform error correction operations on the conference text data to generate meeting minutes; The method of using the decoder of the preset intelligent decoding engine to decode the conference voice data according to the first probability value and the second probability value to obtain conference text data includes: The conference voice data is decoded by using a decoding function of the decoder according to the first probability value and the second probability value to obtain conference text data, wherein the decoding function is: oh * =argmax ω (logP(ω|o)+λlogP LM (ω)+γlen(ω)); Where P(ω|o) is the first probability value, λ is the weight of the language model, and P LM (ω) is the second probability value, and γlen(ω) is the length of the conference voice data.
8. An electronic device, It is characterized in that The method comprises a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method for generating meeting minutes as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, It is characterized in that It stores a computer program, which, when executed by a processor, implements the method for generating meeting minutes as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent conference summary generation method and system
CN110717031A
Voice recognition method and device, computer equipment and storage medium
CN112102815A