Methods and apparatus for generating meeting minutes, storage media and electronic devices
By detecting and segmenting change points in conference audio and combining this with voiceprint feature matching technology, the problem of inaccurate speaker voice recognition in traditional conference recordings has been solved, achieving high-precision conference record generation.
Patent Information
- Application Number
- CN202510043916.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-01-10
AI Technical Summary
In the traditional meeting minutes generation process, the audio files contain a mix of human voices, making it impossible to accurately distinguish the speaker's voice content, resulting in inaccurate meeting minutes.
By acquiring meeting audio, detecting audio change points for segmentation, using voiceprint feature matching technology to identify the speaker's identity, and generating meeting minutes based on audio features.
It improved the accuracy of speaker audio recognition, resulting in the generation of accurate meeting minutes.
Smart Images

Figure CN119943053B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computers, and more specifically, to a method and apparatus for generating meeting minutes, a storage medium, and an electronic device. Background Technology
[0002] Traditional meeting minutes primarily involve recording audio using audio recording devices and storing the recordings to preserve the original meeting material. After the meeting, manual transcription is required, a time-consuming, inefficient, and error-prone process. Especially with a large number of participants and frequent, dense speaking, the recordings become cluttered and difficult to distinguish, resulting in inaccurate minutes. In other words, existing technologies suffer from inaccurate meeting minutes generation.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This invention provides a method and apparatus for generating meeting minutes, a storage medium, and an electronic device to at least solve the technical problem of inaccurate meeting minutes content generation.
[0005] According to one aspect of the present invention, a method for generating meeting minutes is provided, comprising: acquiring a first audio that matches a target meeting activity, wherein the target meeting activity includes multiple participants, and the first audio is object audio generated by at least one participant; if at least one audio change point is detected in the first audio, segmenting the first audio based on the audio change point to obtain a second audio, wherein the audio difference between a first reference audio located before the audio change point and a second reference audio located after the audio change point in the first audio is greater than or equal to a target threshold; performing feature matching sequentially in multiple voiceprint feature sets according to a target comparison order based on the audio features of the second audio, wherein the multiple voiceprint feature sets correspond one-to-one with multiple object identifier sets, the voiceprint feature sets include voiceprint features that match at least one object identifier in the object identifier sets respectively, and the target comparison order corresponds to the matching probability of the object identifiers of the participants in the object identifier sets; if a target voiceprint feature matching the audio features of the second audio is determined from the multiple voiceprint feature sets, generating target content in the meeting minutes based on the object identifier matching the target voiceprint feature and the target text identified from the second audio.
[0006] According to another aspect of the present invention, a meeting record generation apparatus is also provided, comprising: an acquisition unit, configured to acquire a first audio matching a target meeting activity, wherein the target meeting activity includes multiple participants, and the first audio is object audio generated by at least one participant; a segmentation unit, configured to segment a second audio from the first audio based on an audio change point detected from the first audio, wherein the audio difference between a first reference audio located before the audio change point and a second reference audio located after the audio change point is greater than or equal to a target threshold; and a matching unit, configured to match according to the second audio... The audio features of the audio are sequentially matched in multiple voiceprint feature sets according to the target comparison order. Each set of voiceprint features corresponds one-to-one with a set of object identifiers. Each voiceprint feature set includes voiceprint features that match at least one object identifier in the object identifier set. The target comparison order corresponds to the matching probability of the object identifiers of the meeting participants in the object identifier set. The generation unit is used to generate the target content in the meeting minutes based on the object identifier that matches the target voiceprint feature and the target text identified from the second audio, after determining the target voiceprint feature that matches the audio features of the second audio from the multiple voiceprint feature sets.
[0007] Optionally, the generation unit includes: a first matching module, configured to perform feature matching in a first voiceprint feature set based on the audio features of the second audio, wherein the first voiceprint feature set corresponds to a first object identifier set, and the first object identifier set includes object identifiers pre-configured according to the target meeting activity; and a second matching module, configured to perform feature matching in at least one second voiceprint feature set if no voiceprint feature matching the audio features is found in the first voiceprint feature set, wherein the at least one second voiceprint feature set corresponds to at least one second object identifier set, and the second object identifier set is used to indicate the object set determined according to the object organizational structure.
[0008] Optionally, the above-mentioned meeting record generation apparatus further includes an adding unit, configured to obtain at least one object identifier matching the target meeting activity from the activity description information of the target meeting activity; add the voiceprint features matched by each of the at least one object identifier to a first voiceprint feature set; and, in response to determining at least one target object set from the object organizational structure, add the voiceprint features matched by each of the object identifiers included in the at least one target object set to the first voiceprint feature set.
[0009] Optionally, the above-mentioned meeting record generation device further includes a traversal unit, configured to: traverse multiple voiceprint feature sets and perform the following steps: if the current voiceprint feature set includes a first voiceprint feature, configure a first matching weight for the first voiceprint feature, wherein the first voiceprint feature is a voiceprint feature that matches the audio feature of a third audio, and the third audio is the previous audio object that completed audio matching before the second audio; if the current voiceprint feature set includes at least one second voiceprint feature, configure a second matching weight for each of the at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches the audio feature of a fourth audio, and the fourth audio is the previous audio object that completed historical audio matching before the second audio; wherein the first matching weight is greater than at least one second matching weight, and the second matching weight of each of the at least one second voiceprint feature is positively correlated with the matching frequency corresponding to each of the second voiceprint features, and the first matching weight and the second matching weight are used to indicate the matching order of their respective corresponding voiceprint features.
[0010] Optionally, the aforementioned meeting record generation apparatus further includes a target content generation unit, configured to: determine the audio features of the second audio as features to be processed and configure candidate labels for the second audio if no audio features matching the audio features of the second audio are determined from multiple sets of audio features; perform clustering processing on the multiple objects audio and their respective features to be processed to obtain multiple sets of audio features, wherein each set of audio features corresponds to a candidate label; and generate the target content in the meeting record based on the multiple sets of audio features and their respective candidate labels.
[0011] Optionally, the above-mentioned meeting record generation device further includes a change point determination unit, configured to: perform noise reduction preprocessing on the first audio to obtain a first reference audio, wherein the first signal-to-noise ratio of the first reference audio is higher than the second signal-to-noise ratio of the first audio; segment the first reference audio to obtain multiple audio segments; obtain the feature change rate corresponding to each of the multiple audio segments based on the prosodic features corresponding to each of the multiple audio segments; and determine at least one audio change point from the first audio based on the feature change rate corresponding to each of the multiple audio segments.
[0012] Optionally, the aforementioned change point determination unit is further configured to: determine the feature change rate higher than the change rate threshold as the reference feature change rate; take the reference audio segment before the time point corresponding to the reference feature change rate as the third reference audio, and take the reference audio segment after the time point corresponding to the reference feature change rate as the fourth reference audio; obtain the reference audio difference degree between the third reference audio and the fourth reference audio; and, if the reference audio difference degree is greater than or equal to the target threshold, take the time point corresponding to the reference feature change rate as the audio change point.
[0013] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored therein, wherein the computer program is configured to execute the above-described method for generating meeting minutes when running.
[0014] According to another aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program / instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program / instructions from the computer-readable storage medium, and executes the computer program / instructions, causing the computer device to perform the meeting record generation method as described above.
[0015] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the above-described method for generating meeting minutes through the computer program.
[0016] In the above implementation, a first audio recording matching the target meeting activity can be obtained, wherein the target meeting activity includes multiple participants, and the first audio recording is the object audio generated by at least one participant. Then, if at least one audio change point is detected in the first audio recording, a second audio recording is obtained by segmenting the first audio recording based on the audio change point. The audio difference between a first reference audio recording located before the audio change point and a second reference audio recording located after the audio change point in the first audio recording is greater than or equal to a target threshold. Subsequently, based on the audio features of the second audio recording, feature matching is performed sequentially in multiple voiceprint feature sets according to a target comparison order. Each set of voiceprint features corresponds one-to-one with a set of object identifiers. Each voiceprint feature set includes voiceprint features that match at least one object identifier in the object identifier set. The target comparison order corresponds to the matching probability of the object identifiers in the object identifier set that include the participants. This achieves the generation of target content in the meeting minutes based on the object identifier matching the target voiceprint feature and the target text identified from the second audio recording, after determining the target voiceprint feature matching the audio features of the second audio recording from multiple voiceprint feature sets. By detecting and segmenting audio changes to obtain multiple speaker audios and performing feature matching, the accuracy of speaker audio recognition is improved, thereby solving the technical problem of inaccurate meeting record generation. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0018] Figure 1 This is a hardware structure block diagram of an optional meeting recording server device according to an embodiment of the present invention;
[0019] Figure 2 This is a flowchart of an optional method for generating meeting minutes according to an embodiment of the present invention;
[0020] Figure 3 This is a flowchart of another optional method for generating meeting minutes according to an embodiment of the present invention;
[0021] Figure 4 This is a schematic diagram of an optional meeting record generation device according to an embodiment of the present invention;
[0022] Figure 5 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a server device for a data processing method according to an embodiment of this application. For example... Figure 1 As shown, the server device may include one or more ( Figure 1Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0026] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to a method for adjusting the read voltage in the memory according to an embodiment of this application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0027] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0028] As an optional implementation method, such as Figure 2 As shown, the method for generating the above meeting minutes includes the following steps:
[0029] S202, Obtain the first audio that matches the target meeting activity, wherein the target meeting activity includes multiple participants, and the first audio is the object audio generated by at least one participant;
[0030] S204, if at least one audio change point is detected in the first audio, the second audio is obtained by segmenting the first audio based on the audio change point, wherein the audio difference between the first reference audio located before the audio change point and the second reference audio located after the audio change point in the first audio is greater than or equal to a target threshold.
[0031] S206, based on the audio features of the second audio, feature matching is performed sequentially in multiple voiceprint feature sets according to the target comparison order. The multiple voiceprint feature sets correspond one-to-one with multiple object identifier sets. Each voiceprint feature set includes voiceprint features that match at least one object identifier in the object identifier set. The target comparison order corresponds to the matching probability of the object identifiers of the participating objects in the object identifier set.
[0032] S208, if a target voiceprint feature matching the audio features of the second audio is determined from multiple voiceprint feature sets, the target content in the meeting record is generated based on the object identifier matching the target voiceprint feature and the target text identified from the second audio.
[0033] In the above embodiment, step S202 is used to acquire conference audio. In step S202, the target conference activity is a currently ongoing meeting, in which multiple participants exist simultaneously, i.e., the attendees. The audio generated by these attendees is the first audio. It is understood that the first audio can be captured using audio acquisition devices such as a microphone array; for example, the first audio can be captured using a microphone array according to a preset audio format and sampling rate.
[0034] It should be noted that step S204 above is used to segment the acquired first audio according to audio change points. The detection of these audio change points can be achieved through a deep neural network model, such as a Long Short-Term Memory (LSTM) model, a Bidirectional Long Short-Term Memory (BiLSTM) model, a Gated Recurrent Unit (GRU), or a Deep Belief Network formed by stacking multiple Restricted Boltzmann Machines. Other RNN models adept at processing sequential data can also be used; no specific limitation is made here. The deep neural network model determines the switching points between different speakers, i.e., the aforementioned audio change points, by identifying feature changes in the first audio, such as acoustic features, tone patterns, or speaking rhythm. It is understood that these audio change points indicate significant changes in the characteristics of the speech signal in the first audio, which typically coincide with the start of speaking by different participants.
[0035] Specifically, the first audio data is segmented, dividing the long audio file into several shorter segments. For example, a Speaker Change Detection (SCD) algorithm based on an RNN network can be used to segment the first audio data according to the speaker's transitions. After segmentation, each audio segment corresponds to a speaker's continuous and complete speech within a certain period. Specifically, the prosody features of the speech signal are first extracted, including but not limited to rise and fall rate features, pitch features, energy features, and duration features. These features can be extracted using speech signal processing libraries (such as librosa or Praat in Python).
[0036] Optionally, the extracted features can be further filtered, for example, features that can reflect changes in the speaker, such as pitch change rate and energy change rate.
[0037] It should be noted that after the features are extracted, they need to be normalized to ensure that features of all types can be processed by the neural network on a uniform scale.
[0038] The normalized features are input into the neural network model. It can be understood that the neural network model here is a trained neural network model, which can be the aforementioned LSTM (Long Short-Term Memory Network) or GRU (Gated Recurrent Unit), or it can be a self-designed and trained RNN model.
[0039] It should be noted that in the process of designing an RNN model, it is necessary to ensure that the RNN model includes an input layer for receiving the extracted prosody features; at least one hidden layer for capturing dynamic changes in the feature sequence; and an output layer for generating the speaker probability distribution at each time step.
[0040] Furthermore, the speaker probability distribution at each time step can be obtained using the aforementioned RNN model. Further analysis of the probability distribution determines the boundaries of speaker variation. It should be noted that the boundaries of speaker variation can be determined by detecting significant changes in the probability distribution. For example, if a drastic change occurs in the probability distribution at a certain time, that time is considered the boundary of speaker variation.
[0041] It should be noted that after determining the boundary of the speaker's change, it is necessary to further determine the audio change point based on the audio difference before and after the boundary. Specifically, the segmentation operation is triggered only when the audio difference on both sides of the speaker's change boundary exceeds the aforementioned target threshold. In other words, the speaker's change boundary where the audio difference on both sides exceeds the aforementioned target threshold is determined as the aforementioned audio change point. The purpose of doing this is to accurately identify speaker transitions, reduce incorrect and missed segments, and improve segmentation accuracy. The difference can be calculated based on various acoustic features, such as MFCC (Mel frequency cepstral coefficients), fundamental frequency, energy distribution, etc. By comparing the differences of these features on both sides of the change point, it is determined whether the segmentation standard has been met.
[0042] Furthermore, the first audio is segmented based on detected audio change points, resulting in multiple smaller audio segments, each corresponding to a speaker's continuous speech. Post-segmentation processing is then performed, including smoothing the segmentation results to reduce erroneous segmentation, and manual verification and adjustment to ensure accuracy. These smaller audio segments, after post-segmentation processing, are then used as the second audio.
[0043] In this embodiment, when audio change points are detected in the first audio, the original conference audio (first audio) is segmented into multiple independent audio segments (second audio) based on the detected audio change points. Each segment corresponds to the continuous speech of a speaker. The segmented second audio ensures that the audio segments reflect the complete speech of a single speaker while avoiding the problems of blurred voiceprint features or omission of change points caused by excessively long audio segments.
[0044] In step S206 above, feature matching is performed sequentially in multiple voiceprint feature sets according to the audio features of the second audio and the target comparison order, in order to identify the speaker's identity from the segmented audio segment (second audio).
[0045] It should be noted that in step S206 above, the aforementioned voiceprint feature set was collected and preprocessed in the previous voiceprint database establishment step, and contains voice feature information of different participants. The extraction and comparison of voiceprint features can be implemented by machine learning algorithms, such as the aforementioned deep neural network (DNN) or support vector machine (SVM). By using machine learning algorithms such as deep neural networks or support vector machines, feature vectors reflecting individual voice characteristics are extracted from audio signals for subsequent comparison work.
[0046] Then, step S208 is executed. If a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, the target content in the meeting record is generated based on the object identifier matching the target voiceprint feature and the target text identified from the second audio.
[0047] In step S208 above, when the voiceprint feature of the second audio successfully matches a certain voiceprint feature in the voiceprint database, the system will determine the corresponding object identifier, i.e., the speaker's identity, based on the matched voiceprint feature. Next, the second audio is recognized and transcribed into text to form "target text," which is then used to generate the target content in the meeting minutes.
[0048] Specifically, a pre-trained audio-to-text conversion model is first used to convert each audio segment (the second audio segment) into text. This audio-to-text conversion model can be a deep neural network-based model, such as LSTM or a more advanced Transformer model, capable of handling long-duration speech input and performing high-precision speech recognition. It should be noted that the text output by the audio-to-text conversion model includes the start and end times of each sentence to facilitate subsequent text integration and sorting.
[0049] Next, the transcribed text and matched object identifiers are stored and managed using a pre-created data structure. This data structure can be a list or a dictionary, where each element or key-value pair represents the recognition result of an audio segment, containing information such as a timestamp, text content, and speaker identifier.
[0050] Next, timeline integration is performed. Specifically, based on the start and end times of the audio segments, the identified text content is arranged in chronological order to ensure the continuity and time accuracy of the meeting minutes. Furthermore, speaker information is embedded by inserting a matching object identifier into the data entry of each text segment, forming a record format of "Speaker: Speech Content".
[0051] Finally, generate the target content for the meeting minutes. It should be noted that the meeting minutes can be generated using a template. Optionally, the template includes key information such as the time, speaker ID, and content of their remarks. The template supports custom formatting, allowing for adjustments to the report's layout and details to suit individual needs.
[0052] The meeting minutes template is populated with text in the aforementioned "Speaker: Speech Content" format to generate the target content in the meeting minutes. Optionally, the meeting minutes include the speech content corresponding to each audio segment, as well as the accurate timestamp of that content and the identified speaker. Optionally, the meeting minutes may also include a meeting summary and keyword information automatically generated based on the target text.
[0053] Through the above-described embodiments of this application, a first audio recording matching a target meeting activity is first obtained, wherein the target meeting activity includes multiple participants, and the first audio recording is the object audio generated by at least one participant. Then, if at least one audio change point is detected in the first audio recording, a second audio recording is obtained by segmenting the first audio recording based on the audio change point. The audio difference between a first reference audio recording located before the audio change point and a second reference audio recording located after the audio change point in the first audio recording is greater than or equal to a target threshold. Subsequently, based on the audio features of the second audio recording, feature matching is performed sequentially in multiple voiceprint feature sets according to a target comparison order. Each multiple voiceprint feature set corresponds one-to-one with a multiple object identifier set. Each voiceprint feature set includes voiceprint features that match at least one object identifier in the object identifier set. The target comparison order corresponds to the matching probability of the object identifiers in the object identifier set that include the participants. This achieves the generation of target content in the meeting minutes based on the object identifier matching the target voiceprint feature and the target text identified from the second audio recording, after determining the target voiceprint feature that matches the audio features of the second audio recording from multiple voiceprint feature sets. By detecting and segmenting audio changes to obtain multiple speaker audios and performing feature matching, the accuracy of speaker audio recognition is improved, thereby solving the technical problem of inaccurate meeting record generation.
[0054] As an optional implementation, the above-mentioned feature matching based on the audio features of the second audio, in accordance with the target comparison order, sequentially performing feature matching in multiple voiceprint feature sets includes:
[0055] S1, based on the audio characteristics of the second audio, feature matching is performed in the first voiceprint feature set, wherein the first voiceprint feature set corresponds to the first object identifier set, and the first object identifier set includes object identifiers pre-configured according to the target meeting activity;
[0056] S2, if no voiceprint feature matching the audio feature is found in the first voiceprint feature set, feature matching is performed in at least one second voiceprint feature set, wherein the at least one second voiceprint feature set corresponds to at least one second object identifier set, and the second object identifier set is used to indicate the object set determined according to the object organizational structure.
[0057] It is understandable that steps S1 to S2 above constitute a search strategy in the voiceprint matching process. Specifically, in step S1, during the initial stage of voiceprint matching, a search is preferentially performed within the first voiceprint feature set, which corresponds to the first object identifier set. This first voiceprint feature set includes pre-configured attendee voiceprint features based on the specific meeting scenario. Specifically, based on information provided by the meeting organizer, such as a meeting invitation list, a first voiceprint feature set containing the voiceprint features of each attendee on the invitation list is configured. It is understood that this first voiceprint feature set represents a preliminary voiceprint search range. By first searching the first feature set, attendees who speak frequently can be quickly located, thereby reducing unnecessary computational burden and improving matching speed.
[0058] It should be noted that if no matching voiceprint feature is found in the first voiceprint feature set, the search scope needs to be expanded. In this case, at least one second voiceprint feature set should be checked. It is understood that the aforementioned second voiceprint feature set is determined based on a broader or more specific organizational structure, such as the company as a whole, various departments within the company, external partners, or specific cross-departmental or cross-company teams. In other words, through step S2 above, a search can be conducted in a wider voiceprint database, ensuring that all speakers can be identified as much as possible even when participant information is incomplete.
[0059] As an optional implementation, before performing feature matching from the audio features of the second audio in the first voiceprint feature set, the above further includes at least one of the following:
[0060] S1, obtain at least one object identifier that matches the target meeting activity from the activity description information of the target meeting activity; add the voiceprint features that match each of the at least one object identifier to the first voiceprint feature set;
[0061] S2, in response to determining at least one set of target objects from the object organization structure, adds the voiceprint features that match the object identifiers included in the at least one set of target objects to the first set of voiceprint features.
[0062] Steps S1 to S2 described above constitute the method for determining the first voiceprint feature set. In step S1, attendee information is first automatically extracted from the activity description information of the target meeting. This information may include a meeting invitation list, a list of expected participants, or any document that mentions attendees; no specific limitations are imposed here. Attendees' identifiers (such as name, employee ID, etc.) are matched against an existing voiceprint feature database, and the voiceprint features corresponding to these identifiers are added to a temporary voiceprint feature library (the first voiceprint feature set).
[0063] It is worth noting that step S1 above is an automatic addition method, while step S2 above is a manual addition method. In step S2, one or more target object sets are determined based on the object's organizational structure (such as the company's departmental structure, project team members, etc.). For example, when the meeting involves a specific department or working group, the voiceprint features corresponding to the object identifiers in the target object set corresponding to the involved department or group are added to the first voiceprint feature set. This is to avoid the first voiceprint feature set being incomplete due to attendees not appearing on the meeting invitation list.
[0064] This implementation allows for dynamic adjustment of the voiceprint database search scope based on the actual participants in the meeting, prioritizing the search of voiceprint features from known participants, thereby significantly improving recognition speed and reducing computational resource consumption. Furthermore, since the voiceprint database is customized for the specific meeting scenario, the accuracy of voiceprint matching is also enhanced.
[0065] It should be noted that each voiceprint feature in the first voiceprint feature set in the aforementioned embodiment comes from the existing total voiceprint library. Each voiceprint feature in the total voiceprint library can be obtained by automatic identification, clustering, and labeling based on historical meeting recordings.
[0066] Specifically, the first step is to denoise the recordings of the concluded historical meetings. The denoising method can use a Deep Spectral Residual Network (DSRN). The audio denoising process is explained below.
[0067] The input speech signal is transformed to the spectral domain using a Fast Fourier Transform (FFT), and the spectral data is divided into frames of several milliseconds in length to facilitate frame-by-frame processing by the DSRN network. The spectral data is then input into the trained DSRN model, which outputs the predicted residual of the noise spectrum.
[0068] The predicted noise residue is subtracted from the input spectral data to obtain the denoised spectral data. Then, the inverse Fourier transform (iSTFT) is used to convert the denoised spectral data back to the time domain signal, thus obtaining the denoised speech signal.
[0069] The denoised speech signal is then segmented so that each segment contains continuous speech from a single speaker. The segmentation method is described in steps S204 to S206, and will not be repeated here.
[0070] Furthermore, unsupervised clustering is performed on each segmented audio segment. Specifically, the speaker features of each segmented audio segment are extracted first. The extraction method can be MFCC (Mel Frequency Cepstral Coefficients), LPCC (Linear Predictive Cepstral Coefficients), etc., without specific limitations here.
[0071] In the case of MFCC extraction, the audio signal is converted to the frequency domain, and the Mel filter bank is applied to enhance the pitch and tone components in the speech signal. Then, Discrete Cosine Transform (DCT) is performed to compress and extract the main features of the spectral information.
[0072] When using LPCC as the extraction method, future sample values are predicted and extracted from the audio signal based on linear predictive analysis. Then, features are obtained by performing a discrete cosine transform on the logarithmic energy of the prediction error. Specifically, pre-emphasis is first performed, then the prediction error is estimated using linear predictive coding (LPC), and finally, a DCT transform is performed on the logarithmic energy of the prediction error to extract the linear predictive cepstral coefficients.
[0073] Next, an unsupervised learning algorithm is used to cluster the extracted voiceprint features to distinguish different speakers. The unsupervised learning algorithm mentioned above can be K-Means, GMM-HMM (Gaussian Mixture Model-Hidden Markov Model), DBSCAN (density-based clustering algorithm), or hierarchical clustering, etc., without specific limitations.
[0074] For each clustering result, automatic labeling is performed using speaker identifiers (speaker's name and employee ID). Optionally, for clusters that cannot be determined by automatic labeling, a manual review process is introduced, where professionals verify the identities and label the clusters based on the audio content and meeting minutes.
[0075] Furthermore, the voiceprint features from the labeled clustering results are stored in the aforementioned total voiceprint database and associated with the corresponding object identifiers (name, employee number, etc.).
[0076] Through the above implementation method, the construction of the total voiceprint database is realized, so that the first feature set can obtain voiceprint features from the total voiceprint database.
[0077] As an optional implementation, before performing feature matching sequentially in multiple speaker feature sets according to the target comparison order based on the audio features of the second audio, the method further includes:
[0078] Traverse multiple voiceprint feature sets and perform the following steps:
[0079] S1, if the current voiceprint feature set includes the first voiceprint feature, configure the first voiceprint feature with the first matching weight, wherein the first voiceprint feature is the voiceprint feature that matches the audio feature of the third audio, and the third audio is the previous audio object that has completed audio matching before the second audio.
[0080] S2, if the current voiceprint feature set includes at least one second voiceprint feature, configure a second matching weight for each of the at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches the audio feature of the fourth audio, and the fourth audio is the object audio that has completed historical audio matching before the second audio.
[0081] Wherein, the first matching weight is greater than at least one second matching weight, and the second matching weight of each of the at least one second voiceprint feature is positively correlated with the matching frequency of each of the second voiceprint features. The first matching weight and the second matching weight are used to indicate the matching order of their respective voiceprint features.
[0082] It should be noted that steps S1 and S2 are not sequential. Step S1 is executed when the current voiceprint feature set includes a first voiceprint feature, and step S2 is executed when the current voiceprint feature set includes at least one second voiceprint feature.
[0083] Regarding step S1 above, it is worth noting that, since speakers in continuous meeting dialogues typically exhibit consistent speaking characteristics, the person who spoke most recently is likely to speak again. Therefore, a higher first-match weight is assigned to the most recently matched voiceprint feature (the first voiceprint feature). Setting a higher first-match weight means that the first voiceprint feature will be given priority in subsequent matching processes; in the later matching process, the first voiceprint feature will be matched first with the audio segment to be identified (the second audio).
[0084] Regarding step S2 above, it should be noted that for other voiceprint features (second voiceprint features) that have been matched before but not in the most recent match, matching weights (second matching weights) are also assigned to them. It can be understood that if a second voiceprint feature matches a previous historical audio segment (fourth audio), it indicates that the speaker corresponding to the second voiceprint feature may also be an active speaker in the meeting.
[0085] Specifically, a second matching weight is assigned to each of the second voiceprint features, and the second matching weight is positively correlated with the historical matching frequency of the corresponding second voiceprint feature. It can be understood that the more frequently a participant has spoken in the history of the meeting, the higher the second matching weight assigned to their corresponding second voiceprint feature.
[0086] As an optional implementation, after performing feature matching sequentially in multiple voiceprint feature sets according to the target comparison order based on the audio features of the second audio, the method further includes:
[0087] S1, if no voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, the audio feature of the second audio is determined as the feature to be processed, and candidate labels are configured for the second audio.
[0088] S2, after obtaining multiple object audios and their respective features to be processed, cluster the multiple features to be processed to obtain multiple audio feature sets, where each audio feature set corresponds to a candidate label;
[0089] S3 generates the target content in the meeting minutes based on multiple audio feature sets and their corresponding candidate labels.
[0090] It should be noted that in step S1 above, when an audio segment (the second audio) in the meeting recording does not find a matching voiceprint feature in the existing voiceprint feature set after undergoing the voiceprint feature matching process, it indicates that the speaker information of that audio segment has not yet been recognized or registered by the system. At this time, the system will not directly discard this information, but will mark its audio features as "features to be processed," meaning that the audio features of this part of the audio need further analysis and processing. Simultaneously, for subsequent processing and recognition, the system configures the aforementioned candidate label for this audio segment.
[0091] Understandably, step S2 above is used to perform cluster analysis on unmatched features. It should be noted that during a meeting, the voiceprint features of multiple audio segments may not be directly matched to the known set of voiceprint features. When multiple such audio clips and their respective features to be processed are collected, step S2 is executed for cluster analysis. Specifically, clustering algorithms (such as K-means clustering, hierarchical clustering, DBSCAN, etc.) are used to cluster all features to be processed, forming multiple audio feature sets. It is worth noting that the audio features within each audio feature set have a high degree of similarity.
[0092] Then, step S3 is executed to generate meeting content based on the clustering results of step S2. Specifically, after performing clustering analysis in step S2, the target content of the meeting record is constructed based on the generated multiple audio feature sets and corresponding candidate labels.
[0093] It should be noted that in step S3, each audio feature set serves as a recording of a potential speaker's speech, and the aforementioned candidate tags act as temporary identifiers for that audio feature set. The audio segments and transcribed text within each audio feature set are integrated to form the speech content for each potential speaker, which is then organized, annotated, and finally merged into the meeting minutes.
[0094] Through the above implementation method, even if there are speakers in the meeting who have not pre-registered, the system can automatically classify and record their speech content based on the similarity of their voice characteristics, ensuring the integrity of the meeting records and the accuracy of the information.
[0095] As an optional implementation, before segmenting the second audio from the first audio based on the audio change point, if at least one audio change point is detected from the first audio, the method further includes:
[0096] S1, perform noise reduction preprocessing on the first audio to obtain the first reference audio, wherein the first signal-to-noise ratio of the first reference audio is higher than the second signal-to-noise ratio of the first audio;
[0097] S2, the first reference audio is segmented to obtain multiple audio segments;
[0098] S3, based on the prosodic features corresponding to each of the multiple audio segments, obtain the feature change rate corresponding to each of the multiple audio segments;
[0099] S4, determine at least one audio change point from the first audio based on the characteristic change rate corresponding to each of the multiple audio segments.
[0100] It should be noted that steps S1 to S4 above are used to determine the audio change points, which will be explained in detail below:
[0101] In step S1 above, before further analysis of the first audio, it is first subjected to noise reduction preprocessing to obtain the first reference audio. Specifically, this is achieved by applying noise suppression algorithms, such as spectral subtraction, Wiener filtering, or deep learning-based denoising algorithms, such as DSRN mentioned in the previous embodiment. The noise reduction preprocessing in step S1 improves the clarity of the first audio and reduces the interference of background noise on subsequent analysis. The preprocessed first reference audio has a higher signal-to-noise ratio (SNR) than the first audio, meaning that the speech signal of the first reference audio is stronger relative to the background noise, which is beneficial for the accurate execution of subsequent steps.
[0102] Then, step S2 is performed to segment the first reference audio into multiple audio segments. Specifically, the noise-reduced first reference audio is segmented into multiple shorter audio segments. It should be noted that the segmentation process is based on a time window.
[0103] In step S3 above, for each segmented audio segment, its prosodic features are analyzed. It should be noted that prosodic features reflect information such as the rhythm, pitch, and intensity of speech. Furthermore, the prosodic feature change rate of each audio segment is calculated. It should be noted that the above feature change rate can be the derivative or difference of the feature value, that is, the above feature change rate is obtained by calculating the derivative or difference of the feature value.
[0104] Optionally, step S4 above, which determines at least one audio change point from the first audio based on the characteristic change rate corresponding to each of the multiple audio segments, may specifically include the following steps:
[0105] S4-1, the feature change rate that is higher than the change rate threshold is determined as the reference feature change rate;
[0106] S4-2, the reference audio segment before the time point corresponding to the reference feature change rate is taken as the third reference audio, and the reference audio segment after the time point corresponding to the reference feature change rate is taken as the fourth reference audio.
[0107] S4-3, Obtain the reference audio difference between the third and fourth reference audio;
[0108] S4-4, when the reference audio difference is greater than or equal to the target threshold, the time point corresponding to the rate of change of the reference feature is taken as the audio change point.
[0109] In step S4-1 above, the aforementioned rate of change threshold is a pre-set threshold used to distinguish between significant feature changes and slight fluctuations. All feature rates of change exceeding the aforementioned rate of change threshold are identified as the aforementioned reference feature rates of change. It is understood that the time points corresponding to the aforementioned reference feature rates of change indicate potential audio change points, i.e., time points where possible speaker transitions or significant changes in tone may occur.
[0110] After determining the reference feature change rate, step S4-2 above is then performed. The portion before the time point corresponding to each reference feature change rate is defined as the third reference audio, which contains the sound information before the change point. Correspondingly, the audio segment after the change point is defined as the fourth reference audio, which contains the sound information after the change point. It should be noted that the determination of the third and fourth reference audio can be based on the time points corresponding to the reference feature change rate. Specifically, the third reference audio corresponding to a reference feature change rate is from the time point corresponding to the previous reference feature change rate to the time point corresponding to the current reference feature change rate, and the fourth reference audio is from the time point corresponding to the current reference feature change rate to the time point corresponding to the next reference feature change rate.
[0111] Then, step S4-3 is executed to calculate the reference audio difference between the third reference audio and the fourth reference audio. The calculation method for the reference audio difference is the same as the calculation method for the audio difference in step S204 of the previous embodiment, and will not be repeated here.
[0112] Finally, perform step S4-4 above to check the reference audio difference between all third and fourth reference audios. If the reference audio difference between a pair of third and fourth reference audios is greater than or equal to the preset target threshold, it can be determined that the change point between them is valid, that is, an audio change has occurred at that time point. At this time, the time point corresponding to the corresponding reference feature change rate is taken as the audio change point.
[0113] Figure 3 This is a flowchart of another optional method for generating meeting minutes according to an embodiment of the present invention, which is described below in conjunction with... Figure 3 This section explains the method for generating a complete meeting record.
[0114] In execution Figure 3 Before step S302, voiceprint information preparation is required. Specifically, firstly, voiceprint information of specific individuals (within the group or a small group of participants) is collected in advance. By collecting a certain length of high-quality audio, it is converted into a digital signal, and acoustic feature information (voiceprint) is extracted from the audio and stored.
[0115] Next, the stored feature information is sharded into databases. After voiceprint collection, the voiceprints are initially sharded to avoid slow queries and errors caused by matching the same audio segment over an excessively large area. Specifically, the sharding method can be based on departments, groups, or teams. In essence, each person's voiceprint information is stored in a tree-structured database relationship. The root node of this tree structure is the entire group database, the next level node is the first-level department database, and subsequent node levels correspond to second-level department databases, business department databases, and team databases, respectively. Furthermore, the same person can exist in multiple databases at the same level. For example, a person might be in another department. For instance, if Zhang San from business department A is seconded to business department B, then both business department A and business department B's business department databases will contain Zhang San's voiceprint information.
[0116] Optionally, in step S390, voiceprint pre-registration involves retrieving the list of attendees for this meeting from the group's voiceprint database based on the actual attendees, defining the scope of attendees (default options include groups, first-level departments, second-level departments, etc.), and using this scope as the search range for voiceprint matching in subsequent audio files. It should be noted that without step S390, the entire group's voiceprint database would be used as the search scope for voiceprint matching in subsequent audio files. The purpose of step S390 is to search the search range during voiceprint matching to improve response speed and reduce the processing time for speaker differentiation.
[0117] After preparing the voiceprint information, execute Figure 3 In step S302, the meeting recording is initiated. It should be noted that in order to ensure the quality and effect of subsequent speech recognition and voiceprint comparison, a professional microphone is used for audio recording.
[0118] S304, Audio Preprocessing. This step performs noise reduction, echo removal, and voice enhancement on the recorded meeting audio. Understandably, to ensure the quality of subsequent processing, step S304 uses mathematical calculations to perform noise reduction, echo removal, and voice enhancement on the audio, making the spoken voice as clear as possible. Optionally, MFCC noise reduction is used. Specifically, the two aligned signals are differentially calculated to remove common ambient noise. The formula for differential calculation is:
[0119]
[0120] in, It is the differential signal. and These are the aligned signals from the two microphones.
[0121] S306, Speaker transition detection. Specifically, the audio is then segmented based on speaker switching. In this step, the audio is segmented using an RNN-based model. The specific segmentation method is described in step S204 above and will not be repeated here.
[0122] S308 performs speech-to-speech conversion segment by segment. Each segmented audio piece is transcribed using ASR. During the transcription process, a single audio piece is transcribed into multiple clauses. For example, if speaker 1 speaks 26 sentences in this part, it will be transcribed into 26 corresponding clauses.
[0123] S310, Align the transcription results. Reassemble the clauses according to the segmentation of each piece until each long text consists of continuous speech from the same speaker.
[0124] S312, Speaker Matching: This involves matching each segmented continuous speech fragment with a speaker. The matching method is voiceprint comparison, and the specific comparison process includes:
[0125] S312-1 First, search the temporary sub-database (the pre-registered voiceprint database obtained in step S390), department, and group sub-databases. If a speaker is matched, return the pre-registered speaker information and the similarity of the match.
[0126] S312-2, If no results are found in the pre-registration database, search the temporary anonymous database (a new temporary anonymous database is created for this meeting to store speakers who have not pre-registered their voiceprint information, i.e., temporary attendees, such as attendees from other companies).
[0127] S312-3, If no result is found in the temporary anonymous database, register the speaker as a temporary speaker and assign a number, and mark the speaker of this sentence with this number so that other sentences can match the role in process b;
[0128] S312-4 If no match is found and registration is successful, it indicates that the audio segment may have been caused by insufficient acoustic information due to speaking too briefly, or low audio quality due to being too far from the microphone, resulting in a failure to match the voiceprint information. In this case, it is marked as pending processing.
[0129] It should be noted that acoustic clustering is used for matching the audio to be processed. During the clustering process, both the labeled audio data (successfully matched to specific roles in steps S312-1 to S312-3) and the audio segment to be processed are used as input to analyze the characteristics of the audio, improving the accuracy of acoustic matching and ensuring that the audio can be efficiently assigned to the corresponding speaker. Through the above combined processing techniques, the actual speaker information for each segment of a speech in a conference audio session can be quickly and accurately identified in large-scale audio data.
[0130] S314, Organize the results sentence by sentence and paragraph by paragraph, and label the speakers. After the aforementioned steps, namely text transcription, segmentation, and speaker matching of the meeting audio, each speech in the meeting audio is matched with a speaker role. At this point, the relative time of each text segment and its audio in the meeting audio (the start and end times of the audio segment), as well as detailed data such as the speaker and the number of words in the clauses, are assembled.
[0131] S316, Compile into a report. Based on the assembly results, i.e., using the assembly results from the previous step S314, a full-text meeting report is generated. The report can be compiled based on information such as the relative time of clauses, the transcription results of clauses, and the speakers of clauses, forming a complete report of "who said what at what time". On this basis, the full-text information can also be used to intelligently generate meeting minutes, meeting summaries, and other information, making it easier to summarize and record.
[0132] Optionally, after step S316 above, the generated report can be further processed, such as removing sensitive words and correcting proper nouns, etc., without specific limitations.
[0133] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0134] According to another aspect of the present invention, an apparatus for generating meeting minutes for implementing the above-described method for generating meeting minutes is also provided. For example... Figure 4 As shown, the device includes: an acquisition unit 402, configured to acquire a first audio that matches a target meeting activity, wherein the target meeting activity includes multiple participants, and the first audio is object audio generated by at least one participant; a segmentation unit 404, configured to segment a second audio from the first audio based on the audio change point when at least one audio change point is detected in the first audio, wherein the audio difference between a first reference audio located before the audio change point and a second reference audio located after the audio change point in the first audio is greater than or equal to a target threshold; and a matching unit 406, configured to match the second audio according to the audio characteristics of the second audio. Following the target comparison order, feature matching is performed sequentially in multiple voiceprint feature sets. Each voiceprint feature set corresponds one-to-one with a set of object identifiers. Each voiceprint feature set includes voiceprint features that match at least one object identifier in the set of object identifiers. The target comparison order corresponds to the matching probability of the object identifiers of the meeting participants in the set of object identifiers. The generation unit 408 is used to generate the target content in the meeting minutes based on the object identifier that matches the audio features of the second audio, after determining the target voiceprint feature that matches the audio features of the second audio from multiple voiceprint feature sets, and the target text identified from the second audio.
[0135] Optionally, the generation unit 408 includes: a first matching module, configured to perform feature matching in a first voiceprint feature set based on the audio features of the second audio, wherein the first voiceprint feature set corresponds to a first object identifier set, and the first object identifier set includes object identifiers pre-configured according to the target meeting activity; and a second matching module, configured to perform feature matching in at least one second voiceprint feature set if no voiceprint feature matching the audio features is found in the first voiceprint feature set, wherein the at least one second voiceprint feature set corresponds to at least one second object identifier set, and the second object identifier set is used to indicate the object set determined according to the object organizational structure.
[0136] Optionally, the above-mentioned meeting record generation apparatus further includes an adding unit, configured to obtain at least one object identifier matching the target meeting activity from the activity description information of the target meeting activity; add the voiceprint features matched by each of the at least one object identifier to a first voiceprint feature set; and, in response to determining at least one target object set from the object organizational structure, add the voiceprint features matched by each of the object identifiers included in the at least one target object set to the first voiceprint feature set.
[0137] Optionally, the above-mentioned meeting record generation device further includes a traversal unit, configured to: traverse multiple voiceprint feature sets and perform the following steps: if the current voiceprint feature set includes a first voiceprint feature, configure a first matching weight for the first voiceprint feature, wherein the first voiceprint feature is a voiceprint feature that matches the audio feature of a third audio, and the third audio is the previous audio object that completed audio matching before the second audio; if the current voiceprint feature set includes at least one second voiceprint feature, configure a second matching weight for each of the at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches the audio feature of a fourth audio, and the fourth audio is the previous audio object that completed historical audio matching before the second audio; wherein the first matching weight is greater than at least one second matching weight, and the second matching weight of each of the at least one second voiceprint feature is positively correlated with the matching frequency corresponding to each of the second voiceprint features, and the first matching weight and the second matching weight are used to indicate the matching order of their respective corresponding voiceprint features.
[0138] Optionally, the aforementioned meeting record generation apparatus further includes a target content generation unit, configured to: determine the audio features of the second audio as features to be processed and configure candidate labels for the second audio if no audio features matching the audio features of the second audio are determined from multiple sets of audio features; perform clustering processing on the multiple objects audio and their respective features to be processed to obtain multiple sets of audio features, wherein each set of audio features corresponds to a candidate label; and generate the target content in the meeting record based on the multiple sets of audio features and their respective candidate labels.
[0139] Optionally, the above-mentioned meeting record generation device further includes a change point determination unit, configured to: perform noise reduction preprocessing on the first audio to obtain a first reference audio, wherein the first signal-to-noise ratio of the first reference audio is higher than the second signal-to-noise ratio of the first audio; segment the first reference audio to obtain multiple audio segments; obtain the feature change rate corresponding to each of the multiple audio segments based on the prosodic features corresponding to each of the multiple audio segments; and determine at least one audio change point from the first audio based on the feature change rate corresponding to each of the multiple audio segments.
[0140] Optionally, the aforementioned change point determination unit is further configured to: determine the feature change rate higher than the change rate threshold as the reference feature change rate; use the reference audio segment before the time point corresponding to the reference feature change rate as the third reference audio, and the reference audio segment after the time point corresponding to the reference feature change rate as the fourth reference audio; obtain the reference audio difference degree between the third reference audio and the fourth reference audio; and, if the reference audio difference degree is greater than or equal to the target threshold, use the time point corresponding to the corresponding reference feature change rate as the audio change point. Optionally, in this embodiment, the embodiments to be implemented by the above-mentioned unit modules can refer to the above-mentioned method embodiments, and will not be repeated here.
[0141] According to another aspect of the present invention, an electronic device for implementing the above-described method for generating meeting minutes is also provided. This electronic device may be... Figure 5 The terminal device or server shown. This embodiment uses this electronic device as an example for illustration. Figure 5 As shown, the electronic device includes a memory 502 and a processor 504. The memory 502 stores a computer program, and the processor 504 is configured to execute the steps of any of the above method embodiments through the computer program.
[0142] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0143] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0144] S1, Obtain the first audio that matches the target meeting activity, wherein the target meeting activity includes multiple participants, and the first audio is the object audio generated by at least one participant;
[0145] S2, if at least one audio change point is detected in the first audio, the second audio is obtained by segmenting the first audio based on the audio change point, wherein the audio difference between the first reference audio located before the audio change point and the second reference audio located after the audio change point in the first audio is greater than or equal to a target threshold.
[0146] S3, based on the audio features of the second audio, feature matching is performed sequentially in multiple voiceprint feature sets according to the target comparison order. The multiple voiceprint feature sets correspond one-to-one with multiple object identifier sets. Each voiceprint feature set includes voiceprint features that match at least one object identifier in the object identifier set. The target comparison order corresponds to the matching probability of the object identifiers of the participating objects in the object identifier set.
[0147] S4, after determining the target voiceprint feature that matches the audio features of the second audio from multiple voiceprint feature sets, generate the target content in the meeting record based on the object identifier matched by the target voiceprint feature and the target text identified from the second audio.
[0148] Alternatively, as those skilled in the art will understand, Figure 5 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 5 This does not limit the structure of the aforementioned electronic devices or electronic equipment. For example, electronic devices or electronic equipment may also include components that are more... Figure 5 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 5 The different configurations shown.
[0149] The memory 502 can be used to store software programs and modules, such as the program instructions / modules corresponding to the meeting record generation method and apparatus in this embodiment of the invention. The processor 504 executes various functional applications and data processing by running the software programs and modules stored in the memory 502, thereby realizing the meeting record generation method described above. The memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 502 may further include memory remotely located relative to the processor 504, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 502 may be used, but is not limited to, to store various elements in the scene screen, meeting record generation information, and other information. As an example, such as... Figure 5 As shown, the memory 502 may include, but is not limited to, the acquisition unit 402, the segmentation unit 404, the matching unit 406, and the generation unit 408 in the meeting record generation device. Furthermore, it may include, but is not limited to, other module units in the meeting record generation device, which will not be elaborated upon in this example.
[0150] Optionally, the aforementioned transmission device 506 is used to receive or send data via a network. Specific examples of the network may include wired and wireless networks. In one example, the transmission device 506 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 506 is a Radio Frequency (RF) module used to communicate with the Internet wirelessly. Furthermore, the aforementioned electronic device also includes: a display 508 for displaying a virtual scene in the interface; and a connection bus 510 for connecting the various module components in the aforementioned electronic device.
[0151] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.
[0152] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.
[0153] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0154] According to one aspect of this application, a computer-readable storage medium is provided, from which a processor of a computer device reads computer instructions, and executes the computer instructions, causing the computer device to perform the aforementioned method for generating meeting minutes. Optionally, in this embodiment, the aforementioned computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0155] S1, Obtain the first audio that matches the target meeting activity, wherein the target meeting activity includes multiple participants, and the first audio is the object audio generated by at least one participant;
[0156] S2, if at least one audio change point is detected in the first audio, the second audio is obtained by segmenting the first audio based on the audio change point, wherein the audio difference between the first reference audio located before the audio change point and the second reference audio located after the audio change point in the first audio is greater than or equal to a target threshold.
[0157] S3, based on the audio features of the second audio, feature matching is performed sequentially in multiple voiceprint feature sets according to the target comparison order. The multiple voiceprint feature sets correspond one-to-one with multiple object identifier sets. Each voiceprint feature set includes voiceprint features that match at least one object identifier in the object identifier set. The target comparison order corresponds to the matching probability of the object identifiers of the participating objects in the object identifier set.
[0158] S4, after determining the target voiceprint feature that matches the audio features of the second audio from multiple voiceprint feature sets, generate the target content in the meeting record based on the object identifier matched by the target voiceprint feature and the target text identified from the second audio.
[0159] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0160] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0161] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. It should be understood that the disclosed client can be implemented in other ways in the several embodiments provided in this application. The device embodiments described above are merely illustrative. For example, the division of the units described above is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of units or modules, and may be electrical or other forms.
[0162] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.
[0163] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for generating meeting minutes, characterized in that, include: Obtain a first audio that matches a target meeting activity, wherein the target meeting activity includes multiple participants, and the first audio is an object audio generated by at least one of the participants; If at least one audio change point is detected in the first audio, a second audio is obtained by segmenting the first audio based on the audio change point, wherein the audio difference between the first reference audio located before the audio change point and the second reference audio located after the audio change point is greater than or equal to a target threshold. Based on the audio features of the second audio, feature matching is performed sequentially in multiple voiceprint feature sets according to the target comparison order. Each of the multiple voiceprint feature sets corresponds one-to-one with a multiple object identifier set. Each voiceprint feature set includes voiceprint features that match at least one object identifier in the object identifier set. The target comparison order corresponds to the matching probability of the object identifiers of the participating objects included in the object identifier set. If a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, the target content in the meeting record is generated based on the object identifier matching the target voiceprint feature and the target text identified from the second audio. Based on the audio features of the second audio, before performing feature matching sequentially in multiple voiceprint feature sets according to the target comparison order, the process further includes: Traversing multiple sets of voiceprint features, the following steps are performed: If the current voiceprint feature set includes a first voiceprint feature, a first matching weight is configured for the first voiceprint feature, wherein the first voiceprint feature is a voiceprint feature that matches the audio feature of a third audio, and the third audio is the previous audio object that completed audio matching before the second audio; if the current voiceprint feature set includes at least one second voiceprint feature, a second matching weight is configured for each of the at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches the audio feature of a fourth audio, and the fourth audio is the previous audio object that completed historical audio matching before the second audio; wherein the first matching weight is greater than at least one second matching weight, and the second matching weight of each of the at least one second voiceprint feature is positively correlated with the matching frequency corresponding to each of the second voiceprint features, and the first matching weight and the second matching weight are used to indicate the matching order of their respective corresponding voiceprint features.
2. The method according to claim 1, characterized in that, The step of performing feature matching sequentially in multiple voiceprint feature sets according to the target comparison order based on the audio features of the second audio includes: Based on the audio characteristics of the second audio, feature matching is performed in the first voiceprint feature set, wherein the first voiceprint feature set corresponds to the first object identifier set, and the first object identifier set includes object identifiers pre-configured according to the target meeting activity; If no voiceprint feature matching the audio feature is found in the first voiceprint feature set, feature matching is performed in at least one second voiceprint feature set, wherein at least one second voiceprint feature set corresponds to at least one second object identifier set, and the second object identifier set is used to indicate the object set determined according to the object organizational structure.
3. The method according to claim 2, characterized in that, Before performing feature matching from the audio features of the second audio in the first voiceprint feature set, at least one of the following is also included: From the activity description information of the target meeting activity, obtain at least one object identifier that matches the target meeting activity; add the voiceprint features that match each of the at least one object identifier to the first voiceprint feature set; In response to identifying at least one set of target objects from the object organizational structure, the voiceprint features matching the object identifiers included in at least one set of target objects are added to the first set of voiceprint features.
4. The method according to claim 1, characterized in that, Based on the audio features of the second audio, and following the target comparison order, feature matching is performed sequentially in multiple voiceprint feature sets, and the process further includes: If no voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, the audio feature of the second audio is determined as the feature to be processed, and a candidate label is configured for the second audio. Given multiple object audios and their respective features to be processed, clustering is performed on the multiple features to be processed to obtain multiple audio feature sets, wherein each audio feature set corresponds to a candidate label; The target content in the meeting record is generated based on multiple sets of audio features and their corresponding candidate tags.
5. The method according to claim 1, characterized in that, If at least one audio change point is detected in the first audio, before segmenting the second audio from the first audio based on the audio change point, the method further includes: The first audio is subjected to noise reduction preprocessing to obtain a first reference audio, wherein the first signal-to-noise ratio of the first reference audio is higher than the second signal-to-noise ratio of the first audio; The first reference audio is segmented to obtain multiple audio segments; Based on the prosodic features corresponding to each of the multiple audio segments, obtain the feature change rate corresponding to each of the multiple audio segments; Based on the characteristic change rate corresponding to each of the multiple audio segments, at least one audio change point is determined from the first audio.
6. The method according to claim 5, characterized in that, Determining at least one audio change point from the first audio based on the feature change rate corresponding to each of the plurality of audio segments includes: The rate of change of the feature that is higher than the rate of change threshold is determined as the reference rate of change of the feature; The reference audio segment before the time point corresponding to the reference feature change rate is used as the third reference audio, and the reference audio segment after the time point corresponding to the reference feature change rate is used as the fourth reference audio. Obtain the reference audio difference between the third reference audio and the fourth reference audio; If the reference audio difference is greater than or equal to the target threshold, the time point corresponding to the rate of change of the reference feature is taken as the audio change point.
7. A device for generating meeting minutes, characterized in that, include: The acquisition unit is used to acquire a first audio that matches the target meeting activity, wherein the target meeting activity includes multiple participants, and the first audio is an object audio generated by at least one of the participants; A segmentation unit is configured to segment a second audio from a first audio based on an audio change point when at least one audio change point is detected in the first audio, wherein the audio difference between a first reference audio located before the audio change point and a second reference audio located after the audio change point is greater than or equal to a target threshold. The matching unit is used to perform feature matching in multiple voiceprint feature sets in sequence according to the audio features of the second audio and the target comparison order. The multiple voiceprint feature sets correspond one-to-one with multiple object identifier sets. The voiceprint feature set includes voiceprint features that match at least one object identifier in the object identifier set. The target comparison order corresponds to the matching probability of the object identifiers of the participating objects in the object identifier set. The generation unit is configured to, when a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, generate target content in the meeting minutes based on the object identifier matching the target voiceprint feature and the target text identified from the second audio. The meeting record generation device is further configured to traverse multiple sets of voiceprint features and perform the following steps: when the current set of voiceprint features includes a first voiceprint feature, configure a first matching weight for the first voiceprint feature, wherein the first voiceprint feature is a voiceprint feature that matches the audio feature of a third audio, and the third audio is the previous audio object that completed audio matching before the second audio; when the current set of voiceprint features includes at least one second voiceprint feature, configure a second matching weight for at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches the audio feature of a fourth audio, and the fourth audio is the previous audio object that completed historical audio matching before the second audio; wherein the first matching weight is greater than at least one second matching weight, and the second matching weight of each of the at least one second voiceprint feature is positively correlated with the matching frequency corresponding to each of the second voiceprint features, and the first matching weight and the second matching weight are used to indicate the matching order of the corresponding voiceprint features.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method according to any one of claims 1 to 6.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 6 through the computer program.
Citation Information
Patent Citations
Audio processing method, readable storage medium, program product and electronic equipment
CN118072722A
Intelligent voice recognition and analysis method and system
CN118571247A