Conference record generation method and device, storage medium and electronic equipment

By detecting the audio change points in the conference audio for segmentation and matching features in the voiceprint feature set, the problem of inaccurate generation of conference records is solved, and more accurate conference records is achieved.

CN119943053AActive Publication Date: 2025-05-06INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Patent Information

Application Number
CN202510043916.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

In the prior art, the content of the meeting minutes is inaccurate, especially when there are a large number of participants, the voices in the recording file are mixed, and it is impossible to accurately distinguish the voice content of the speaker.

Method used

By obtaining the audio matching with the target conference activity, detecting the audio change points for segmentation, and matching features in the voiceprint feature set based on the segmented audio features to generate accurate conference records.

Benefits of technology

It improves the accuracy of speaker audio recognition, solves the problem of inaccurate generation of meeting minutes, and the generated meeting minutes are more accurate and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943053A_ABST
    Figure CN119943053A_ABST
Patent Text Reader

Abstract

The invention discloses a conference record generation method and device, a storage medium and electronic equipment. The method comprises the following steps: under the condition that at least one audio change point is detected from a first audio, segmenting the first audio based on the audio change point to obtain a second audio; according to the audio features of the second audio, feature matching is sequentially carried out in the multiple voiceprint feature sets according to a target comparison sequence, and the target comparison sequence corresponds to the matching probability of the object identification of the conference participating object in the object identification set; and under the condition that a target voiceprint feature matched with the audio feature of the second audio is determined from the plurality of voiceprint feature sets, generating target content in the conference record according to an object identifier matched with the target voiceprint feature and a target text obtained by identifying the second audio. According to the invention, the technical problem of inaccurate content generation of the conference record is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to a method and device for generating conference records, a storage medium and an electronic device. Background Art

[0002] Traditional meeting records mainly use recording equipment to record the meeting audio, store the recording files to retain the original meeting materials, and need to manually organize the meeting content after the meeting. The sorting process is time-consuming, inefficient, and has a high error rate. Especially in the case of a large number of participants, due to the large number of participants and the dense content of the speeches in the meeting, the voices in the recording files are mixed, and it is impossible to accurately distinguish the voice content of the speaker, resulting in inaccurate meeting records. In other words, there is a technical problem of inaccurate meeting record content generation in the prior art.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present invention provide a method and device for generating a conference record, a storage medium and an electronic device, so as to at least solve the technical problem of inaccurate generation of conference record content.

[0005] According to one aspect of an embodiment of the present invention, a method for generating a conference record is provided, comprising: obtaining a first audio that matches a target conference activity, wherein the target conference activity includes a plurality of participants, and the first audio is an object audio generated by at least one participant; when at least one audio change point is detected from the first audio, obtaining a second audio from the first audio based on the audio change point, wherein the audio difference between a first reference audio located before the audio change point in the first audio and a second reference audio located after the audio change point is greater than or equal to a target threshold; performing feature matching in sequence in a plurality of voiceprint feature sets according to an audio feature of the second audio and a target comparison order, wherein the plurality of voiceprint feature sets correspond one-to-one to a plurality of object identification sets, the voiceprint feature sets include voiceprint features that match at least one object identification in the object identification set, and the target comparison order corresponds to the matching probability of the object identification of the participant in the object identification set; when a target voiceprint feature that matches the audio feature of the second audio is determined from the plurality of voiceprint feature sets, generating a target content in the conference record according to the object identification matched by the target voiceprint feature and a target text obtained by recognizing the second audio.

[0006] According to another aspect of an embodiment of the present invention, a device for generating a conference record is also provided, including: an acquisition unit, used to acquire a first audio that matches a target conference activity, wherein the target conference activity includes multiple participants, and the first audio is an object audio generated by at least one participant; a segmentation unit, used to segment the first audio to obtain a second audio based on the audio change point when at least one audio change point is detected from the first audio, wherein the audio difference between a first reference audio located before the audio change point in the first audio and a second reference audio located after the audio change point is greater than or equal to a target threshold; a matching unit, used to obtain a second audio based on the second reference audio located before the audio change point in the first audio The audio features of the audio are matched in sequence in multiple voiceprint feature sets according to a target comparison order, wherein the multiple voiceprint feature sets correspond one-to-one to the multiple object identification sets, the voiceprint feature sets include voiceprint features that match at least one object identification in the object identification set, and the target comparison order corresponds to the matching probability of the object identification of the attending object in the object identification set; a generating unit is used to generate target content in the meeting minutes according to the object identification that matches the target voiceprint feature and the target text obtained by recognizing the second audio when a target voiceprint feature that matches the audio features of the second audio is determined from the multiple voiceprint feature sets.

[0007] Optionally, the above-mentioned generation unit includes: a first matching module, used to perform feature matching in a first voiceprint feature set based on the audio features of the second audio, wherein the first voiceprint feature set corresponds to a first object identification set, and the first object identification set includes object identifications pre-configured according to the target conference activity; a second matching module, used to perform feature matching in at least one second voiceprint feature set if no voiceprint features matching the audio features are found in the first voiceprint feature set, wherein the at least one second voiceprint feature set corresponds to at least one second object identification set respectively, and the second object identification set is used to indicate an object set determined according to the object organization structure.

[0008] Optionally, the above-mentioned conference record generating device also includes an adding unit, which is used to obtain at least one object identifier matching the target conference activity from the activity description information of the target conference activity; add the voiceprint features matching each of the at least one object identifier to the first voiceprint feature set; in response to determining at least one target object set from the object organization structure, add the voiceprint features matching each of the object identifiers included in the at least one target object set to the first voiceprint feature set.

[0009] Optionally, the above-mentioned conference record generation device also includes a traversal unit, which is used to: traverse multiple voiceprint feature sets and perform the following steps: when the current voiceprint feature set includes a first voiceprint feature, configure a first matching weight for the first voiceprint feature, wherein the first voiceprint feature is a voiceprint feature that matches the audio feature of a third audio, and the third audio is the object audio that completed the audio matching before the second audio; when the current voiceprint feature set includes at least one second voiceprint feature, configure a second matching weight for at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches the audio feature of a fourth audio, and the fourth audio is the object audio that completed the historical audio matching before the second audio; wherein the first matching weight is greater than at least one second matching weight, and the second matching weight of each of the at least one second voiceprint features is positively correlated with the matching frequency corresponding to each of the second voiceprint features, and the first matching weight and the second matching weight are used to indicate the matching order of the corresponding voiceprint features.

[0010] Optionally, the above-mentioned conference record generation device also includes a target content generation unit, which is used to determine the audio features of the second audio as features to be processed and configure candidate tags for the second audio when no voiceprint features matching the audio features of the second audio are determined from multiple voiceprint feature sets; when multiple object audios and their respective features to be processed are obtained, cluster the multiple features to be processed to obtain multiple audio feature sets, wherein each audio feature set corresponds to a candidate tag; and generate target content in the conference record based on the multiple audio feature sets and their respective corresponding candidate tags.

[0011] Optionally, the above-mentioned meeting record generation device also includes a change point determination unit, which is used to: perform noise reduction preprocessing on the first audio to obtain a first reference audio, wherein a first signal-to-noise ratio of the first reference audio is higher than a second signal-to-noise ratio of the first audio; divide the first reference audio to obtain multiple audio segments; obtain a feature change rate corresponding to each of the multiple audio segments based on rhythmic features corresponding to each of the multiple audio segments; and determine at least one audio change point from the first audio based on the feature change rate corresponding to each of the multiple audio segments.

[0012] Optionally, the above-mentioned change point determination unit is also used to: determine a feature change rate that is higher than a change rate threshold as a reference feature change rate; use a reference audio segment before a time point corresponding to the reference feature change rate as a third reference audio, and use a reference audio segment after a time point corresponding to the reference feature change rate as a fourth reference audio; obtain a reference audio difference between the third reference audio and the fourth reference audio; and when the reference audio difference is greater than or equal to a target threshold, use the time point corresponding to the corresponding reference feature change rate as an audio change point.

[0013] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for generating meeting records when running.

[0014] According to another aspect of the embodiment of the present application, a computer program product is provided, the computer program product including a computer program / instruction, the computer instruction being stored in a computer-readable storage medium. A processor of a computer device reads the computer program / instruction from the computer-readable storage medium, and the processor executes the computer program / instruction, so that the computer device executes the method for generating the meeting minutes as described above.

[0015] According to another aspect of an embodiment of the present invention, there is provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the method for generating meeting minutes through the computer program.

[0016] In the above implementation, a first audio matching the target conference activity can be obtained, wherein the target conference activity includes multiple participants, and the first audio is an object audio generated by at least one participant; and then, when at least one audio change point is detected from the first audio, a second audio is obtained by segmenting the first audio based on the audio change point, wherein the audio difference between the first reference audio located before the audio change point in the first audio and the second reference audio located after the audio change point is greater than or equal to a target threshold; and then, according to the audio features of the second audio, feature matching is performed in sequence in a plurality of voiceprint feature sets in accordance with a target comparison order, wherein the plurality of voiceprint feature sets correspond one-to-one to the plurality of object identification sets, the voiceprint feature sets include voiceprint features that match at least one object identification in the object identification set, and the target comparison order corresponds to the matching probability of the object identification of the participant in the object identification set; thereby, when a target voiceprint feature matching the audio feature of the second audio is determined from a plurality of voiceprint feature sets, the target content in the conference record is generated according to the object identification matched by the target voiceprint feature and the target text obtained by recognizing the second audio. By detecting and segmenting the audio change points to obtain multiple speaker audios and performing feature matching, the speaker audio recognition accuracy is improved, thus solving the technical problem of inaccurate meeting record content generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0018] Figure 1 is a hardware structure block diagram of an optional conference record server device according to an embodiment of the present invention;

[0019] Figure 2 is a flow chart of an optional method for generating conference records according to an embodiment of the present invention;

[0020] Figure 3 is a flowchart of another optional method for generating conference records according to an embodiment of the present invention;

[0021] Figure 4 is a schematic structural diagram of an optional device for generating conference records according to an embodiment of the present invention;

[0022] Figure 5 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 1 is a hardware structure block diagram of a server device of a data processing method according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the above-mentioned server device may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0026] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as a computer program corresponding to a method for adjusting a read voltage in a memory in an embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to a server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0027] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0028] As an optional implementation, Figure 2 As shown, the method for generating the above-mentioned meeting minutes includes the following steps:

[0029] S202, obtaining a first audio that matches a target conference activity, wherein the target conference activity includes a plurality of conference participants, and the first audio is an object audio generated by at least one of the conference participants;

[0030] S204, when at least one audio change point is detected from the first audio, segmenting the first audio based on the audio change point to obtain a second audio, wherein an audio difference between a first reference audio before the audio change point in the first audio and a second reference audio after the audio change point is greater than or equal to a target threshold;

[0031] S206, performing feature matching in a plurality of voiceprint feature sets in sequence according to a target comparison order based on the audio feature of the second audio, wherein the plurality of voiceprint feature sets correspond one-to-one to the plurality of object identifier sets, the voiceprint feature sets include voiceprint features that match at least one object identifier in the object identifier set, and the target comparison order corresponds to a matching probability of the object identifier of the conference participant included in the object identifier set;

[0032] S208, when a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, target content in the meeting record is generated according to an object identifier matched with the target voiceprint feature and a target text obtained by recognizing the second audio.

[0033] In the above implementation, the above step S202 is used to collect conference audio. In the above step S202, the above target conference activity is a currently ongoing conference, and there are multiple participants in the above target conference activity at the same time, namely the above participants. The object audio generated by the above participants is the above first audio. It is understandable that the above first audio can be captured by an audio acquisition device such as a microphone array, for example, the first audio is captured by a microphone array according to a preset audio format and sampling rate.

[0034] It should be noted that the above step S204 is used to segment the acquired first audio according to the audio change points. The above audio change points can be detected by a deep neural network model, such as a long short-term memory model (LSTM), a bidirectional long short-term memory model (BiLSTM), a gated recurrent unit (GRU), or a deep belief network formed by stacking multiple restricted Boltzmann machines. It can also be other RNN models that are good at processing sequence data, which are not specifically limited here. The deep neural network model determines the switching points of different speakers, i.e., the above audio change points, by identifying feature changes in the first audio, such as acoustic features, tone patterns, or speaking rhythm. It can be understood that the above audio change points indicate that the features of the speech signal in the first audio have changed significantly, and the change usually coincides with the beginning of speaking by different participants.

[0035] Specifically, the first audio data is segmented, and the long audio file is divided into several shorter segments. For example, the first audio data can be segmented according to the speaker's conversion through the RNN network-based speaker change detection algorithm (SCD, Speaker Change Detection), and each audio segment after segmentation corresponds to a continuous and complete speech of a speaker within a period of time. Specifically, the rise and fall rate (prosody) features of the speech signal are first extracted, including but not limited to the rise and fall rate features, fundamental frequency (pitch) features, energy (energy) features, duration (duration) features, etc. These features can be extracted through a speech signal processing library (such as Python's librosa or Praat).

[0036] Optionally, the extracted features may be further screened, for example, features that may reflect changes in the speaker, such as pitch change rate, energy change rate, etc., may be screened out.

[0037] It should be noted that after the features are extracted, they need to be normalized to ensure that each type of feature can be processed by the neural network on a uniform scale.

[0038] The normalized features are input into the neural network model. It can be understood that the neural network model here is a trained neural network model, which can be the LSTM (long short-term memory network) or GRU (gated recurrent unit) mentioned above, or it can be a self-designed and trained RNN model.

[0039] It should be noted that in the process of independently designing the RNN model, it is necessary to ensure that the RNN model includes an input layer for receiving the extracted prosody features; at least one hidden layer for capturing dynamic changes in the feature sequence; and an output layer for the output layer to generate the probability distribution of the speaker at each moment.

[0040] Furthermore, the probability distribution of speakers at each moment can be obtained through the above RNN model. Further, the probability distribution can be analyzed to determine the boundary of speaker change. It should be noted that the boundary of speaker change can be realized by detecting significant changes in probability distribution. For example, if the probability distribution changes dramatically at a certain moment, this moment is considered to be the boundary of speaker change.

[0041] It should be pointed out that after determining the boundary of the speaker change, it is necessary to further determine the above-mentioned audio change point based on the audio difference before and after the boundary of the speaker change. Specifically, the segmentation operation will be triggered only when the difference between the audio on both sides of the speaker change boundary exceeds the above-mentioned target threshold. In other words, the speaker change boundary where the difference between the audio on both sides exceeds the above-mentioned target threshold is determined as the above-mentioned audio change point. The purpose of doing this is to accurately identify the speaker transition, reduce miscuts and missed cuts, and improve the accuracy of segmentation. The calculation of the difference can be based on a variety of acoustic features, such as MFCC (Mel frequency cepstral coefficients), fundamental frequency, energy distribution, etc. By comparing the differences between these features on both sides of the change point, it is determined whether the segmentation standard has been met.

[0042] Furthermore, the first audio is segmented according to the detected audio change points to obtain a plurality of smaller audio segments, each of which corresponds to the continuous speech content of a speaker. A post-segmentation processing operation is further performed, including smoothing the segmentation results to reduce mis-segmentation, and performing manual verification and adjustment to ensure the accuracy of the segmentation. The plurality of smaller audio segments that have undergone the post-segmentation processing operation are then used as the second audio.

[0043] Through this embodiment, when an audio change point is detected from the first audio, the original conference audio (first audio) is divided into multiple independent audio segments (second audio) based on the detected audio change point, and each segment corresponds to a continuous speech of a speaker. The segmented second audio can ensure that the audio segment can reflect the complete speech of a single speaker, and avoid the problem of blurred voiceprint features or missing change points caused by too long audio segments.

[0044] In the above step S206, feature matching is performed in sequence in multiple voiceprint feature sets according to the audio features of the second audio in a target comparison order, so as to identify the identity of the speaker from the segmented audio segment (the second audio).

[0045] It should be noted that in the above step S206, the above voiceprint feature set is collected and preprocessed in the early stage through the voiceprint library establishment step, and contains the voice feature information of different participants. The extraction and comparison of voiceprint features can be achieved by machine learning algorithms, such as the aforementioned deep neural network (DNN) or support vector machine (SVM). The feature vector reflecting the individual voice characteristics is extracted from the audio signal through machine learning algorithms such as deep neural network or support vector machine for subsequent comparison.

[0046] Then, the above step S208 is executed. When a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, the target content in the meeting minutes is generated according to the object identifier matched by the target voiceprint feature and the target text obtained by recognizing the second audio.

[0047] In the above step S208, when the voiceprint feature of the second audio successfully matches a voiceprint feature in the voiceprint library, the system will determine the corresponding object identifier, that is, the identity of the speaker, based on the matched voiceprint feature. Then, the second audio is recognized and transcribed into text to form a "target text", thereby generating the target content in the meeting minutes.

[0048] Specifically, each audio clip (second audio) is first converted into text using a trained audio-to-text conversion model. The audio-to-text conversion model can be a model based on a deep neural network, such as LSTM or a more advanced Transformer model, which can process long-term voice input and perform high-precision voice recognition. It should be noted that the text output by the above audio-to-text conversion model contains the start time and end time of each sentence to facilitate subsequent text integration and sorting.

[0049] After that, the transcribed text and the matched object identifiers are stored and managed through a pre-created data structure. The above data structure can be a list or a dictionary, where each element or key-value pair represents the recognition result of an audio clip, including information such as timestamp, text content, and speaker identifier.

[0050] Then, the timeline is integrated. Specifically, the recognized text content is arranged in chronological order according to the start and end time of the audio clip to ensure the coherence and time accuracy of the meeting record. The speaker information is further embedded. In the data entry of each text clip, the matching object identifier is inserted to form a record format of "speaker: speech content".

[0051] Finally, the target content in the meeting minutes is generated. It should be noted that the above meeting minutes can be generated in the form of a template. Optionally, the template of the meeting minutes contains key information such as time, speaker identification and speech content. The template supports custom formats to adjust the layout and details of the report according to needs.

[0052] The meeting minutes template is filled with text content in the format of "speaker: speech content" to generate the target content in the meeting minutes. Optionally, the meeting minutes contain the speech content corresponding to each audio clip, as well as the accurate timestamp of the content and the identified speaker. Optionally, the above-mentioned meeting minutes may also include the meeting summary and keyword information automatically generated based on the target text.

[0053] Through the above-mentioned implementation mode of the present application, firstly, a first audio matching the target conference activity is obtained, wherein the target conference activity includes multiple participants, and the first audio is an object audio generated by at least one participant; then, when at least one audio change point is detected from the first audio, a second audio is obtained by segmenting the first audio based on the audio change point, wherein the audio difference between the first reference audio located before the audio change point in the first audio and the second reference audio located after the audio change point is greater than or equal to the target threshold; then, according to the audio features of the second audio, feature matching is performed in sequence in multiple voiceprint feature sets in accordance with a target comparison order, wherein the multiple voiceprint feature sets correspond one-to-one to the multiple object identification sets, the voiceprint feature sets include voiceprint features that match at least one object identification in the object identification set, and the target comparison order corresponds to the matching probability of the object identification of the participant in the object identification set; thereby, when a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, the target content in the conference record is generated according to the object identification matched by the target voiceprint feature and the target text obtained by recognizing the second audio. By detecting and segmenting the audio change points to obtain multiple speaker audios and performing feature matching, the speaker audio recognition accuracy is improved, thus solving the technical problem of inaccurate meeting record content generation.

[0054] As an optional implementation, the above-mentioned sequentially performing feature matching in multiple voiceprint feature sets according to the audio feature of the second audio in a target comparison order includes:

[0055] S1, performing feature matching in a first voiceprint feature set according to an audio feature of a second audio, wherein the first voiceprint feature set corresponds to a first object identifier set, and the first object identifier set includes object identifiers pre-configured according to a target conference activity;

[0056] S2, when no voiceprint feature matching the audio feature is found in the first voiceprint feature set, feature matching is performed in at least one second voiceprint feature set, wherein the at least one second voiceprint feature set corresponds to at least one second object identification set, and the second object identification set is used to indicate an object set determined according to the object organization structure.

[0057] It can be understood that the above steps S1 to S2 are a search strategy in the voiceprint matching process. Specifically, in the above step S1, in the preliminary stage of voiceprint matching, the first voiceprint feature set will be searched first, and the first voiceprint feature set corresponds to the first object identification set, wherein the first voiceprint feature set includes the voiceprint features of the participants pre-configured according to the specific meeting scenario. Specifically, based on the information provided by the conference organizer, such as the conference invitation list, the first voiceprint feature set including the voiceprint features of each participant on the conference invitation list is configured. It can be understood that the first voiceprint feature set is a preliminary voiceprint search range. By first searching the first feature set, participants who speak frequently can be quickly located, thereby reducing unnecessary computing burden and improving matching speed.

[0058] It should be noted that, if no voiceprint feature matching the audio feature is found in the first voiceprint feature set, it is necessary to expand the search scope of the voiceprint feature, and at this time, check at least one second voiceprint feature set. It is understandable that the above second voiceprint feature set is determined based on a broader or more specific object organizational structure, such as the company as a whole, departments at all levels within the company, external partners of the company, or specific cross-departmental or cross-company groups. In other words, through the above step S2, a search can be performed in a broader voiceprint library to ensure that all speakers can be identified as much as possible even when the participant information is incomplete.

[0059] As an optional implementation manner, before performing feature matching in the first voiceprint feature set based on the audio feature of the second audio, the method further includes at least one of the following:

[0060] S1, obtaining at least one object identifier matching the target conference activity from the activity description information of the target conference activity; adding the voiceprint features matching the at least one object identifier to the first voiceprint feature set;

[0061] S2: In response to determining at least one target object set from the object organization structure, adding voiceprint features that match the object identifiers included in the at least one target object set to the first voiceprint feature set.

[0062] The above steps S1 to S2 are the method for determining the above first voiceprint feature set. In the above step S1, the information of the attendees is first automatically extracted from the activity description information of the target meeting. This information may include a meeting invitation list, a list of expected participants, or any document mentioning the attendees, which is not specifically limited here. The attendee identifiers (such as name, employee ID, etc.) in the attendee information are matched with the existing voiceprint feature database, and the voiceprint features corresponding to these identifiers are added to a temporary voiceprint feature library (the first voiceprint feature set).

[0063] It is worth noting that the above step S1 is an automatic adding method, while the above step S2 is a manual adding method. In the above step S2, one or more target object sets are determined based on the object organizational structure (such as the company's department structure, project team members, etc.). For example, when the meeting involves a specific department or working group, the voiceprint features corresponding to the object identifier in the target object set corresponding to the department or group involved are added to the first voiceprint feature set. This is to avoid the incompleteness of the first voiceprint feature set caused by the participant not appearing in the meeting invitation list.

[0064] Through this implementation, the search scope of the voiceprint library can be dynamically adjusted according to the actual participants of the meeting, and the voiceprint features of these known participants can be searched first, thereby significantly improving the recognition speed and reducing the consumption of computing resources. At the same time, since the voiceprint library is customized according to the specific meeting scene, the accuracy of voiceprint matching is also improved.

[0065] It should be noted that each voiceprint feature in the first voiceprint feature set in the aforementioned embodiment is derived from an existing total voiceprint library, and each voiceprint feature in the total voiceprint library can be obtained by automatic identification, clustering, and labeling based on historical conference recordings.

[0066] Specifically, we first denoise the recordings of the ended historical meetings. The denoising method can use the Deep Spectral Residual Network (DSRN). The following is an explanation of audio denoising.

[0067] The input speech signal is converted to the spectrum domain through Fast Fourier Transform (FFT), and the spectrum data is divided into frames of several milliseconds in length so that the DSRN network can process them frame by frame. The spectrum data is input into the trained DSRN model, and the DSRN model outputs the predicted residual of the noise spectrum.

[0068] The predicted noise residual is subtracted from the input spectrum data to obtain the denoised spectrum data. Then, the denoised spectrum data is converted back to the time domain signal using the inverse Fourier transform (iSTFT) to obtain the denoised speech signal.

[0069] The denoised speech signal is segmented so that each segmented audio segment contains a continuous speech of a speaker. The segmentation method is referred to in steps S204 to S206, which will not be described in detail here.

[0070] Furthermore, unsupervised clustering is performed on each audio segment after segmentation. Specifically, the voiceprint features in each audio segment after segmentation are first extracted. The extraction method can be MFCC (Mel Frequency Cepstral Coefficients), LPCC (Linear Predictive Cepstral Coefficients), etc., which are not specifically limited here.

[0071] When the extraction method is MFCC, the audio signal is converted into the frequency domain, a Mel filter bank is applied to enhance the pitch and timbre components in the speech signal, and then a discrete cosine transform (DCT) is performed to compress and extract the main features of the spectral information.

[0072] When the extraction method is LPCC, based on linear prediction analysis, future sample values ​​are predicted and extracted from the audio signal, and then features are obtained by discrete cosine transforming the logarithmic energy of the prediction error. The specific process is to first perform pre-emphasis, then estimate the prediction error through linear prediction coding (LPC), and then perform DCT transform on the logarithmic energy of the prediction error to extract the linear prediction cepstrum coefficients.

[0073] Afterwards, the extracted voiceprint features are clustered using an unsupervised learning algorithm to distinguish different speakers. The unsupervised learning algorithm may be K-Means, GMM-HMM (Gaussian Mixture Model-Hidden Markov Model), DBSCAN (density-based clustering algorithm) or hierarchical clustering, etc., which are not specifically limited here.

[0074] Each clustering result is automatically labeled with the speaker ID (speaker's name, work number). Optionally, for clusters that cannot be determined by automatic labeling, a manual review process is introduced, and professionals confirm and label the identity based on the audio content and meeting records.

[0075] Furthermore, the voiceprint features in the labeled clustering results are stored in the above-mentioned total voiceprint library and associated with the corresponding object identifier (name, work number, etc.).

[0076] Through the above implementation, the construction of the total voiceprint library is realized, so that the first feature set can obtain voiceprint features from the total voiceprint library.

[0077] As an optional implementation, according to the audio feature of the second audio, before performing feature matching in sequence in multiple voiceprint feature sets in accordance with the target comparison order, the method further includes:

[0078] Traverse multiple voiceprint feature sets and perform the following steps:

[0079] S1, when the current voiceprint feature set includes a first voiceprint feature, configuring a first matching weight for the first voiceprint feature, wherein the first voiceprint feature is a voiceprint feature that matches an audio feature of a third audio, and the third audio is the last audio object that has completed audio matching before the second audio;

[0080] S2, when the current voiceprint feature set includes at least one second voiceprint feature, respectively configure a second matching weight for the at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches an audio feature of a fourth audio, and the fourth audio is an object audio for which historical audio matching is completed before the second audio;

[0081] Among them, the first matching weight is greater than at least one second matching weight, the second matching weight of at least one second voiceprint feature is positively correlated with the matching frequency corresponding to each second voiceprint feature, and the first matching weight and the second matching weight are used to indicate the matching order of the corresponding voiceprint features.

[0082] It should be noted that the above steps S1 and S2 themselves do not constitute a sequential execution relationship. The above step S1 is executed when the current voiceprint feature set includes the first voiceprint feature, and the above step S2 is executed when the current voiceprint feature set includes at least one second voiceprint feature.

[0083] Regarding the above step S1, it is worth noting that since speakers usually have continuous speech characteristics in a continuous conference conversation, the person who spoke most recently is likely to speak again. Then, a higher first matching weight is configured for the most recently matched voiceprint feature (first voiceprint feature). Setting a higher first matching weight means that in the next matching, the first voiceprint feature will be given priority, and in the subsequent matching process, the first voiceprint feature will be first tried to match with the audio segment to be identified (second audio).

[0084] Regarding the above step S2, it should be noted that for other voiceprint features (second voiceprint features) that have been matched before but not the most recent match, matching weights (second matching weights) are also configured for them. It can be understood that the second voiceprint feature matches a previous historical audio clip (fourth audio), indicating that the speaker corresponding to the second voiceprint feature may also be an active speaker in the meeting.

[0085] Specifically, a second matching weight is configured for each second voiceprint feature, and the second matching weight is positively correlated with the historical matching frequency of the corresponding second voiceprint feature. It can be understood that the more frequently a participant speaks in the conference history, the higher the second matching weight assigned to the corresponding second voiceprint feature.

[0086] As an optional implementation, after performing feature matching in sequence in multiple voiceprint feature sets according to the audio feature of the second audio in a target comparison order, the method further includes:

[0087] S1, when no voiceprint feature matching the audio feature of the second audio is determined from the multiple voiceprint feature sets, determining the audio feature of the second audio as a feature to be processed, and configuring a candidate tag for the second audio;

[0088] S2, when a plurality of object audios and respective features to be processed of the plurality of object audios are obtained, clustering the plurality of features to be processed to obtain a plurality of audio feature sets, wherein each audio feature set corresponds to a candidate label;

[0089] S3, generating target content in the meeting record according to multiple audio feature sets and their corresponding candidate tags.

[0090] It should be noted that in the above step S1, when a certain audio clip (second audio) in the conference recording has not found a matching voiceprint feature in the existing voiceprint feature set after the voiceprint feature matching process, it indicates that the speaker information of the audio clip has not been recognized or registered by the system. At this time, the system will not directly abandon this information, but mark its audio features as "pending features", which means that the audio features of this part of the audio need further analysis and processing. At the same time, for subsequent processing and identification, the system configures a candidate tag for the audio clip.

[0091] It is understandable that the above step S2 is used to perform cluster analysis on unmatched features. It should be noted that during the meeting, there may be multiple audio clips whose voiceprint features cannot be directly matched to the known voiceprint feature set. When multiple such object audios and their respective features to be processed are collected, the above step S2 will be executed for cluster analysis. Specifically, all features to be processed are clustered through a clustering algorithm (such as K-means clustering, hierarchical clustering, DBSCAN, etc.) to form multiple audio feature sets. It is worth noting that the audio features within each audio feature set are highly similar.

[0092] Then, the above step S3 is executed to generate the meeting content according to the clustering result of step S2. Specifically, after completing the clustering analysis in step S2, the target content of the meeting record is constructed according to the generated multiple audio feature sets and corresponding candidate tags.

[0093] It should be noted that in step S3, each audio feature set is a speech record of a potential speaker, and the above candidate tags are used as temporary identification tags corresponding to the audio feature set. The audio clips and transcribed texts in each audio feature set are integrated to form the speech content of each potential speaker, which is then sorted and annotated and finally merged into the meeting record.

[0094] Through the above implementation, even if there are speakers who have not pre-registered in the meeting, the system can automatically classify and record the speech content according to the similarity of their voice features, thereby ensuring the integrity of the meeting minutes and the accuracy of the information.

[0095] As an optional implementation manner, when at least one audio change point is detected from the first audio, before the second audio is obtained by segmenting the first audio based on the audio change point, the method further includes:

[0096] S1, performing noise reduction preprocessing on a first audio to obtain a first reference audio, wherein a first signal-to-noise ratio of the first reference audio is higher than a second signal-to-noise ratio of the first audio;

[0097] S2, dividing the first reference audio into multiple audio segments;

[0098] S3, acquiring a feature change rate corresponding to each of the multiple audio clips according to the prosodic features corresponding to each of the multiple audio clips;

[0099] S4: Determine at least one audio change point from the first audio according to the characteristic change rates corresponding to the multiple audio segments.

[0100] It should be noted that the above steps S1 to S4 are used to determine the audio change point, which is described in detail below:

[0101] In the above step S1, before further analyzing the first audio, it is first subjected to noise reduction preprocessing to obtain a first reference audio. Specifically, it is achieved by applying a noise suppression algorithm, such as spectral subtraction, Wiener filtering, or a denoising algorithm based on deep learning, such as the DSRN mentioned in the above embodiment. Through the noise reduction preprocessing of step S1, the clarity of the first audio is improved and the interference of background noise on subsequent analysis is reduced. The first reference audio obtained after preprocessing has a higher signal-to-noise ratio (SNR) than the first audio, that is, the speech signal of the first reference audio is stronger than the background noise, which is conducive to the accurate execution of subsequent steps.

[0102] Then, the above step S2 is performed to segment the first reference audio to obtain multiple audio segments. Specifically, the first reference audio obtained after noise reduction is segmented into multiple shorter audio segments. It should be noted that the segmentation process is performed based on a time window.

[0103] In the above step S3, for each audio segment obtained by segmentation, its prosodic features are analyzed. It should be noted that the prosodic features reflect information such as the rhythm, pitch and intensity of the speech. Further, the prosodic feature change rate of each audio segment is calculated. It should be noted that the above feature change rate can be the derivative or difference of the feature value, that is, the above feature change rate is obtained by calculating the derivative or difference of the feature value.

[0104] Optionally, the above step S4, determining at least one audio change point from the first audio according to the characteristic change rates corresponding to the multiple audio segments, may specifically include the following steps:

[0105] S4-1, determining a feature change rate that is higher than a change rate threshold as a reference feature change rate;

[0106] S4-2, using the reference audio segment before the time point corresponding to the reference feature change rate as the third reference audio, and using the reference audio segment after the time point corresponding to the reference feature change rate as the fourth reference audio;

[0107] S4-3, obtaining a reference audio difference between the third reference audio and the fourth reference audio;

[0108] S4-4, when the reference audio difference is greater than or equal to the target threshold, the time point corresponding to the reference feature change rate is taken as the audio change point.

[0109] In the above step S4-1, the above change rate threshold is a pre-set threshold, which is used to distinguish significant feature changes from slight fluctuations. All feature change rates that are higher than the above change rate threshold are determined as the above reference feature change rates. It can be understood that the time point corresponding to the above reference feature change rate indicates a potential audio change point, that is, a time point at which a speaker may switch or a significant change in intonation occurs.

[0110] After determining the reference feature change rate, the above step S4-2 is then performed. The portion before the time point corresponding to each reference feature change rate is defined as the third reference audio, and the third reference audio contains the sound information before the change point. Correspondingly, the audio segment after the change point is defined as the fourth reference audio, which contains the sound information after the change point. It should be noted that the determination method of the above third reference audio and fourth reference audio can be determined according to the time point corresponding to the reference feature change rate. Specifically, the third reference audio corresponding to a reference feature change rate is from the time point corresponding to the previous reference feature change rate to the time point corresponding to the current reference feature change rate, and the fourth reference audio is from the time point corresponding to the current reference feature change rate to the time point corresponding to the next reference feature change rate.

[0111] Then, step S4-3 is executed to calculate the reference audio difference between the third reference audio and the fourth reference audio. The calculation method of the reference audio difference is referred to the calculation method of the audio difference in step S204 in the above embodiment, which will not be repeated here.

[0112] Finally, execute the above step S4-4 to check the reference audio differences between all the third and fourth reference audios. If the reference audio difference between a pair of third and fourth reference audios is greater than or equal to the preset target threshold, it can be concluded that the change point between them is valid, that is, an audio change occurs at this time point. At this time, the time point corresponding to the corresponding reference feature change rate is used as the audio change point.

[0113] Figure 3 is a flowchart of another optional method for generating conference records according to an embodiment of the present invention. Figure 3 The method for generating a complete meeting record is explained.

[0114] In execution Figure 3 Before step S302 in the above, voiceprint information preparation is required. Specifically, first, voiceprint information of specific people (inside the group or a small group of participants) is collected in advance, and high-quality audio of a certain length is collected and converted into a digital signal. Acoustic feature information (voiceprint) is extracted from the audio and the feature information is stored.

[0115] After that, the stored feature information is divided into libraries. After the voiceprint is collected, the voiceprint is initially divided into libraries to avoid the situation where the same audio segment is used to match the voiceprint in a too large range, resulting in slow query and easy matching errors. The specific library division method can be divided into libraries according to the scope of departments, groups, groups, etc. Specifically, each person's voiceprint information is stored in a tree-structured library relationship. The root node of the above tree structure is the whole group library, the next level node is the first-level department library, and the subsequent node levels correspond to the second-level department library, the business department library, and the group library respectively. And the same person can have multiple libraries at the same level. For example, a person may be in another department. For example, Zhang San from Business Department A is seconded to Business Department B. At this time, the business department libraries of Business Department A and Business Department B each contain Zhang San's voiceprint information.

[0116] Optionally, in step S390, voiceprint pre-registration, based on the actual participants of the meeting, find the list of participants of this meeting from the group's voiceprint library, and define the scope of the participants (the default selection is group, first-level department, second-level department, etc.), and use this scope as the scope of subsequent audio file matching voiceprints. It should be noted that if there is no step 390, the entire group's voiceprint library will be used as the search range for subsequent audio file matching voiceprints. It should be noted that the purpose of step S390 is to search the search range during voiceprint search to improve the response speed and reduce the processing time for speaker differentiation.

[0117] After preparing the voiceprint information, execute Figure 3 In step S302, a conference recording is initiated. It should be noted that in order to ensure the quality and effect of subsequent speech recognition and voiceprint comparison, a professional microphone is used for recording audio.

[0118] S304, audio preprocessing. The recorded conference audio is subjected to noise reduction, echo removal, and voice enhancement. It is understandable that in order to ensure the subsequent processing effect, step S304 performs noise reduction, echo removal, voice enhancement, etc. on the audio through mathematical calculations to make the voice of people speaking in the audio as clear as possible. Optionally, MFCC noise reduction is used, specifically, the two aligned signals are differentially calculated to remove the common ambient sound. The differential calculation formula is:

[0119]

[0120] in, is the signal after differentiation, and are the aligned signals of the two microphones.

[0121] S306, speaker switching detection. Specifically, the audio is segmented according to the switching of the speakers. In this step, the audio is segmented based on the RNN model. The specific segmentation method is referred to the aforementioned step S204 and will not be described in detail here.

[0122] S308, perform voice conversion segment by segment. Each segmented audio segment is transcribed through ASR, and during the transcription process, a single audio segment is transcribed into multiple clauses. For example, if speaker 1 speaks 26 sentences in this section, it is transcribed into 26 corresponding clauses.

[0123] S310, aligning the transcription results. According to the segmentation of each segment, the clauses are reassembled until each long text is the continuous speech content of the same speaker.

[0124] S312, speaker matching, speaker matching is performed on each continuous voice segment after segmentation, and the matching method is voiceprint comparison. The specific comparison process is divided into:

[0125] S312-1, first search in the temporary sub-database (the pre-registered voiceprint database obtained in step S390), department, and group sub-databases. If a speaker is matched, the pre-registered speaker information and the matching similarity are returned;

[0126] S312-2, if the search in the pre-registration database fails, search in the temporary anonymous database (a new temporary anonymous database is created for this conference to store speakers who have not pre-registered their voiceprint information, i.e. temporary participants, such as participants from outside companies, etc.);

[0127] S312-3, if the temporary anonymous database is not searched, register as a temporary speaker and assign a number, and mark the speaker of this sentence with this number, so that other sentences can be matched with the role in process b;

[0128] S312-4: If the search is never successful and the registration is not successful, it means that the acoustic information of this audio segment may not be rich enough due to the short speech, or the audio quality is too low due to the distance from the microphone. In this case, the voiceprint information cannot be successfully matched. In this case, it is marked as pending.

[0129] It should be noted that for the audio to be processed, the acoustic clustering method is adopted for matching. In the process of the clustering method, the marked audio data (that has been successfully matched to a specific role in the stages from step S312-1 to step S312-3) and the audio segment to be processed are used as clustering inputs to analyze the characteristics of the audio to be processed, improve the accuracy of the acoustic matching, and ensure that the audio to be processed can be efficiently attributed to the corresponding speaker. Through the above-mentioned combined technical processing, the actual speaker information of each speech in the conference audio can be quickly and accurately identified in large-scale audio data.

[0130] S314, sort the results sentence by sentence and segment by segment, and mark the speakers. After the above steps, i.e., the conference audio is transcribed, segmented, speaker matched, etc., each speech in the conference audio is matched to the speaker role. At this time, the relative time of each text and audio of the text in the conference audio (the start time and end time of the audio segment), as well as the speaker, the number of words in the clause, and other detailed data are assembled.

[0131] S316, compile into a report. Based on the assembly result, that is, using the assembly result of the previous step S314, a full-text report of the meeting is formed. The report can form a complete report of "who said what at what time" based on the relative time of the clauses, the transcription results of the clauses, the speakers of the clauses and other information. On this basis, the full-text information can also be used to intelligently generate meeting minutes, meeting summaries and other information, which is more convenient for summarizing and recording.

[0132] Optionally, after the above step S316, the generated report may be further subjected to text post-processing, such as removing sensitive words, correcting proper nouns, etc., which are not specifically limited here.

[0133] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0134] According to another aspect of the embodiments of the present invention, there is also provided a device for generating conference records for implementing the above-mentioned method for generating conference records. Figure 4 As shown, the device includes: an acquisition unit 402, which is used to acquire a first audio that matches a target conference activity, wherein the target conference activity includes multiple participants, and the first audio is an object audio generated by at least one participant; a segmentation unit 404, which is used to segment the first audio based on the audio change point to obtain a second audio when at least one audio change point is detected from the first audio, wherein the audio difference between the first reference audio located before the audio change point in the first audio and the second reference audio located after the audio change point is greater than or equal to the target threshold; a matching unit 406, which is used to match the second audio according to the audio features of the second audio. According to the target comparison order, feature matching is performed in sequence in multiple voiceprint feature sets, wherein the multiple voiceprint feature sets correspond one-to-one to the multiple object identification sets, the voiceprint feature sets include voiceprint features that match at least one object identification in the object identification set, and the target comparison order corresponds to the matching probability of the object identification of the attending object in the object identification set; a generation unit 408 is used to generate target content in the meeting minutes according to the object identification that matches the target voiceprint feature and the target text obtained by recognizing the second audio when a target voiceprint feature that matches the audio feature of the second audio is determined from the multiple voiceprint feature sets.

[0135] Optionally, the generation unit 408 includes: a first matching module, used to perform feature matching in a first voiceprint feature set based on audio features of the second audio, wherein the first voiceprint feature set corresponds to a first object identification set, and the first object identification set includes object identifications pre-configured according to the target conference activity; a second matching module, used to perform feature matching in at least one second voiceprint feature set if no voiceprint features matching the audio features are found in the first voiceprint feature set, wherein at least one second voiceprint feature set corresponds to at least one second object identification set, and the second object identification set is used to indicate an object set determined according to the object organization structure.

[0136] Optionally, the above-mentioned conference record generating device also includes an adding unit, which is used to obtain at least one object identifier matching the target conference activity from the activity description information of the target conference activity; add the voiceprint features matching each of the at least one object identifier to the first voiceprint feature set; in response to determining at least one target object set from the object organization structure, add the voiceprint features matching each of the object identifiers included in the at least one target object set to the first voiceprint feature set.

[0137] Optionally, the above-mentioned conference record generation device also includes a traversal unit, which is used to: traverse multiple voiceprint feature sets and perform the following steps: when the current voiceprint feature set includes a first voiceprint feature, configure a first matching weight for the first voiceprint feature, wherein the first voiceprint feature is a voiceprint feature that matches the audio feature of a third audio, and the third audio is the object audio that completed the audio matching before the second audio; when the current voiceprint feature set includes at least one second voiceprint feature, configure a second matching weight for at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches the audio feature of a fourth audio, and the fourth audio is the object audio that completed the historical audio matching before the second audio; wherein the first matching weight is greater than at least one second matching weight, and the second matching weight of each of the at least one second voiceprint features is positively correlated with the matching frequency corresponding to each of the second voiceprint features, and the first matching weight and the second matching weight are used to indicate the matching order of the corresponding voiceprint features.

[0138] Optionally, the above-mentioned conference record generation device also includes a target content generation unit, which is used to determine the audio features of the second audio as features to be processed and configure candidate tags for the second audio when no voiceprint features matching the audio features of the second audio are determined from multiple voiceprint feature sets; when multiple object audios and their respective features to be processed are obtained, cluster the multiple features to be processed to obtain multiple audio feature sets, wherein each audio feature set corresponds to a candidate tag; and generate target content in the conference record based on the multiple audio feature sets and their respective corresponding candidate tags.

[0139] Optionally, the above-mentioned meeting record generation device also includes a change point determination unit, which is used to: perform noise reduction preprocessing on the first audio to obtain a first reference audio, wherein a first signal-to-noise ratio of the first reference audio is higher than a second signal-to-noise ratio of the first audio; divide the first reference audio to obtain multiple audio segments; obtain a feature change rate corresponding to each of the multiple audio segments based on rhythmic features corresponding to each of the multiple audio segments; and determine at least one audio change point from the first audio based on the feature change rate corresponding to each of the multiple audio segments.

[0140] Optionally, the above-mentioned change point determination unit is also used to: determine the feature change rate that is higher than the change rate threshold as the reference feature change rate; use the reference audio segment before the time point corresponding to the reference feature change rate as the third reference audio, and use the reference audio segment after the time point corresponding to the reference feature change rate as the fourth reference audio; obtain the reference audio difference between the third reference audio and the fourth reference audio; when the reference audio difference is greater than or equal to the target threshold, use the time point corresponding to the corresponding reference feature change rate as the audio change point. Optionally, in this embodiment, the embodiments to be implemented by the above-mentioned various unit modules can refer to the above-mentioned various method embodiments, which will not be repeated here.

[0141] According to another aspect of the embodiment of the present invention, an electronic device for implementing the above-mentioned method for generating conference records is also provided. The electronic device may be Figure 5 The terminal device or server shown in the figure. This embodiment is described by taking the electronic device as a terminal device as an example. Figure 5 As shown, the electronic device includes a memory 502 and a processor 504. The memory 502 stores a computer program, and the processor 504 is configured to execute the steps in any of the above method embodiments through the computer program.

[0142] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0143] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0144] S1, obtaining a first audio that matches a target conference activity, wherein the target conference activity includes a plurality of participants, and the first audio is an object audio generated by at least one participant;

[0145] S2, when at least one audio change point is detected from the first audio, segmenting the first audio based on the audio change point to obtain a second audio, wherein an audio difference between a first reference audio before the audio change point in the first audio and a second reference audio after the audio change point is greater than or equal to a target threshold;

[0146] S3, based on the audio features of the second audio, in accordance with a target comparison order, sequentially performing feature matching in a plurality of voiceprint feature sets, wherein the plurality of voiceprint feature sets correspond one-to-one to a plurality of object identifier sets, the voiceprint feature sets include voiceprint features that match at least one object identifier in the object identifier set, and the target comparison order corresponds to a matching probability of the object identifier of the participant in the object identifier set;

[0147] S4, when a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, target content in the meeting minutes is generated according to an object identifier matched with the target voiceprint feature and a target text obtained by recognizing the second audio.

[0148] Alternatively, a person skilled in the art may understand that: Figure 5 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone, a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, etc. Figure 5 The electronic device and the electronic equipment described above are not limited in structure. Figure 5 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 5 Different configurations are shown.

[0149] Among them, the memory 502 can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for generating meeting records in the embodiment of the present invention. The processor 504 executes various functional applications and data processing by running the software programs and modules stored in the memory 502, that is, realizing the above-mentioned method for generating meeting records. The memory 502 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 502 may further include a memory remotely arranged relative to the processor 504, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 502 may be specifically used, but is not limited to, to store information such as various elements in the scene screen, generation information of the meeting records, and so on. As an example, such as Figure 5 As shown, the memory 502 may include, but is not limited to, the acquisition unit 402, the segmentation unit 404, the matching unit 406, and the generation unit 408 in the device for generating the conference record. In addition, it may also include, but is not limited to, other module units in the device for generating the conference record, which will not be described in detail in this example.

[0150] Optionally, the transmission device 506 is used to receive or send data via a network. Specific examples of the above-mentioned network may include wired networks and wireless networks. In one example, the transmission device 506 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 506 is a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly. In addition, the above-mentioned electronic device also includes: a display 508, which is used to display the virtual scene in the interface; and a connection bus 510, which is used to connect the various module components in the above-mentioned electronic device.

[0151] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes in the form of network communication. Among them, the nodes may form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0152] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program / instruction, the computer program / instruction comprising a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, various functions provided by the embodiments of the present application are executed.

[0153] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0154] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for generating meeting minutes. Optionally, in this embodiment, the above-mentioned computer-readable storage medium can be configured to store a computer program for executing the following steps:

[0155] S1, obtaining a first audio that matches a target conference activity, wherein the target conference activity includes a plurality of participants, and the first audio is an object audio generated by at least one participant;

[0156] S2, when at least one audio change point is detected from the first audio, segmenting the first audio based on the audio change point to obtain a second audio, wherein an audio difference between a first reference audio before the audio change point in the first audio and a second reference audio after the audio change point is greater than or equal to a target threshold;

[0157] S3, based on the audio features of the second audio, in accordance with a target comparison order, sequentially performing feature matching in a plurality of voiceprint feature sets, wherein the plurality of voiceprint feature sets correspond one-to-one to a plurality of object identifier sets, the voiceprint feature sets include voiceprint features that match at least one object identifier in the object identifier set, and the target comparison order corresponds to a matching probability of the object identifier of the participant in the object identifier set;

[0158] S4, when a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, target content in the meeting minutes is generated according to an object identifier matched with the target voiceprint feature and a target text obtained by recognizing the second audio.

[0159] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0160] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to perform all or part of the steps of the above methods of various embodiments of the present invention.

[0161] In the above-mentioned embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the above-mentioned units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0162] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the various embodiments of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0163] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for generating conference records, characterized in that: include: Acquire a first audio that matches a target conference activity, wherein the target conference activity includes a plurality of meeting participants, and the first audio is an object audio generated by at least one of the meeting participants; In a case where at least one audio change point is detected from the first audio, second audio is obtained by segmenting the first audio based on the audio change point, wherein an audio difference between a first reference audio before the audio change point in the first audio and a second reference audio after the audio change point is greater than or equal to a target threshold; According to the audio features of the second audio, feature matching is performed in sequence in a plurality of voiceprint feature sets in a target comparison order, wherein the plurality of voiceprint feature sets correspond one-to-one to a plurality of object identifier sets, the voiceprint feature sets include voiceprint features that match at least one object identifier in the object identifier sets, and the target comparison order corresponds to a matching probability of the object identifier of the participant in the object identifier set; When a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets, target content in the meeting minutes is generated based on an object identifier matched with the target voiceprint feature and a target text obtained by recognizing the second audio.

2. The method according to claim 1, characterized in that The performing feature matching in sequence in a plurality of voiceprint feature sets according to the audio feature of the second audio in a target comparison order includes: According to the audio feature of the second audio, feature matching is performed in a first voiceprint feature set, wherein the first voiceprint feature set corresponds to a first object identifier set, and the first object identifier set includes object identifiers preconfigured according to the target conference activity; If a voiceprint feature matching the audio feature is not found in the first voiceprint feature set, feature matching is performed in at least one second voiceprint feature set, wherein at least one of the second voiceprint feature sets corresponds to at least one second object identification set, and the second object identification set is used to indicate an object set determined according to an object organization structure.

3. The method according to claim 2, characterized in that Before performing feature matching in the first voiceprint feature set based on the audio feature of the second audio, the method further includes at least one of the following: Acquire at least one object identifier matching the target conference activity from the activity description information of the target conference activity; add the voiceprint feature matching each of the at least one object identifier to the first voiceprint feature set; In response to determining at least one target object set from the object organization structure, the voiceprint features that match the object identifiers included in at least one of the target object sets are added to the first voiceprint feature set.

4. The method according to claim 2, characterized in that: According to the audio feature of the second audio, before performing feature matching in a plurality of voiceprint feature sets in sequence according to a target comparison order, the method further includes: Traverse the plurality of voiceprint feature sets and perform the following steps: In a case where the current voiceprint feature set includes a first voiceprint feature, configuring a first matching weight for the first voiceprint feature, wherein the first voiceprint feature is a voiceprint feature that matches the audio feature of a third audio, and the third audio is the last object audio that has completed audio matching before the second audio; In a case where the current voiceprint feature set includes at least one second voiceprint feature, configuring a second matching weight for at least one second voiceprint feature, wherein the second voiceprint feature is a voiceprint feature that matches the audio feature of a fourth audio, and the fourth audio is an object audio that has completed historical audio matching before the second audio; Among them, the first matching weight is greater than at least one of the second matching weights, the second matching weight of at least one of the second voiceprint features is positively correlated with the matching frequency corresponding to each of the second voiceprint features, and the first matching weight and the second matching weight are used to indicate the matching order of the corresponding voiceprint features.

5. The method according to claim 1, characterized in that After performing feature matching in a plurality of voiceprint feature sets in sequence according to the target comparison order based on the audio feature of the second audio, the method further includes: If no voiceprint feature matching the audio feature of the second audio is determined from the plurality of voiceprint feature sets, determining the audio feature of the second audio as a feature to be processed, and configuring a candidate tag for the second audio; When a plurality of object audios and the respective features to be processed of the plurality of object audios are obtained, clustering the plurality of features to be processed to obtain a plurality of audio feature sets, wherein each of the audio feature sets corresponds to one of the candidate tags; The target content in the meeting record is generated according to the multiple audio feature sets and the corresponding candidate tags.

6. The method according to claim 1, characterized in that In a case where at least one audio change point is detected from the first audio, before the second audio is obtained by segmenting the first audio based on the audio change point, the method further includes: Performing noise reduction preprocessing on the first audio to obtain a first reference audio, wherein a first signal-to-noise ratio of the first reference audio is higher than a second signal-to-noise ratio of the first audio; Segmenting the first reference audio to obtain multiple audio segments; According to the rhythmic features corresponding to the plurality of audio clips, respectively, acquiring the feature change rates corresponding to the plurality of audio clips; At least one audio change point is determined from the first audio according to the characteristic change rates corresponding to each of the plurality of audio segments.

7. The method according to claim 6, characterized in that The step of determining at least one audio change point from the first audio according to the feature change rates corresponding to each of the plurality of audio segments comprises: Determining the characteristic change rate that is higher than the change rate threshold as a reference characteristic change rate; Using the reference audio segment before the time point corresponding to the reference feature change rate as the third reference audio, and using the reference audio segment after the time point corresponding to the reference feature change rate as the fourth reference audio; Acquire a reference audio difference between the third reference audio and the fourth reference audio; When the reference audio difference is greater than or equal to the target threshold, the time point corresponding to the corresponding reference feature change rate is used as the audio change point.

8. A device for generating conference records, characterized in that: include: An acquiring unit, configured to acquire a first audio matching a target conference activity, wherein the target conference activity includes a plurality of meeting participants, and the first audio is an object audio generated by at least one of the meeting participants; a segmentation unit, configured to, when at least one audio change point is detected from the first audio, segment the first audio based on the audio change point to obtain a second audio, wherein an audio difference between a first reference audio located before the audio change point in the first audio and a second reference audio located after the audio change point is greater than or equal to a target threshold; a matching unit, configured to perform feature matching in sequence in a plurality of voiceprint feature sets according to an audio feature of the second audio and in a target comparison order, wherein the plurality of voiceprint feature sets correspond one-to-one to a plurality of object identifier sets, the voiceprint feature sets include voiceprint features that match at least one object identifier in the object identifier set, respectively, and the target comparison order corresponds to a matching probability of the object identifier of the participant in the object identifier set; A generating unit is used to generate target content in the meeting record based on an object identifier matched with the target voiceprint feature and a target text obtained by recognizing the second audio, when a target voiceprint feature matching the audio feature of the second audio is determined from multiple voiceprint feature sets.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 7 when executed.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

11. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.

Citation Information

Patent Citations

  • Conference record generation method based on voice recognition, device and storage medium

    CN110335612A

  • Voice print recognition method and device, electronic equipment and storage medium

    CN113707182A

  • Conference data processing method and related equipment

    CN114762039A

  • Audio signal processing method, device, system and medium, and conference recording and presenting method, device and system

    CN114792522A

  • Audio processing method, readable storage medium, program product and electronic equipment

    CN118072722A

Cited By

  • Information processing systems, information processing methods, and programs

    JP7897673B1