Audio signal processing method, device and electronic equipment
By performing speech recognition and sound source localization on audio signals and using a microphone array to obtain direction-of-arrival spectrogram information for smoothing, the problem of recording errors in scenarios where multiple people are speaking is solved, efficient and accurate speaker identification and text separation are achieved, and the quality of meeting records is improved.
Patent Information
- Application Number
- CN202011133819.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2040-10-21
AI Technical Summary
In scenarios where multiple people are speaking, existing technologies require a lot of manual work and may result in recording errors or omissions. Existing voice recognition products cannot effectively identify changes in speakers, resulting in low recording accuracy.
By performing speech recognition and sound source localization on the audio signal, the microphone array is used to obtain the direction of arrival spectrogram information, which is then smoothed to determine the sound source localization result, and the text is separated according to the position of the speaker change event.
It improves the accuracy and efficiency of meeting records, reduces the workload of manual labeling, can promptly identify speaker changes and separate speech recognition results, and reduce the error rate.
Smart Images

Figure CN114387970B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio signal recognition, and in particular to an audio signal processing method, device and electronic device. Background Art
[0002] In scenarios where multiple people are speaking, such as meetings and court hearings, there is often a need to record the content of the meeting. Traditionally, a dedicated clerk is required to take notes on site, recording the specific content of the speeches and the corresponding speakers. The recording method is usually for the clerk to type the speech content heard on site into a computer or other computer device. However, this places high demands on the clerk's professional ability and concentration. Once the speaker changes or someone "interrupts the conversation" at a certain moment, the speaker information needs to be changed in a timely manner and the speech content needs to be recorded. As a result, recording errors or omissions may occur.
[0003] Although some speech recognition-related products exist in the existing technology, they can usually only convert the collected voice signals into text. A clerk is required to segment the text strings, mark the specific speaker information, etc. Therefore, a lot of manual operation is still required, and the error rate is still high.
[0004] Therefore, how to further reduce the workload of meeting recorders in scenarios where multiple people are speaking and improve the accuracy of the recorded content has become a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] The present application provides an audio signal processing method, device, and electronic device that can improve the efficiency and accuracy of meeting minutes and reduce the workload of meeting minute-taking staff.
[0006] This application provides the following solutions:
[0007] A method for processing an audio signal, comprising:
[0008] Perform speech recognition and sound source localization on audio signals collected in a multi-person speaking scenario; when localizing the sound source of the audio signal, perform the following processing on a frame-by-frame basis:
[0009] Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum;
[0010] Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0011] The occurrence location of the speaker change event is determined according to the sound source localization results of the plurality of signal frames, and the text obtained by the speech recognition is separated according to the occurrence location of the speaker change event.
[0012] A sound source localization method, comprising:
[0013] determining an audio signal to be processed;
[0014] Obtaining direction-of-arrival spectrogram information of a current signal frame and a target number of signal frames before and after the current signal frame in the audio signal;
[0015] Smoothing a matrix spectrum composed of direction of arrival spectrum information of the signal frame and a target number of signal frames before and after it;
[0016] The sound source localization result of the current signal frame is determined according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame.
[0017] A method for generating meeting minutes, comprising:
[0018] Perform speech recognition and sound source localization on audio signals collected in a conference scenario where multiple people are speaking. When localizing the sound source of the audio signal, perform the following processing on a frame-by-frame basis:
[0019] Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum;
[0020] Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0021] Determining the occurrence location of a speaker change event based on sound source localization results of multiple signal frames, and separating the text obtained by speech recognition based on the occurrence location of the speaker change event;
[0022] A meeting record of the meeting is generated based on the separated multiple text segments.
[0023] A live video processing method, comprising:
[0024] Perform speech recognition and sound source localization on audio signals collected from a live video broadcast of multiple speakers in the same space. The following processing is performed on each frame of the audio signal during sound source localization:
[0025] Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum;
[0026] Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0027] Determining the occurrence location of a speaker change event based on sound source localization results of multiple signal frames, and separating the text obtained by speech recognition based on the occurrence location of the speaker change event;
[0028] According to the time axis information corresponding to the separated multiple text segments, the text segments are added to the video image collected in the live video scene to generate a live video image with subtitles.
[0029] An audio signal processing device, comprising:
[0030] The recognition and positioning unit is used to perform speech recognition and sound source localization on the audio signal collected in a multi-person speaking scenario; wherein, when the recognition and positioning unit localizes the sound source of the audio signal, it includes the following subunits:
[0031] a signal spectrogram processing subunit, configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after it, to form a matrix spectrogram, and to perform smoothing on the matrix spectrogram;
[0032] A positioning result determination subunit is used to determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0033] The recognition text processing unit is used to determine the occurrence location of the speaker change event according to the sound source localization results of multiple signal frames, and separate the text obtained by speech recognition according to the occurrence location of the speaker change event.
[0034] A pickup comprises the audio signal processing device.
[0035] A sound source localization device, comprising:
[0036] an audio signal determining unit, configured to determine an audio signal to be processed;
[0037] a directional spectrogram determining unit, configured to obtain direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after the current signal frame in the audio signal;
[0038] a smoothing processing unit, configured to perform smoothing processing on a matrix spectrogram composed of direction of arrival spectrogram information of the signal frame and a target number of signal frames before and after the signal frame;
[0039] The unit result determination unit is used to determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed arrival direction spectrum corresponding to the current signal frame.
[0040] A sound pickup comprises the sound source localization device.
[0041] A device for generating meeting minutes, comprising:
[0042] The recognition and positioning unit is used to perform speech recognition and sound source localization on audio signals collected in a conference scene where multiple people are speaking; wherein the recognition and positioning unit includes subunits when localizing the sound source of the audio signal:
[0043] a signal spectrogram processing subunit, configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after it, to form a matrix spectrogram, and to perform smoothing on the matrix spectrogram;
[0044] A positioning result determination subunit is used to determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0045] a recognition text processing unit, configured to determine the occurrence location of a speaker change event based on the sound source localization results of the plurality of signal frames, and to separate the text obtained by speech recognition based on the occurrence location of the speaker change event;
[0046] The meeting record generating unit is used to generate the meeting record of the meeting according to the multiple separated text segments.
[0047] A live video processing device, comprising:
[0048] The recognition and positioning unit is used to perform speech recognition and sound source positioning on the audio signal collected in the live video broadcast scene of multiple people speaking, wherein the recognition and positioning unit includes subunits when positioning the sound source of the audio signal:
[0049] a signal spectrogram processing subunit, configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after it, to form a matrix spectrogram, and to perform smoothing on the matrix spectrogram;
[0050] A positioning result determination subunit is used to determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0051] a recognition text processing unit, configured to determine the occurrence location of a speaker change event based on the sound source localization results of the plurality of signal frames, and to separate the text obtained by speech recognition based on the occurrence location of the speaker change event;
[0052] The subtitle adding unit is used to add the text segments to the video image collected in the live video scene according to the time axis information corresponding to the separated multiple text segments, so as to generate a live video image with subtitles.
[0053] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0054] Through embodiments of the present application, speech recognition and sound source localization can be performed on audio signals collected in a multi-speaking scenario. During sound source localization, multiple signal frames within a window of a target length can be taken, centered on the current signal frame, and direction-of-arrival spectrogram information for each signal frame can be obtained. This allows for more data to be used in calculations and processing, and smoothing can be performed on this data. The sound source localization result for the current signal frame can then be determined based on the smoothed direction-of-arrival spectrogram corresponding to the current signal frame. This ensures real-time performance because sound source localization can be performed on a per-frame basis. Furthermore, because the direction-of-arrival spectrogram information for multiple signal frames can be included in the calculations and processing, the accuracy of sound source localization is also increased. Based on this high-precision and real-time sound source localization, even in situations such as "interruptions," speaker change events and their corresponding occurrence locations can be detected promptly based on the sound source localization results. Furthermore, the text obtained from speech recognition can be separated based on the location of the speaker change event. In this way, the speech recognition results are no longer a whole paragraph of text content, but are separated according to the location where the speaker changes. This makes it easier to add speaker labels to specific speech recognition results later, improving efficiency and accuracy and reducing the workload of meeting record staff.
[0055] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0057] Figure 1 It is a schematic diagram of the system architecture provided by the embodiment of the present application;
[0058] Figure 2 is a flowchart of the first method provided in an embodiment of the present application;
[0059] Figure 3-1 、 3-2 This is a schematic diagram of a meeting record generation interface provided in an embodiment of the present application;
[0060] Figure 4 is a flow chart of the second method provided in an embodiment of the present application;
[0061] Figure 5 is a flowchart of the third method provided in an embodiment of the present application;
[0062] Figure 6 is a flowchart of the fourth method provided in an embodiment of the present application;
[0063] Figure 7 is a schematic diagram of a first device provided in an embodiment of the present application;
[0064] Figure 8 is a schematic diagram of a second device provided in an embodiment of the present application;
[0065] Figure 9 is a schematic diagram of a third device provided in an embodiment of the present application;
[0066] Figure 10 is a schematic diagram of a fourth device provided in an embodiment of the present application;
[0067] Figure 11 Schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0068] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0069] In an embodiment of the present application, for a multi-person speaking scenario, in addition to performing speech recognition on the collected audio signal, sound source positioning can also be performed. That is to say, in a specific multi-person speaking scenario, for example, a multi-person meeting, a specific speaker can usually speak at his or her own seat, etc., and the speaker's position usually does not change during the meeting. Therefore, it is possible to determine whether there is an event of speaker change by identifying a sudden change in the direction of the sound source. If so, it is also possible to determine the location where the speaker change occurred, and truncate and separate the speech recognition results at the corresponding location. In this way, the speech recognition result is no longer a whole paragraph of text content, but can be separated into multiple segments according to the time point of the speaker change, so as to facilitate subsequent further speaker marking work.
[0070] Alternatively, in an optional manner, the separated text segments can be automatically labeled with speakers based on the sound source localization results, and text segments from the same sound source direction can also be automatically marked with the same label, for example, they can be marked as 1, 2, 3, or ID-type identifiers such as A, B, and C. Subsequently, information such as the speaker's name can be added through manual modification, and the speaker's name corresponding to the same ID can all be automatically modified. In addition, if the correspondence between the position and the speaker's name in a specific scene can be obtained in advance, the text segment corresponding to the specific sound source direction can also be directly added with a specific speaker name identifier, and so on. In this way, the workload of manual labeling can be reduced, and it is also beneficial to improve accuracy and reduce error rates.
[0071] In implementing the above solution, the specific speaker change time in a multi-speaking scenario is not fixed. In some scenes with heated debates, "interruptions" (that is, one person interrupts another) may occur at any time. Therefore, when performing sound source localization, it is important to accurately identify speaker changes, interruptions, and other situations that occur at any time.
[0072] To this end, in an embodiment of the present application, a microphone array can be used to receive audio signals, and calculations and processing can be performed based on the wave direction spectrogram. Specifically, since the wave direction spectrogram corresponds to a multi-dimensional vector, for example, it can be a 360-dimensional vector, each dimension corresponds to a phase difference (angle), so that each signal frame can have 360 values. In addition, for each signal frame, a window of a target length can be added before and after it, for example, the window length can be 9, and so on. In this way, there can be a total of 360*9=3240 values, and more calculations and processing can be done using these values, which can specifically include filtering processing, etc. For example, some unimportant components in the signal can be filtered out through smoothing processing, so that the human voice with higher resolution becomes more prominent, and so on. In this way, the sound source localization results can be made more accurate. Moreover, since the processing is performed in signal frames, and each signal frame is usually in the millisecond level (for example, 20ms), even if a window of the target length is required, the total time is relatively short, usually less than the pronunciation duration of a single word or syllable during the speech process. Therefore, speaker change events can be discovered more promptly and the corresponding occurrence location can be determined.
[0073] In specific implementation, from the perspective of system architecture, the embodiments of the present application may specifically correspond to products such as "smart pickup". Figure 1 Specifically, it can include both hardware and software. The hardware part mainly corresponds to microphone arrays, which are used to collect audio signals in scenarios where multiple people are speaking. In specific implementation, since participants usually sit around the conference table during a meeting, devices such as smart microphones can be placed in the center of the conference table to better perform speech recognition and sound source localization.
[0074] The software is mainly used to process the specific collected audio signals, and can include two modules: speech recognition and sound source localization. Based on the results of speech recognition and sound source localization, the audio signal can be identified as multiple text segments, and the truncation position of each text segment corresponds to the location where the speaker change event occurred. Among them, for the specific sound source localization module, the direction of arrival spectrogram information of the signal frame and its previous and subsequent target data signal frames can be calculated and processed to improve the accuracy and real-time performance of sound source localization, so as to more accurately and promptly detect speaker change events and determine the corresponding location.
[0075] Among them, part of the software can be run in the form of an application on a terminal device such as a personal computer. For example, the above application can be run on a user computer device responsible for recording the meeting, and a microphone array such as an intelligent microphone can be connected to the computer device. In this way, the audio signal collected by the intelligent microphone can be transmitted to the computer device, and the application running in the computer device can perform specific data processing. It can also provide a corresponding interface to display the speech recognition results and the text segments that are disassembled or annotated according to the sound source localization results, and finally generate the corresponding meeting minutes.
[0076] The specific implementation method provided in the embodiments of the present application is introduced in detail below.
[0077] Example 1
[0078] This embodiment is directed to Figure 1 The multi-person speaking scenario shown in the figure provides an audio signal processing method, see Figure 2 , the method may specifically include:
[0079] S210: Performing speech recognition and sound source localization on audio signals collected in a multi-person speaking scenario; wherein, when performing sound source localization on the audio signals, the following processing is performed on signal frames in the audio signals:
[0080] Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum;
[0081] The sound source localization result of the current signal frame is determined according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame.
[0082] Among them, regarding speech recognition, specifically, it is to allow the machine to convert the speech signal into corresponding text through the recognition and understanding process. Regarding its specific implementation method, it will not be described in detail in the embodiments of this application.
[0083] Regarding sound source localization, in the embodiments of the present application, it mainly refers to sound source localization based on a microphone array. Specifically, it can mainly determine the direction of the sound source (which can be represented by an angle). Since the positions of the microphone and each speaker are usually fixed in scenes such as meetings, when the speaker changes, the direction angle of the specific sound reaching the microphone will change. Therefore, it is possible to detect the direction of the sound source to determine whether a speaker change occurs during the speech. For example, the seat of a certain participant A is 45 degrees relative to the microphone reference angle (for example, 0 degrees to the north, etc.), and the seat of participant B is 90 degrees relative to the microphone reference angle. If at a certain moment t during the speech of participant A, participant B suddenly interrupts A and starts speaking, at this time, through sound source localization, it can be determined that the sound source localization result of a certain signal frame is 45 degrees, and the sound source localization result of the next signal frame suddenly changes to 90 degrees. Therefore, the position of the signal frame can be determined as the position where the speaker has changed, and so on.
[0084] Among them, in order to be able to accurately and timely perform the above-mentioned sound source positioning, in an embodiment of the present application, the wave direction spectrum information of the current signal frame and the target number of signal frames before and after it (for example, 4 before and after) can be first obtained to form a matrix spectrogram, wherein, since the wave direction spectrum can include numerical values corresponding to multiple angles, more calculations and processing can be performed based on these data. Among them, it can include smoothing the matrix spectrogram. That is, by smoothing the information of the current signal frame and the target number of signal frames before and after it, some unimportant components (for example, non-human voices) are filtered out, and important components (for example, sounds with higher human voice resolution) are more prominent. In this way, the result of sound source positioning can be made more accurate.
[0085] There are various possible definitions of signal frames. For example, every 20ms of audio data can be defined as a signal frame, ensuring the real-time performance of sound source localization. The length of the selected window can also be set based on actual needs. However, it should not be too long, as this may affect the real-time performance of the sound source localization results, nor too short, as this may affect the smoothing effect. Therefore, in an optional embodiment, the window length can be set to 9. In other words, a window of length 9 can be selected by taking 4 frames forward and 4 frames backward, centered on the current frame. This window selection method can be used for each current frame.
[0086] Furthermore, a direction of arrival spectrogram can be obtained for each signal frame in the window. There are various ways to obtain a direction of arrival spectrogram. For example, a technique such as gcc-phat (generalized cross-correlation) can be used to calculate a 360-dimensional vector representing 0-360 degrees for a specific signal frame. This vector represents the direction of arrival spectrogram for that signal frame. This generates a 360-dimensional vector for each signal frame. For a window length of 9, a 360*9 matrix spectrogram can be obtained.
[0087] There are several ways to perform smoothing. For example, in one approach, the corresponding kurtosis can be calculated for the current signal frame and a target number of signal frames before and after it. Then, the matrix spectrogram can be smoothed using a target filter and the kurtosis information. Specifically, for each signal frame in the matrix spectrogram, the corresponding kurtosis (kurtosis, a characteristic number that characterizes the height of the peak of the probability density distribution curve at the mean value. Intuitively, kurtosis reflects the sharpness of the peak; high kurtosis means that the increased variance is caused by low-frequency extreme differences above or below the mean value) is calculated, and nine kurtosis values are obtained. There are also various filter options. For example, one optional filter can be a Kalman filter, which can be used to smooth the 360*9 matrix spectrogram using this Kalman filter and the nine kurtosis values as weights. Kalman filtering is an algorithm that uses a linear system state equation and observation data to optimally estimate the system state. Because the observation data is affected by noise and interference in the system, the system state estimation process can also be considered a filtering process.
[0088] After the smoothing process is completed, the angle corresponding to the value that meets the condition (such as the maximum value) in the current frame of the smoothed matrix spectrogram can be taken, and the angle can be determined as the sound source localization result of the current frame.
[0089] S220: Determine the occurrence location of the speaker change event according to the sound source localization results of the multiple signal frames, and separate the text obtained by speech recognition according to the occurrence location of the speaker change event.
[0090] Through the method provided in the embodiments of the present application, a specific sound source localization result can be obtained for each signal frame. After obtaining the sound source localization result for each signal frame, the difference between the sound source localization result of the current signal frame and the sound source localization result of the previous signal frame can be determined. If the difference is greater than a threshold, a speaker change event can be determined, and the location of the current signal frame can be determined as the location of the speaker change event. A speaker change event refers to an event in which the speaker changes. For example, from time t1 to time t2, A is speaking, and from time t2, B begins speaking. A speaker change event occurs at time t2, and accordingly, the location on the timeline at time t2 is the location of the speaker change event. Since A's location is different from B's, the directions of the sound waves generated by their respective speeches reaching the microphone are different. The embodiments of the present application identify the location of time t2 by analyzing the direction of arrival spectrogram.
[0091] The thresholds can be set based on actual circumstances and can be dynamically configured. For example, a threshold setting option can be provided in a specific client interface. This allows different thresholds to be set based on the size of the conference room, the density of attendees, etc. For example, if the density of attendees is greater, the phase differences between adjacent participants relative to their microphones will be closer, and in this case, a smaller threshold can be set.
[0092] In addition, during the specific implementation, the location of the speaker change event determined by sound source positioning can also be verified by extracting the speaker's timbre features. That is to say, different speakers often have different timbres. Therefore, in the process of detecting speaker change events by sound source positioning, this feature can also be used to verify the detection results of the speaker change events. For example, the speaker change event is detected at a certain moment t2 by the sound source positioning method in the embodiment of the present application. After that, the timbre features in several signal frames before and after the moment t2 can be extracted and compared. If there are indeed obvious differences, the credibility of the detection results of the sound source positioning method can be further improved. If, through the comparison of the timbre features, it is found that the timbre features before and after the moment t2 have not changed significantly, then when outputting the separated text, a mark can be added at the position to remind subsequent editors to further confirm the position, etc.
[0093] In the above manner, since the location corresponding to the speaker change event can be identified, the specific speech recognition result can be truncated into multiple text segments, and the specific clerk or other user can subsequently associate the text segment with the specific speaker name and other information. In other words, the result of speech recognition of the audio signal is the text content of the entire paragraph. Of course, these text contents can also be associated with the time axis of the specific audio signal. In the embodiment of the present application, by identifying the speaker change time point in the audio signal, it is possible to determine at which time points the speaker change occurred, and then according to the position of these time points on the time axis, the entire text content of the speech recognition can be separated into multiple text segments, each of which can be a sentence, or multiple sentences, and so on.
[0094] Alternatively, to further reduce the clerk's workload, the separated text can be automatically labeled with a speaker ID. As previously mentioned, tags can be added to specific text segments, and the specific tag content can include the speaker's ID, for example. The speaker's ID can be determined during the audio signal processing process, and the same tag can be added to different text segments corresponding to the same speaker. That is, during a multi-person speaking session, the same speaker may speak multiple times in different time periods. Since the scenarios in the embodiments of this application are typically meetings, speakers typically speak from their own seats. Therefore, when the same speaker speaks in different time periods, the sound source localization results for the signal frames generated within the specific time period should be the same or within the same range. This principle allows not only the location of a speaker change event to be identified, but also the identification of speakers corresponding to different time periods. Specifically, multiple time periods can be separated based on the location of the speaker change event, and the sound source localization results for multiple signal frames within the same time period can be statistically analyzed to determine the range of the sound source localization results for each time period. Then, based on the similarity of the ranges between the different time periods, the speakers corresponding to the multiple time periods can be identified as the same person. Afterwards, the same label is added to the text segments corresponding to different time segments of the same speaker.
[0095] Specifically, from a program implementation perspective, when a speaker begins speaking, the sound source localization result corresponding to the speaker can be identified (the volume can be a certain angle value), and it can be determined whether the user is speaking for the first time. If the speaker has not spoken before, a new ID is assigned to the speaker and added as the speaker label for the currently separated text segment. If the speaker has spoken before, the previously assigned ID can be added as the label for the currently separated text segment.
[0096] Specifically, there are multiple ways to determine whether a speaker is speaking for the first time. For example, after the speaker change is first identified, the correspondence between the ID assigned to the user and the sound source localization result can be saved. When the speaker change is subsequently identified again, it is determined whether the sound source localization result corresponding to the signal frame at the location where the change occurred appears in the previously saved correspondence. If it does not appear, it proves that it is the first time speaking, and a new ID is assigned to it. Otherwise, it is not the first time speaking, and the ID corresponding to the sound source localization result can be used to mark the separated text segment.
[0097] Of course, since a signal frame corresponds to milliseconds, a speaker's speech involves multiple signal frames, and each signal frame can correspond to a sound source localization result. However, it's almost impossible for a user to remain completely still while speaking. Furthermore, inherent errors in sound source localization mean that even if the speaker hasn't changed, the sound source localization results for each frame may not be exactly the same. However, as long as the speaker hasn't changed, the sound source localization results for multiple signal frames will be within a certain range (if they deviate too much, or even exceed a threshold, it will be determined that the speaker has changed), for example, all around 45 degrees. Therefore, after identifying a speaker change event and before the speaker changes again, the sound source localization results corresponding to the signal frames within that time period can be counted to determine the range of the sound source localization results for the corresponding speaker. The corresponding relationship between the speaker ID and the range information of the sound source localization results is then stored. For example, the sound source localization result for a speaker might be between 43 and 47 degrees, etc. In other words, participants who have spoken are assigned a speaker ID, which is then associated with a sound source localization result range. In this way, after detecting the new speaker change time, it can be determined whether the sound source localization result corresponding to the signal frame where the change occurs appears in the interval range of the sound source localization result corresponding to a certain speaker ID. If so, it means that this speaker is not speaking for the first time, and the newly separated text segment can be associated with the speaker ID. Otherwise, this speaker is speaking for the first time, so a new ID can be reassigned to the speaker, and so on.
[0098] In this way, in addition to separating multiple text segments according to the location where the speaker change occurs, a speaker identification can also be added to the text segment. Moreover, although this identification can be marked in the form of an ID, since "same person judgment" can also be performed, the text segments corresponding to the same speaker can also be marked as the same speaker ID. In this way, when editing the name of a specific speaker later, one of the speaker IDs is edited and modified to the name of the speaker, etc., and the labels of other text segments corresponding to the same speaker ID can also be automatically modified. Specifically, an operation option for editing the label of the text segment can be provided. After receiving the editing result of one of the text segments, other text segments with the same label corresponding to the text segment can be determined, and the labels of the other text segments can be modified to the editing result. In this way, batch editing of information such as the speaker's name can be achieved, thereby further improving efficiency and reducing the probability of errors.
[0099] For example, Figure 3-1 As shown in FIG, it is a recognition result interface diagram of the application under a specific implementation mode. As can be seen from the figure, the embodiment of the present application can recognize a specific audio signal as text, and cut the text at the position where the speaker changes according to the sound source localization result, generate multiple text segments, and add speaker labels for different text segments. When different text segments correspond to the same speaker, the same label is also added. For example, Figure 3-1 In the example shown, text segment 1 and text segment 3 both correspond to user 2, text segment 2 corresponds to user 1, and so on. In addition, operation options for editing labels can also be provided for specific text segments, for example, Figure 3-1 As shown at 31 in FIG, the user can modify the label of the text segment through this operation option, and can realize batch modification of multiple different text segments associated with the same label. Figure 3-2 As shown, assuming that the user edits the label of text segment 1 to "Zhang San", the label corresponding to text segment 3 can also be automatically modified to "Zhang San" at the same time, and so on.
[0100] In addition, during specific implementation, recommended tags can also be provided based on the pre-acquired sound feature information of different speakers and their corresponding speaker identifiers. For example, specifically, the sound features of the participants of a certain meeting can be extracted in advance, and the corresponding speaker identifiers can be saved. In this way, assuming that the sound source localization results determine that a speaker change has occurred at a certain moment t2, the sound features of the signal frame after that moment can be extracted, and the previously saved sound features can be retrieved based on the extracted sound features. If the sound features of one of them match, the speaker identifier corresponding to the sound features can be used to provide recommended tags. For example, the speaker identifier saved in the specific index can be the speaker's name, work number, and other identity information, and this name or work number information can be used to provide recommended tags. In this way, if the identification is correct, the workload of subsequent manual editing can be further reduced.
[0101] It should be noted that, in a specific implementation, the audio signal processing method provided in the embodiments of the present application can have a variety of specific application scenarios, for example, it can be a conference scenario where multiple people speak (including a discussion meeting, a court trial meeting, etc.). Alternatively, it can also include a live video broadcast scenario where multiple people speak, etc. Regardless of the specific scenario, multiple speakers can be located in the same space, so that speaker change events can be detected by sound source localization.
[0102] Among them, for situations where there are associated video signals such as live video scenes with multiple people speaking, in an optional implementation, after separating the text obtained by the speech recognition, the time axis information corresponding to the multiple separated text segments (for a specific video, the video signal, audio signal, and text information obtained by specific speech recognition can all correspond to the same time axis) can be added to the associated video image to generate a video image with subtitles. The above-mentioned process of adding subtitles can be carried out during the live broadcast process, targeting the live data stream collected in real time. Therefore, specific speech recognition and sound source positioning, etc., can all be carried out based on the live data stream.
[0103] In summary, through the embodiments of the present application, speech recognition and sound source localization can be performed on audio signals collected in a multi-speaking scenario. When performing sound source localization, multiple signal frames within a window of a target length can be taken, centered on the current signal frame, and the direction of arrival spectrogram information for each signal frame can be obtained. This allows for more data to be used in calculations and processing, and smoothing can be performed on this basis. The sound source localization result for the current signal frame can then be determined based on the smoothed direction of arrival spectrogram corresponding to the current signal frame. In this way, since sound source localization can be performed on a per-frame basis, real-time performance is ensured. Furthermore, since the direction of arrival spectrogram information for multiple signal frames can be included in the calculations and processing, the accuracy of sound source localization is also increased. Based on this high-precision and real-time sound source localization, even in situations such as "interruptions," speaker change events and their corresponding occurrence locations can be detected promptly based on the sound source localization results, and the text obtained from speech recognition can then be separated based on the location of the speaker change event. In this way, the speech recognition result is no longer a whole paragraph of text content, but is separated according to the location where the speaker changes. This makes it easier to add speaker labels to specific speech recognition results later, improving efficiency and accuracy.
[0104] Example 2
[0105] In the above embodiment 1, the information processing method in a specific multi-person speaking scenario is introduced, which involves a specific sound source localization method, and the sound source localization method can also be used in other application scenarios. For this reason, in the embodiment 2 of the present application, a sound source localization method is provided separately, see Figure 4 , the method may specifically include:
[0106] S410: Determine an audio signal to be processed;
[0107] There may be multiple types of audio signals to be processed, for example, an audio signal collected in real time in a certain scene, or a recording result, etc.
[0108] S420: Obtain direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after the current signal frame in the audio signal;
[0109] S430: performing smoothing processing on a matrix spectrogram composed of direction of arrival spectrogram information of the signal frame and a target number of signal frames before and after the signal frame;
[0110] S440: Determine a sound source localization result of the current signal frame according to an angle corresponding to a value that meets a target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame.
[0111] Example 3
[0112] This embodiment 3 provides a specific application solution for a conference scenario where multiple people speak. Specifically, this embodiment 3 provides a method for generating conference records, see Figure 5 , the method may include:
[0113] S510: Perform speech recognition and sound source localization on audio signals collected in a conference scenario where multiple people are speaking. When localizing the sound source of the audio signals, the following processing is performed on signal frames in the audio signals:
[0114] Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum;
[0115] Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0116] S520: Determine the occurrence location of the speaker change event based on the sound source localization results of the multiple signal frames, and separate the text obtained by speech recognition based on the occurrence location of the speaker change event;
[0117] S530: Generate a meeting record of the meeting according to the separated multiple text segments.
[0118] During specific implementation, labels may be added to the multiple separated text segments, and the labels are used to represent speaker identifications.
[0119] In addition, the same label can be added to different text segments corresponding to the same speaker.
[0120] Furthermore, an operation option for editing the label of the text segment can be provided; after receiving the editing result of one of the text segments, other text segments corresponding to the same label as the text segment are determined, and the labels of the other text segments are modified to the editing result, and the editing result corresponds to the identity information of the speaker.
[0121] Example 4
[0122] This fourth embodiment provides a live video processing method for application in live video scenarios. Figure 6 , the method may include:
[0123] S610: Performing speech recognition and sound source localization on audio signals collected in a live video broadcast scenario in which multiple speakers are speaking, where the multiple speakers are located in the same space; when localizing the sound source of the audio signals, performing the following processing on signal frames in the audio signals:
[0124] Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum;
[0125] Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0126] S620: Determine the occurrence location of the speaker change event based on the sound source localization results of the multiple signal frames, and separate the text obtained by speech recognition according to the occurrence location of the speaker change event;
[0127] S630: Adding the text segments to the video image collected in the live video scene according to the timeline information corresponding to the separated multiple text segments to generate a live video image with subtitles.
[0128] For the parts not described in detail in the above-mentioned embodiments 2 to 4, please refer to the description in the above-mentioned embodiment 1, and no further details will be given here.
[0129] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described herein within the scope permitted by applicable laws and regulations, subject to the requirements of applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).
[0130] Corresponding to the first embodiment, the present application also provides an audio signal processing device, see Figure 7 , the apparatus may include:
[0131] The recognition and positioning unit 710 is used to perform speech recognition and sound source localization on the audio signal collected in a multi-person speaking scenario. The recognition and positioning unit includes the following subunits when localizing the sound source of the audio signal:
[0132] The signal spectrogram processing subunit 711 is configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of the current signal frame and a target number of signal frames before and after it to form a matrix spectrogram, and perform smoothing on the matrix spectrogram;
[0133] The positioning result determination subunit 712 is configured to determine the sound source localization result of the current signal frame based on the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0134] The recognized text processing unit 720 is configured to determine the occurrence location of the speaker change event based on the sound source localization results of the plurality of signal frames, and to separate the text obtained by the speech recognition according to the occurrence location of the speaker change event.
[0135] In specific implementation, the signal spectrum processing unit can be used to:
[0136] For the current signal frame and a target number of signal frames before and after it, the corresponding kurtosis is calculated respectively; and the matrix spectrogram is smoothed using a target filter and the kurtosis information.
[0137] The text recognition processing unit may be specifically used for:
[0138] Determine a difference between a sound source localization result of the current signal frame and a sound source localization result of the previous signal frame; if the difference is greater than a target threshold, determine that a speaker change event has occurred, and determine the position of the current signal frame as the position where the speaker change event has occurred.
[0139] The target threshold may be dynamically set according to the spatial area associated with the multi-person speaking scene.
[0140] In a specific implementation, the device may further include:
[0141] The label adding unit is used to add labels to the separated text segments, wherein the labels are used to represent speaker identifications.
[0142] Specifically, the label adding unit can be used to:
[0143] Add the same label to different text segments corresponding to the same speaker.
[0144] In a specific implementation, the device may further include:
[0145] a statistical unit, configured to separate a plurality of time periods according to the occurrence location of the speaker change event, and to perform statistics on the sound source localization results of a plurality of signal frames in the same time period, and to determine an interval range of the sound source localization results in each time period;
[0146] A same-person judgment unit, configured to perform same-person judgment on the speakers corresponding to the multiple time periods based on the similarity of the interval ranges between different time periods;
[0147] The label adding unit can be used to:
[0148] The same label is added to the text segments corresponding to different time segments of the same speaker.
[0149] In addition, the device may further include:
[0150] an operation option providing unit, configured to provide an operation option for editing the label of the text segment;
[0151] The editing unit is configured to, after receiving an editing result on one of the text segments, determine other text segments corresponding to the same label as the text segment, and modify the labels of the other text segments to the editing result.
[0152] Furthermore, the device may further include:
[0153] The tag recommendation unit is used to provide recommended tags based on pre-acquired voice feature information of different speakers and their corresponding speaker identifiers.
[0154] The verification unit is used to verify the occurrence position of the speaker change event determined by sound source localization by extracting the speaker's timbre characteristics.
[0155] The multi-person speaking scenario includes a conference scenario where multiple speakers are present, wherein multiple speakers are located in the same space.
[0156] In addition, the audio signal may also be associated with a video signal;
[0157] At this time, the device may further include:
[0158] The subtitle adding unit is used to separate the text obtained by the speech recognition and add the subtitles to the associated video image according to the time axis information corresponding to the multiple separated text segments to generate a video image with subtitles.
[0159] The multi-person speaking scene includes a live video broadcast scene of multi-person speaking, in which multiple speakers are located in the same space.
[0160] In addition, an embodiment of the present application further provides a pickup, which may include the aforementioned audio signal processing device.
[0161] Corresponding to the second embodiment, the present application embodiment also provides a sound source localization device, see Figure 8 , the apparatus may include:
[0162] The audio signal determination unit 810 is configured to determine an audio signal to be processed;
[0163] a directional spectrogram determining unit 820, configured to obtain direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after the current signal frame in the audio signal;
[0164] A smoothing processing unit 830 is configured to perform smoothing processing on a matrix spectrogram consisting of direction of arrival spectrogram information of the signal frame and a target number of signal frames before and after the signal frame;
[0165] The unit result determination unit 840 is configured to determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame.
[0166] In addition, an embodiment of the present application further provides a microphone, which may include the aforementioned sound source localization device.
[0167] Corresponding to the third embodiment, the present application embodiment also provides a device for generating meeting records, see Figure 9 , the apparatus may include:
[0168] The recognition and positioning unit 910 is used to perform speech recognition and sound source localization on audio signals collected in a conference scenario where multiple people are speaking. The recognition and positioning unit includes the following subunits when localizing the sound source of the audio signals:
[0169] The signal spectrogram processing subunit 911 is configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of the current signal frame and a target number of signal frames before and after it to form a matrix spectrogram, and perform smoothing on the matrix spectrogram;
[0170] The positioning result determination subunit 912 is configured to determine the sound source localization result of the current signal frame based on the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0171] A recognition text processing unit 920 is configured to determine the location of a speaker change event based on the sound source localization results of the plurality of signal frames, and to separate the text obtained by speech recognition based on the location of the speaker change event;
[0172] The meeting record generating unit 930 is configured to generate a meeting record of the meeting according to the separated multiple text segments.
[0173] In a specific implementation, the device may further include:
[0174] The label adding unit is used to add labels to the separated text segments, wherein the labels are used to represent speaker identifications.
[0175] Specifically, the label adding unit can be used to:
[0176] Add the same label to different text segments corresponding to the same speaker.
[0177] In addition, the device may further include:
[0178] an operation option providing unit, configured to provide an operation option for editing the label of the text segment;
[0179] The editing unit is configured to, after receiving an editing result on one of the text segments, determine other text segments corresponding to the same label as the text segment, and modify the labels of the other text segments to the editing result.
[0180] Corresponding to the fourth embodiment, the present application embodiment also provides a live video processing device, see Figure 10 , the apparatus may include:
[0181] The recognition and positioning unit 1010 is used to perform speech recognition and sound source localization on audio signals collected in a live video broadcast scenario where multiple people are speaking. The recognition and positioning unit includes subunits when localizing the sound source of the audio signal:
[0182] The signal spectrogram processing subunit 1011 is configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of the current signal frame and a target number of signal frames before and after it, to form a matrix spectrogram, and perform smoothing on the matrix spectrogram;
[0183] The positioning result determination subunit 1012 is configured to determine the sound source localization result of the current signal frame based on the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame;
[0184] A recognized text processing unit 1020 is configured to determine the location of a speaker change event based on the sound source localization results of the plurality of signal frames, and to separate the text obtained by speech recognition based on the location of the speaker change event;
[0185] The subtitle adding unit 1030 is configured to add the text segments to the video image collected in the live video scene according to the time axis information corresponding to the separated multiple text segments, so as to generate a live video image with subtitles.
[0186] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0187] And an electronic device comprising:
[0188] one or more processors; and
[0189] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.
[0190] in, Figure 11 The electronic device architecture is shown as an example, and may include a processor 1110, a video display adapter 1111, a disk drive 1112, an input / output interface 1113, a network interface 1114, and a memory 1120. The processor 1110, the video display adapter 1111, the disk drive 1112, the input / output interface 1113, the network interface 1114, and the memory 1120 may be communicatively connected via a communication bus 1130.
[0191] Among them, the processor 1110 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in this application.
[0192] The memory 1120 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1120 can store an operating system 1121 for controlling the operation of the electronic device 1100, and a basic input and output system (BIOS) for controlling the low-level operations of the electronic device 1100. In addition, a web browser 1123, a data storage management system 1124, and an audio signal processing system 1125, etc. can also be stored. The above-mentioned audio signal processing system 1125 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided in the present application is implemented by software or firmware, the relevant program code is stored in the memory 1120 and is called and executed by the processor 1110.
[0193] The input / output interface 1113 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0194] The network interface 1114 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).
[0195] The bus 1130 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1110 , the video display adapter 1111 , the disk drive 1112 , the input / output interface 1113 , the network interface 1114 , and the memory 1120 ).
[0196] It should be noted that although the above device only shows the processor 1110, video display adapter 1111, disk drive 1112, input / output interface 1113, network interface 1114, memory 1120, bus 1130, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.
[0197] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0198] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0199] The above describes in detail the audio signal processing method, device, and electronic device provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is intended only to help understand the method and core concept of this application. At the same time, those skilled in the art will appreciate that variations in the specific implementation methods and scope of application may occur based on the concepts of this application. In summary, the contents of this specification should not be construed as limiting this application.
Claims
1. A method for processing an audio signal, characterized in that: include: Perform speech recognition and sound source localization on audio signals collected in a multi-person speaking scenario; when localizing the sound source of the audio signal, perform the following processing on a frame-by-frame basis: Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum; Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame; Determining the occurrence location of a speaker change event based on sound source localization results of multiple signal frames, and separating the text obtained by speech recognition based on the occurrence location of the speaker change event; The method further includes: verifying the occurrence location of the speaker change event determined by sound source localization by extracting the speaker's timbre features; extracting and comparing the timbre features of a preset number of signal frames before and after the occurrence location of the speaker change event; and adding a mark at the occurrence location when performing text separation if the comparison result indicates that the timbre before and after the occurrence location is the same, wherein the mark is used to remind editors to confirm the occurrence location.
2. The method according to claim 1, characterized in that The smoothing process of the matrix spectrum graph includes: For the current signal frame and a target number of signal frames before and after it, respectively calculate the corresponding kurtosis; The matrix spectrogram is smoothed using the target filter and kurtosis information.
3. The method according to claim 1, characterized in that Determining the occurrence location of the speaker change event according to the sound source localization result includes: Determine a difference between a sound source localization result of the current signal frame and a sound source localization result of the previous signal frame; If the difference is greater than the target threshold, it is determined that a speaker change event occurs, and the position of the current signal frame is determined as the position where the speaker change event occurs.
4. The method according to claim 3, characterized in that The target threshold is dynamically set according to the spatial area associated with the multi-person speaking scene.
5. The method according to claim 1, wherein Also includes: A label is added to the separated text segment, where the label is used to represent a speaker ID.
6. The method according to claim 5, characterized in that Adding labels to the separated text segments includes: Add the same label to different text segments corresponding to the same speaker.
7. The method according to claim 6, characterized in that Also includes: Separating multiple time periods according to the occurrence location of the speaker change event, and performing statistics on the sound source localization results of multiple signal frames in the same time period to determine the interval range of the sound source localization results in each time period; Performing peer judgment on the speakers corresponding to the multiple time periods based on the similarity of the interval ranges between different time periods; The method of adding the same label to different text segments corresponding to the same speaker includes: Add the same label to different text segments corresponding to the same speaker in different time periods.
8. The method according to claim 6, characterized in that Also includes: Providing an operation option for editing the label of the text segment; After receiving the editing result of one of the text segments, other text segments corresponding to the same label as the text segment are determined, and the labels of the other text segments are modified to the editing result.
9. The method according to claim 6, characterized in that Also includes: Recommended tags are provided based on the pre-acquired voice feature information of different speakers and their corresponding speaker identifiers.
10. The method according to any one of claims 1 to 8, characterized in that The multi-person speaking scenario includes a conference scenario where multiple people speak, wherein multiple speakers are located in the same space.
11. The method according to any one of claims 1 to 8, characterized in that The audio signal is also associated with a video signal; The method further comprises: After the text obtained by the speech recognition is segmented, time axis information corresponding to the multiple segmented text segments is added to the associated video image to generate a video image with subtitles.
12. The method according to claim 11, characterized in that The multi-person speaking scenario includes a live video broadcast scenario of multi-person speaking, wherein the multiple speakers are located in the same space.
13. A sound source localization method, characterized in that: include: determining an audio signal to be processed; Obtaining direction-of-arrival spectrogram information of a current signal frame and a target number of signal frames before and after the current signal frame in the audio signal; Smoothing a matrix spectrum composed of direction of arrival spectrum information of the signal frame and a target number of signal frames before and after it; Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame; It also includes: verifying the occurrence location of the speaker change event determined by sound source positioning by extracting the speaker's timbre features: extracting the timbre features of a preset number of signal frames before and after the occurrence location corresponding to the speaker change event and comparing them; when the comparison result indicates that the timbre before and after the occurrence location is the same, adding a mark at the occurrence location when performing text separation, wherein the mark is used to remind the editor to confirm the occurrence location.
14. A method for generating meeting minutes, characterized in that: include: Perform speech recognition and sound source localization on audio signals collected in a conference scenario where multiple people are speaking. When localizing the sound source of the audio signal, perform the following processing on a frame-by-frame basis: Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum; Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame; Determining the occurrence location of a speaker change event based on sound source localization results of multiple signal frames, and separating the text obtained by speech recognition based on the occurrence location of the speaker change event; generating a meeting record of the meeting based on the separated multiple text segments; It also includes: verifying the occurrence location of the speaker change event determined by sound source positioning by extracting the speaker's timbre features: extracting the timbre features of a preset number of signal frames before and after the occurrence location corresponding to the speaker change event and comparing them; when the comparison result indicates that the timbre before and after the occurrence location is the same, adding a mark at the occurrence location when performing text separation, wherein the mark is used to remind the editor to confirm the occurrence location.
15. The method according to claim 14, characterized in that Also includes: Labels are added to the separated multiple text segments, where the labels are used to represent speaker identifications.
16. The method according to claim 15, characterized in that Add the same label to different text segments corresponding to the same speaker.
17. The method according to claim 16, characterized in that: Also includes: Providing an operation option for editing the label of the text segment; After receiving the editing result of one of the text segments, other text segments corresponding to the same label as the text segment are determined, and the labels of the other text segments are modified to the editing result, where the editing result corresponds to the identity information of the speaker.
18. A live video processing method, characterized in that: include: Perform speech recognition and sound source localization on audio signals collected from a live video broadcast of multiple speakers in the same space. The following processing is performed on each frame of the audio signal during sound source localization: Obtaining direction of arrival spectrum information of the current signal frame and the target number of signal frames before and after it to form a matrix spectrum, and performing smoothing processing on the matrix spectrum; Determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame; Determining the occurrence location of a speaker change event based on sound source localization results of multiple signal frames, and separating the text obtained by speech recognition based on the occurrence location of the speaker change event; Adding the text segments to a video image captured in the live video scene according to timeline information corresponding to the separated multiple text segments to generate a live video image with subtitles; It also includes: verifying the occurrence location of the speaker change event determined by sound source positioning by extracting the speaker's timbre features: extracting the timbre features of a preset number of signal frames before and after the occurrence location corresponding to the speaker change event and comparing them; when the comparison result indicates that the timbre before and after the occurrence location is the same, adding a mark at the occurrence location when performing text separation, wherein the mark is used to remind the editor to confirm the occurrence location.
19. An audio signal processing device, characterized in that: include: The recognition and positioning unit is used to perform speech recognition and sound source localization on the audio signal collected in a multi-person speaking scenario; wherein, when the recognition and positioning unit localizes the sound source of the audio signal, it includes the following subunits: a signal spectrogram processing subunit, configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after it, to form a matrix spectrogram, and to perform smoothing on the matrix spectrogram; A positioning result determination subunit is used to determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame; The text recognition processing unit is used to determine the occurrence location of the speaker change event based on the sound source localization results of multiple signal frames, and to separate the text obtained by speech recognition according to the occurrence location of the speaker change event. It is also used to verify the occurrence location of the speaker change event determined by sound source localization by extracting the speaker's timbre features: extracting the timbre features of a preset number of signal frames before and after the occurrence location of the speaker change event and comparing them; when the comparison result indicates that the timbre before and after the occurrence location is the same, adding a mark to the occurrence location when performing text separation, wherein the mark is used to remind the editor to confirm the occurrence location.
20. A pickup, characterized in that: Comprising the audio signal processing device as claimed in claim 19.
21. A sound source localization device, characterized in that: include: an audio signal determining unit, configured to determine an audio signal to be processed; a directional spectrogram determining unit, configured to obtain direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after the current signal frame in the audio signal; a smoothing processing unit, configured to perform smoothing processing on a matrix spectrogram composed of direction of arrival spectrogram information of the signal frame and a target number of signal frames before and after the signal frame; a unit result determination unit for determining a sound source localization result for the current signal frame based on the angle corresponding to the value that satisfies the target condition in the smoothed direction of arrival spectrogram corresponding to the current signal frame, and for verifying the location of the speaker change event determined by sound source localization by extracting the speaker's timbre features: extracting and comparing the timbre features of a preset number of signal frames before and after the location of the speaker change event; When the comparison result indicates that the timbre before and after the occurrence position is the same, a mark is added to the occurrence position when performing text separation, wherein the mark is used to remind the editor to confirm the occurrence position.
22. A pickup, characterized in that: Comprising the sound source localization device as claimed in claim 21.
23. A device for generating meeting minutes, characterized in that: include: The recognition and positioning unit is used to perform speech recognition and sound source localization on audio signals collected in a conference scene where multiple people are speaking; wherein the recognition and positioning unit includes subunits when localizing the sound source of the audio signal: a signal spectrogram processing subunit, configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after it, to form a matrix spectrogram, and to perform smoothing on the matrix spectrogram; A positioning result determination subunit is used to determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame; A text recognition processing unit is configured to determine the location of a speaker change event based on sound source localization results of multiple signal frames, separate text obtained by speech recognition based on the location of the speaker change event, and verify the location of the speaker change event determined by sound source localization by extracting speaker timbre features: extracting timbre features of a preset number of signal frames before and after the location of the speaker change event and comparing them; if the comparison result indicates that the timbre before and after the location of the speaker change event is the same, adding a mark to the location of the event when performing text separation, wherein the mark is used to remind editors to confirm the location of the event; The meeting record generating unit is used to generate the meeting record of the meeting according to the multiple separated text segments.
24. A live video processing device, characterized in that: include: The recognition and positioning unit is used to perform speech recognition and sound source positioning on the audio signal collected in the live video broadcast scene of multiple people speaking, wherein the recognition and positioning unit includes subunits when positioning the sound source of the audio signal: a signal spectrogram processing subunit, configured to obtain, based on signal frames in the audio signal, direction of arrival spectrogram information of a current signal frame and a target number of signal frames before and after it, to form a matrix spectrogram, and to perform smoothing on the matrix spectrogram; A positioning result determination subunit is used to determine the sound source localization result of the current signal frame according to the angle corresponding to the value that meets the target condition in the smoothed direction of arrival spectrum corresponding to the current signal frame; A text recognition processing unit is configured to determine the location of a speaker change event based on sound source localization results of multiple signal frames, separate text obtained by speech recognition based on the location of the speaker change event, and verify the location of the speaker change event determined by sound source localization by extracting speaker timbre features: extracting timbre features of a preset number of signal frames before and after the location of the speaker change event and comparing them; if the comparison result indicates that the timbre before and after the location of the speaker change event is the same, adding a mark to the location of the event when performing text separation, wherein the mark is used to remind editors to confirm the location of the event; The subtitle adding unit is used to add the text segments to the video image collected in the live video scene according to the time axis information corresponding to the separated multiple text segments, so as to generate a live video image with subtitles.
25. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 18 are implemented.
26. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Voice processing method, device and system
CN111145753A
Sound signal separation method of double sound sources and sound pickup
CN111429939A