Spectrogram processing method, device, electronic device and storage medium
By determining the time position information of the specified content in the speech recognition results, displaying multiple segmented spectrograms on the same canvas, and using independently adjusted display window parameters and audio track splicing and playback, the problem of cumbersome operations in spectrogram analysis is solved, and the efficiency of voiceprint identification and user experience are improved.
Patent Information
- Application Number
- CN202111529778.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-12-14
AI Technical Summary
During the voiceprint identification process, the existing technology for segmenting and analyzing the same segments of the spectrogram is cumbersome and has low comparison efficiency, especially for the frequent switching display problem caused by the scattered distribution of segments in audio with longer duration.
By searching for specified content in the speech recognition results of the target audio, determining its time position information, and displaying multiple segmented spectrograms in the same canvas, the display effect is optimized using independently adjusted display window parameters, supporting splicing playback in audio tracks and marker display in the target audio.
It simplifies the comparison operation of spectrograms, improves user experience, reduces the need for unified adjustment of display parameters, makes it easier for users to analyze and compare multiple discontinuous segmented spectrograms, and improves the efficiency of voiceprint identification.
Smart Images

Figure CN114360585B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing, and more specifically, to a method, device, electronic device and storage medium for processing a spectrogram. Background Art
[0002] During the voiceprint identification process, the identification personnel need to compare and analyze the same sound segments in the same audio or different audio in the spectrogram to select the sound segments with high stability as the feature segments.
[0003] In the prior art, when analyzing the segmented spectrograms of the same sound segment in a spectrogram, the appraiser needs to frequently switch back and forth to display the segmented spectrograms corresponding to the same sound segment at different positions in a spectrogram, and perform repeated spectrogram comparisons thereon. For audio with a long duration, the distribution of the same sound segment may be relatively scattered, and the time span is relatively large. Therefore, switching back and forth to display the spectrograms in this way results in cumbersome operations and low comparison efficiency. Summary of the Invention
[0004] In view of the above problems, the embodiments of the present application propose a method, device, electronic device and storage medium for processing a spectrogram to improve the above problems.
[0005] In a first aspect, an embodiment of the present application provides a method for processing a spectrogram, the method comprising: searching for specified content in the speech recognition results of the target audio, and determining the time position information of the sound segment corresponding to the specified content in the target audio; the specified content includes a specified phoneme or a specified text; based on the time position information, determining at least two segmented spectrograms corresponding to the candidate sound segment in the spectrogram of the target audio, the candidate sound segment being the sound segment corresponding to the specified content in the target audio; and displaying the at least two segmented spectrograms in the same canvas, wherein a plurality of display windows are provided in the canvas, each display window being used to display one of the segmented spectrograms, and the display parameters corresponding to each of the display windows can be adjusted individually.
[0006] In a second aspect, an embodiment of the present application provides a spectrogram processing device, comprising: a search module for searching for specified content in the speech recognition results of the target audio, and determining the time position information of the sound segment corresponding to the specified content in the target audio; the specified content includes a specified phoneme or a specified text. A determination module for determining at least two segmented spectrograms corresponding to a candidate sound segment in the spectrogram of the target audio based on the time position information, wherein the candidate sound segment is a sound segment corresponding to the specified content in the target audio. A display module for displaying the at least two segmented spectrograms in the same canvas, wherein a plurality of display windows are provided in the canvas, each display window being used to display one of the segmented spectrograms, and the display parameters corresponding to each of the display windows can be adjusted separately.
[0007] In some embodiments, the spectrogram processing device further includes a track display unit for displaying the at least two segmented spectrograms in a same track so as to continuously play the candidate sound segments corresponding to the at least two segmented spectrograms.
[0008] In some embodiments, the audio track display unit is further configured to splice and display the at least two segmented spectrograms on the same audio track in a time-ordered order based on the time position information corresponding to the at least two segmented spectrograms.
[0009] In some embodiments, the spectrogram processing apparatus further comprises a marking display unit for marking and displaying the segmented spectrogram in the spectrogram of the target audio.
[0010] In some embodiments, the spectrogram processing apparatus further comprises: a detection unit for detecting a selection operation triggered on the marked displayed segmented spectrogram; and a canvas display unit for displaying each selected segmented spectrogram in a display window on the canvas.
[0011] In some embodiments, the spectrogram processing device includes a speech recognition unit for performing speech recognition on the target audio to obtain a speech recognition result of the target audio.
[0012] In some embodiments, the speech recognition unit further includes an active speech detection unit configured to perform active speech detection on the target audio and determine the active speech in the target audio. The speech recognition unit is further configured to perform speech recognition on the active speech in the target audio and obtain a speech recognition result for the target audio.
[0013] In some embodiments, the spectrogram processing device also includes a sorting module for sorting the at least two segmented spectrograms according to preset rules to determine the sorting of the segmented spectrograms; in this embodiment, the display module is further configured to: sort the segmented spectrograms and sequentially display the at least two segmented spectrograms in multiple display windows in the same canvas.
[0014] In some embodiments, the sorting module includes a sorting unit for sorting the at least two segmented spectrograms in descending order of speech stability to determine the segmented spectrogram sorting.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the spectrogram processing method as described above is implemented.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the spectrogram processing method described above is implemented.
[0017] In the solution of the present application, at least two segmented spectrograms corresponding to discontinuous specified content are displayed in the same canvas, so that the user does not need to drag the spectrogram back and forth in the spectrogram of the target audio to compare the segmented spectrograms. At the same time, since the display parameters of each display window can be adjusted separately, adjusting the display parameters of one display window does not affect the segmented spectrograms displayed in other display windows, thereby making it easier for the user to adjust the display parameters corresponding to a certain display window to make the display effect of the displayed segmented spectrogram better, and making it easier for the user to view the details in the segmented spectrogram. Therefore, this solution can solve the problem of cumbersome operations caused by the need to frequently switch in the same spectrogram in the prior art, simplify the operation, and facilitate the user to compare and analyze multiple discontinuous segmented spectrograms.
[0018] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0020] Figure 1AIt is a flowchart of a method for processing a spectrogram according to an exemplary embodiment of the present application.
[0021] Figure 1B FIG. 1 is a flowchart illustrating a method for searching for specified content according to an embodiment of the present invention.
[0022] Figure 2 This is a schematic diagram showing multiple segmented spectrograms displayed on the same canvas according to an embodiment of the present application.
[0023] Figure 3 This is a schematic diagram showing the splicing and display of multiple segmented spectrograms in the same audio track according to an embodiment of the present application.
[0024] Figure 4 This is a flowchart of an embodiment of the present application showing the process of marking and displaying a segmented spectrogram in a spectrogram of a target audio.
[0025] Figure 5 It is a block diagram of a spectrogram processing device according to an exemplary embodiment of the present application.
[0026] Figure 6 FIG. 1 is a hardware structure diagram of an electronic device according to an exemplary embodiment of the present application.
[0027] The above-mentioned drawings have shown clear embodiments of the present invention, which will be described in more detail later. These drawings and textual descriptions are not intended to limit the scope of the present invention in any way, but to illustrate the concept of the present invention to computer technicians in this field through specific embodiments. DETAILED DESCRIPTION
[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0029] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0030] Figure 1AFIG. 0 is a schematic flowchart of a method for processing a spectrogram according to an embodiment of the present application. This method can be executed by an electronic device with computing and processing capabilities, such as a desktop computer, a laptop computer, or other terminal devices. This method can also be interactively executed by a processing system including a server and a terminal. As Figure 1A shown, the method includes the following steps:
[0031] Step 110: Search for specified content in the speech recognition result of the target audio, and determine the time position information of the segment corresponding to the specified content in the target audio; the specified content includes specified phonemes or specified text.
[0032] The target audio can be a continuous audio segment, a long-time audio, or a short-time audio, or it can be an audio composed of multiple audio segments pieced together. There is no limitation here. The number of target audios can be one or multiple.
[0033] A phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzing based on the pronunciation actions in a syllable, one action constitutes one phoneme. Phonemes are divided into two major categories: vowels and consonants. A phoneme is the smallest unit or the smallest speech segment that constitutes a syllable, and it is also a specific physical phenomenon. For example, the Chinese syllable "ā" has only one phoneme "ā"; "ài" has two phonemes, namely "à" and "i"; "dài" has three phonemes, namely "d", "à", and "i". It can be understood that due to different language pronunciations, the pronunciation of the same letter is different. Therefore, the phoneme can be a phoneme in any language. For example, it can be a Chinese phoneme or an English phoneme. There is no limitation here.
[0034] The specified text can be a character, a word, or a phrase, etc. There is no limitation here.
[0035] In some embodiments, before step 110, the method further includes: performing speech recognition on the target audio to obtain the speech recognition result of the target audio.
[0036] In some embodiments, the speech recognition performed can be to recognize the phoneme content corresponding to the target audio. Thus, the speech recognition result of the target audio includes the phoneme content corresponding to the target audio.
[0037] In some embodiments, the speech recognition performed can be to recognize the text content corresponding to the target audio. Thus, the speech recognition result of the target audio includes the text content corresponding to the target audio.
[0038] In some embodiments, the speech recognition performed can include recognizing the phoneme content corresponding to the target audio and recognizing the text content corresponding to the target audio. Correspondingly, the speech recognition result of the target audio includes the phoneme content corresponding to the target audio and the corresponding text content.
[0039] In the solution of the present application, during the speech recognition process of the target audio, not only the phoneme content and / or text content corresponding to the target audio are determined, but also the time position information of each phoneme segment corresponding to each phoneme included in the phoneme content corresponding to the target audio in the target audio is correspondingly determined. Similarly, the time position information of each phoneme segment corresponding to each character (or word or phrase) included in the text content corresponding to the target audio in the target audio is also correspondingly determined.
[0040] In the solution of the present application, searching for specified content in the speech recognition result of the target audio is to locate the specified content in the speech recognition result of the target audio, and further obtain the time position information associated with the located specified content. The time position information associated with the located specified content is the time position information of the phoneme segment corresponding to the specified content in the target audio.
[0041] For example, if the text content corresponding to the target audio is "Today is Sunday and the weather is fine", and if the specified content is the character "天", then searching for the character "天" in the text content corresponding to the target audio can determine the positions of the 3 characters "天" included in the text content corresponding to the target audio. Since the speech recognition result of the target audio further indicates the time position information of each character (or word), thus, after locating the character "天" in the text content corresponding to the target audio, the time position information associated with this character "天" can be correspondingly obtained. The time position information associated with this character "天" is the time position information of the pronunciation segment of the character "天" (i.e., the phoneme segment corresponding to the character "天") in the target audio.
[0042] In the solution of the present application, the speech recognition result of the target audio includes multiple groups of specified content, and the time position information determined through the specified content search is also multiple groups. <==
[0043] The time position information of the phoneme segment corresponding to the specified content indicates the start time and end time of the phoneme segment corresponding to the specified content in the target audio. Since the target audio is continuous in time, therefore, the start time and end time of the phoneme segment in the target audio determine the position of the phoneme segment in the target audio.
[0044] In some embodiments, before step 110, the method further includes: performing active speech detection on the target audio to determine the active speech in the target audio; in this embodiment, the step of performing speech recognition on the target audio to obtain the speech recognition result of the target audio further includes: performing speech recognition on the active speech in the target audio to obtain the speech recognition result of the target audio.
[0045] Voice Activity Detection (VAD), also known as endpoint detection, distinguishes speech segments from non-speech segments (also known as silence) within an audio stream, removing the silence and retaining the speech segments. Therefore, after performing voice activity detection on the target audio, the non-speech segments can be filtered out, while the active speech (i.e., speech segments) in the target audio are retained. During speech recognition, only the active speech in the target audio is recognized, without focusing on the non-speech segments.
[0046] Figure 1B This is a flow chart showing a method for searching for specified content according to an embodiment of the present application. Figure 1B As shown, first obtain the speech recognition result of the target audio through the following steps:
[0047] Step 141: Input target audio, which may be a speech signal in the time domain.
[0048] Step 142: Active speech detection. Thus, the non-speech segments in the target audio are determined through active speech detection, and the non-speech segments are filtered out from the target audio, while the speech segments in the target audio are retained.
[0049] Step 143: Speech recognition: In this step, speech recognition is performed on the speech segment in the target audio.
[0050] Step 144: Output the speech recognition results. The output speech recognition results may include the text content corresponding to the target audio and / or the phoneme content corresponding to the target audio. Furthermore, they may include the temporal position information of the audio segments corresponding to each character or word in the text content within the target audio, and / or the temporal position information of the audio segments corresponding to each phoneme in the phoneme content within the target audio.
[0051] Afterwards, a designated content search is performed in the speech recognition result of the target audio to locate the designated content included in the speech recognition result of the target audio, as well as the time position information of the audio segment corresponding to the designated content in the target audio.
[0052] In the field of voiceprint identification, in order to ensure the accuracy of voiceprint identification, it is necessary to query segments with high feature stability from the audio as feature segments. For a piece of content (such as the character "tian" listed above) in one audio or multiple audios, there may be multiple corresponding audio segments. If one or more audio segments are selected from the audio segments corresponding to a piece of content as feature segments, then this piece of content is the specified content in this solution. Therefore, first search for the specified content from the speech recognition result of the target audio through step 110 above to obtain the time position information associated with the specified content included in the speech recognition result of the target audio. Thus, the position of each audio segment corresponding to the specified content in the target audio can be determined.
[0053] Please continue to refer to Figure 1A , step 120, according to the time position information, determine at least two segmented spectrograms corresponding to the candidate segment in the spectrogram of the target audio, where the candidate segment is the segment corresponding to the specified content in the target audio.
[0054] A spectrogram is a graph representing the variation of the speech spectrum over time. Its vertical axis is frequency, the horizontal axis is time, and the strength of any given frequency component at a given moment is represented by the gray level or the shade of the corresponding point. The darker the color, the stronger the speech energy at that point, and vice versa. Spectrograms are divided into narrowband spectrograms and broadband spectrograms. In some embodiments, the spectrogram of the target audio can be a broadband spectrogram, which can clearly display the formant structure and spectral envelope, and can reflect the rapid time-varying process of the spectrum, so a higher time resolution can be obtained on the broadband spectrogram. In other embodiments, the spectrogram of the target audio can also be a narrowband spectrogram.
[0055] The spectrogram also describes the frequency domain characteristics of each speech frame in chronological order, such as frequency and speech energy, etc. That is, the spectrogram is also time-related. Therefore, based on the time position information corresponding to the specified content obtained in step 110, it can be located correspondingly in the spectrogram of the target audio. The partial spectrogram located in the spectrogram of the target audio based on this time position information is the spectrogram of the segment corresponding to the specified content included in the target audio (i.e., the candidate segment). That is to say, in this application, the segmented spectrogram refers to the spectrogram of the candidate segment.
[0056] As described above, in step 110, at least two sets of time position information are obtained. Thus, in step 120, at least two sets of segmented spectrograms corresponding to the candidate segment are also determined in the spectrogram of the target audio. It can be understood that the number of candidate segments is determined by the number of specified contents included in the speech recognition result corresponding to the target audio.
[0057] Step 130, display at least two segmented spectrograms in the same canvas. Multiple display windows are provided in the canvas, and each display window is used to display a segmented spectrogram. The display parameters corresponding to each display window can be adjusted separately.
[0058] The canvas can be used to manage multiple graphic objects. The display windows in the canvas are used to display image objects. In this solution, the image object can be a segmented spectrogram.
[0059] The display parameters can be the size, contrast, resolution, pixel pitch, chromaticity, background color, color mode for spectrogram display, etc., which are not specifically limited here. It can be understood that when the display parameters corresponding to the display window change, the display effect of the segmented spectrogram displayed in this display window changes accordingly. Therefore, the display effect of the segmented spectrogram in the display window can be adjusted by adjusting the display parameters corresponding to the display window. Moreover, there are differences in the density of formants and the energy of the signal of the voice signals expressed by different segmented spectrograms. Therefore, there are also differences in the display parameters suitable for different segmented spectrograms. Thus, the user can determine the display parameters suitable for the displayed segmented spectrogram by adjusting the display parameters corresponding to the display window, so as to ensure the display effect of the segmented spectrogram in the display window.
[0060] In some embodiments, the number of display windows in a canvas can be preset. In some embodiments, the display windows in the canvas can be newly created according to the user's trigger operation. That is, the user can trigger a window creation operation in the display interface of the canvas, and then a newly added display window will be displayed in the canvas accordingly.
[0061] In some embodiments, the display windows in the canvas can be automatically allocated to the segmented spectrograms corresponding to each time-specified content in sequence according to the chronological order of the appearance of the candidate segments corresponding to the specified content in the target audio.
[0062] For example, if the candidate segment corresponding to “Yes 1” is the segment corresponding to the first “Yes” character in the target audio, then the display window numbered 1 in the canvas is allocated to the segmented spectrogram of the candidate audio corresponding to “Yes 1”; correspondingly, if the candidate segment corresponding to “Yes 2” is the segment corresponding to the second “Yes” character in the target audio, then the display window numbered 2 in the canvas can be allocated to the segmented spectrogram corresponding to “Yes 2”.
[0063] In some embodiments, the display window of the canvas can also be selected for the segmented spectrogram corresponding to the specified content according to the user's trigger operation.
[0064] In some embodiments, the user can also select multiple segmented spectrograms from the at least two segmented spectrograms corresponding to the candidate segments to display on the canvas. Of course, in other embodiments, the at least two segmented spectrograms corresponding to the candidate segments can also be displayed in the display window of the same canvas.
[0065] In the prior art, the display effects of the spectrograms displayed in the canvas can only be adjusted uniformly, that is, uniformly zoomed in or out, etc. Therefore, if the user wants to zoom in on a certain segmented spectrogram, all the spectrograms must be zoomed in, and then the display parameters need to be readjusted to the parameters before adjustment. Since the display parameters of the spectrogram can only be adjusted uniformly, if the spectrogram includes segmented spectrograms of multiple characteristic sound segments, the user needs to repeatedly adjust the display parameters if he needs to analyze the multiple segmented spectrograms. A certain set of display parameters may only be compatible with a certain segmented spectrogram, but not suitable for other segmented spectrograms, resulting in the user having to repeatedly switch between multiple display parameters. The operation is cumbersome and inconvenient for the user to compare segmented spectrograms.
[0066] In this solution, multiple display windows are provided in the canvas, and the display parameters corresponding to each display window can be adjusted separately. Thus, the user can adjust the display effect of the segmented spectrogram in the display window by adjusting the display parameters corresponding to the display window. Thus, the user can determine the display parameters that are adapted to the segmented spectrogram and adjust the display effect of the segmented spectrogram as needed, so that the user can view the details in the segmented spectrogram, such as the distribution position of the resonance peak in the segmented spectrogram, the trend of the resonance peak, the fundamental frequency, the center frequency, the LPC (Linear Predictive Coding) spectrum, etc.
[0067] Since the display parameters of each display window in the canvas can be adjusted independently, that is, if the display parameters of one display window are adjusted, it will not affect the segmented spectrograms displayed in other display windows. Therefore, it can solve the problem in the prior art that the display parameters can only be adjusted uniformly, resulting in the user having to repeatedly switch between multiple sets of display parameters.
[0068] In some embodiments, based on the segmented spectrograms displayed in multiple display windows, since the user can adjust the display effect of the segmented spectrogram by adjusting the display parameters, it is convenient for the user to observe the details of the segmented spectrogram, and then it is convenient for the user to compare the segmented spectrograms in different display windows, so as to determine the segmented spectrogram with stable features, and then determine the candidate segment corresponding to the segmented spectrogram with stable features as the feature segment.
[0069] Figure 2 This is a schematic diagram showing multiple segmented spectrograms in the same canvas according to an embodiment of the present application. Figure 2As shown, in the canvas 200, eight display windows are shown. Each display window shows the segmented spectrogram of a candidate segment corresponding to the character "shi". For the convenience of distinction, each candidate audio segment is named in sequence as: shi 1, shi 2, shi 3,... shi 8.
[0070] Please continue to refer to Figure 2 , in the display interface of the canvas, there are a window creation control 210 and a window deletion control 220. The user can trigger the window creation control 210 to add a display window to the canvas and trigger the window deletion control to delete a display window in the canvas.
[0071] In the canvas, the display parameters of each display window can be adjusted separately to adjust the display effect of the segmented spectrogram in the display window as needed. Among them, the display parameters can include contrast, resolution, pixel pitch, color rendering, etc., which are not limited here.
[0072] The display interface of the canvas provides a mode selection option, and this mode selection option includes a small window mode and a large window mode. Figure 2 shows a schematic diagram of the display window (and the segmented spectrogram shown in the display window) in the small window mode.
[0073] In the large window mode, the size of the canvas remains unchanged, and each display window and the displayed segmented spectrogram will be enlarged. The segmented spectrogram in the large window mode will be shown larger than that in the small window mode. Correspondingly, the number of display windows that can be displayed simultaneously in the display interface of a client will be reduced. In a specific embodiment, the user can select the large window mode or the small window mode according to actual needs. <�
[0074] Figure 2 The display interface of the shown canvas has a sorting mode selection option. Figure 2 The sorting mode selection option in is the increasing mode. In the increasing mode, the segmented spectrograms shown in each display window are sorted in the increasing order of time position from the earliest to the latest. In other embodiments, the decreasing mode can also be selected. In the decreasing mode, the segmented spectrograms are displayed in the display window in the order from the latest to the earliest in terms of time position.
[0075] Please continue to refer to Figure 2 , in the display interface of the canvas, there are also labels for marking, such as Figure 2Four labels, F1, F2, F3, and F4, in it. The label can be a formant label, that is, the formant label can be marked in the displayed segmented spectrogram to indicate the distribution position of the formant shown in the segmented spectrogram through the formant label. Among them, the four labels F1, F2, F3, and F4 can be used to mark different formants in a segmented spectrogram. For example, F1 represents the first formant, F2 represents the second formant, F3 represents the third formant, and F4 represents the fourth formant.
[0076] In each display window, a playback control 230 is provided. The user can trigger and operate the playback control 230 to play the candidate segment corresponding to the current segmented spectrogram, which is convenient for the user to perform listening and discrimination analysis. Further, a first paste control 240 is also provided in the display window. The user can trigger the paste control to paste the segmented spectrogram displayed in the display window.
[0077] In some embodiments, the center frequency of the formant of the spectrogram can be used for candidate audio comparison and analysis. Generally, 4 formants will be selected for comparative analysis. Figure 2 It can be seen that: the center frequency of the formant of the segmented spectrogram named "Yes 6" displayed in the display window (6) in the canvas is the most stable. The center frequencies of the formants of the segmented spectrograms shown in the display windows (2), (4), and (5) are similar to those of the segmented spectrogram shown in the display window (6). The center frequencies of the formants of the segmented spectrograms shown in the display windows (1) and (3) are very different from those of the segmented spectrogram shown in the display window (6). And in the segmented spectrograms shown in the display windows (7) and (8), three formants are similar to those in the display window (6), but one formant is significantly different from the formant of the segmented spectrogram shown in the display window (6). Therefore, through comprehensive comparison, it can be determined that: the differences between the segmented spectrograms shown in the display windows (1), (3), (7), and (8) and the segmented spectrogram shown in the display window (6) are relatively large, and the differences between the segmented spectrograms shown in the display windows (2), (4), and (5) and the segmented spectrogram shown in the display window (6) are relatively small, and the feature stability is good. Therefore, the segments corresponding to the segmented spectrograms shown in the display windows (2), (4)-(6) can be used as feature segments for subsequent voiceprint identification.
[0078] In some embodiments, after step 120, the method further includes: sorting the at least two segmented spectrograms according to a preset rule to determine the segmented spectrogram sorting; in this embodiment, step 130 includes: sequentially displaying the at least two segmented spectrograms in multiple display windows in the same canvas according to the segmented spectrogram sorting.
[0079] In some embodiments, the at least two segmented spectrograms may be sorted in descending order of speech stability to determine the segmented spectrogram sorting.
[0080] For speech signals, the more concentrated the audio energy is near the formant, the higher the stability of the signal. This concentration of audio energy near the formant is reflected in the spectrogram: within a certain range of the formant frequency, the closer the position is to the formant, the darker the color. Conversely, the farther the position is from the formant, the lighter the color. In other words, if a speech signal's audio energy is concentrated near the formant, the color near that formant on the spectrogram will be darker.
[0081] Furthermore, in the grayscale image, the grayscale value corresponding to the color ranges from 0 to 255, where white is 255 and black is 0; therefore, if the segmented spectrogram is converted into a grayscale image, the grayscale value of the darker the color, the smaller the grayscale value, and the grayscale value of the lighter the color, the larger the grayscale value. In this embodiment, in order to compare the speech stability corresponding to all candidate segments in the target speech, the segmented spectrogram corresponding to each candidate segment can be compared. Specifically, the position of the formant frequency can be determined in the corresponding segmented spectrogram based on the formant frequency corresponding to each candidate segment. Then, using the position of the formant frequency in the segmented spectrogram as a reference, the spectrogram area near the formant frequency and the position of the formant frequency in the segmented spectrogram is intercepted as a reference area; then, the average grayscale value of the reference area is calculated. Among them, the first length can be set according to actual needs and is not specifically limited here. It can be understood that the first length is less than the distance between two adjacent formant frequencies in the spectrogram.
[0082] As described above, the lighter the color of a region, the greater the grayscale value. Furthermore, lighter regions indicate a lower concentration of audio energy in that region, i.e., lower speech stability. Therefore, all segmented spectrograms can be sorted in ascending order based on the average grayscale values of the corresponding reference regions. The resulting sorting can be considered the same as sorting the at least two segmented spectrograms in descending order of speech stability.
[0083] It is worth mentioning that a speech signal generally includes four formant frequencies. Therefore, in the process of determining the reference area for each candidate segment, a reference area is determined for each formant frequency, and the average grayscale value is calculated for each reference area. The average grayscale values of all reference areas can then be weighted and summed, and the weighted result is used as the target average grayscale value for sorting. Afterwards, all segmented spectrograms are sorted in ascending order according to the corresponding target average grayscale values to obtain the segmented spectrogram sorting.
[0084] In this embodiment, the at least two segmented spectrograms are sorted and displayed in descending order of speech stability, so that the user can first focus on the candidate segments with higher speech stability.
[0085] In some embodiments, at least two segmented spectrograms can also be sorted and displayed in different display windows in the same canvas according to the time corresponding to the candidate sound segments corresponding to the segmented spectrograms in the target audio, in order from first to last (or from last to first).
[0086] In some embodiments, after determining the order of speech stability from high to low, the segmented spectrogram corresponding to the candidate segment with the highest speech stability can be determined, and then the segmented spectrogram corresponding to the candidate segment with the highest speech stability can be used as the reference segmented spectrogram, and the spectrogram similarity between each other segmented spectrogram and the reference segmented spectrogram can be calculated. After determining that the reference segmented spectrogram is displayed in the first display window on the canvas, the corresponding segmented spectrograms are displayed in other display windows behind the display window where the reference segmented spectrogram is located in order of spectrogram similarity from large to small.
[0087] In some embodiments, the preset rule may be selected by the user, may be a default, or may be a sorting rule that is most frequently used by the user, or a priority sorting rule set by the user, which is not specifically limited here. The preset rule may be, for example, sorting by time, sorting by voice stability, sorting by spectrogram similarity with the reference segmented spectrogram after determining the reference segmented spectrogram, etc., which is not specifically limited here. In some embodiments, after step 120, the method further includes: splicing and displaying at least two segmented spectrograms in the same audio track to continuously play the candidate segments corresponding to the at least two segmented spectrograms.
[0088] In some embodiments, the at least two segmented spectrograms may be displayed together on the same track in a time-ordered order based on the time position information corresponding to the at least two segmented spectrograms. In other embodiments, the at least two segmented spectrograms may be displayed together on the same track in a user-selected arrangement order.
[0089] In some embodiments, in the process of splicing and displaying at least two segmented spectrograms in the same audio track, the candidate segments corresponding to each of the at least two segmented spectrograms are added to the playback file corresponding to the audio track, and according to the arrangement order of the segmented spectrograms in the audio track, the candidate segments corresponding to the at least two segmented spectrograms are spliced and combined in the playback file corresponding to the audio track. Therefore, after the playback operation of the audio track is started, the candidate segments corresponding to the at least two segmented spectrograms displayed in the audio track can be played continuously.
[0090] In the prior art, in order to locate the segmented spectrogram corresponding to the specified content in the spectrogram of the target audio displayed in the audio track, so as to play the candidate segments corresponding to the specified content in the audio track, since the candidate segments corresponding to the specified content are not continuous in the target audio, the user needs to switch positions back and forth in the audio track, which is cumbersome and inconvenient for the user to perform listening analysis.
[0091] In this solution, at least two segmented spectrograms corresponding to the specified content are spliced and displayed in the same audio track, and multiple segmented spectrograms corresponding to the specified content that the user needs to pay attention to are spliced and displayed. Therefore, the user does not need to locate the segmented spectrogram back and forth in the audio track, but can directly play multiple candidate audio segments corresponding to the specified content continuously in the audio track, which is convenient for the user to concentrate on listening and analysis.
[0092] Figure 3 This is a schematic diagram showing a method of displaying multiple segmented spectrograms in the same audio track according to an embodiment of the present application. Figure 3 As shown, in the audio track 300, the ordinate represents the frequency and the abscissa represents the start time of the segmented spectrogram corresponding to the segment. In the audio track 300, the duration of the entire audio track and the time position of the currently playing audio track content are also displayed.
[0093] Please continue reading Figure 3 , a second paste control 310 is provided on the audio track, and the user can trigger the second paste control 310 to paste the segmented spectrogram displayed in the current audio track; the user can drag or click the audio track to select the audio segment to be played.
[0094] In some embodiments, after step 120, the method further includes: marking and displaying the segmented spectrogram in the spectrogram of the target audio.
[0095] The segmented spectrograms corresponding to the candidate segments are marked and displayed in the target audio spectrogram, making it easier for users to find the specific location of the segment corresponding to a segmented spectrogram in the target audio. For example, when the specified content is "yes", the time position information corresponding to all "yes" is marked and displayed in the target audio spectrogram.
[0096] The marking display can be to select and display the segmented spectrogram corresponding to each candidate segment in the spectrogram of the target audio. The color or line type of the frame corresponding to the segmented spectrogram of different candidate segments can be different, so as to facilitate the user to distinguish the segmented spectrograms of the candidate segments. Of course, in other embodiments, in the spectrogram of the target audio, the marking corresponding to the segmented spectrogram corresponding to each candidate segment (for example, the color or line type of the frame) can be the same.
[0097] Since the segmented spectrograms corresponding to the specified content are marked in the spectrogram of the target audio, it is convenient for the user to select one or more segmented spectrograms corresponding to the specified content from the spectrogram of the target audio and add them to the same canvas for display, or add them to the same audio track for playback.
[0098] Figure 4 This is a flow chart of an embodiment of the present application showing the process of marking and displaying segmented spectrograms in the spectrogram of the target audio. Figure 4 As shown, after the step of marking and displaying the segmented spectrogram in the spectrogram of the target audio, the method further includes:
[0099] Step 410: Detect a selection operation triggered on the marked and displayed segmented spectrogram.
[0100] The selection operation may be a click operation (such as a single click, a double click, etc.), a touch operation, or an operation of dragging a mark box, etc., which is not limited here.
[0101] Step 420 : Display each selected segmented spectrogram in a display window on the canvas.
[0102] In this solution, after the segmented spectrogram corresponding to the specified content is marked and displayed in the spectrogram of the target audio, if a selection operation of the segmented spectrogram corresponding to the specified content is detected in the spectrogram of the target audio, the selected segmented spectrogram is added to the display window in the canvas for display, thereby enabling the user to select the segmented spectrogram corresponding to the specified content from the spectrogram of the target audio and add it to the canvas for display as needed.
[0103] Figure 5 is a block diagram of a spectrogram processing device according to an exemplary embodiment of the present application. Figure 5 As shown, the spectrogram processing device 500 includes: a search module 510, a determination module 520 and a display module 530.
[0104] The search module 510 is configured to search for a specified content in the speech recognition results of the target audio and determine the time position information of the sound segment corresponding to the specified content in the target audio; the specified content includes a specified phoneme or a specified text;
[0105] a determination module 520 for determining, in the spectrogram of the target audio according to the time position information, at least two segmented spectrograms corresponding to a candidate segment, where the candidate segment is a segment corresponding to a specified content in the target audio;
[0106] The display module 530 is used to display at least two segmented spectrograms in the same canvas, wherein the canvas is provided with a plurality of display windows, each display window is used to display a segmented spectrogram, and the display parameters corresponding to each of the display windows can be adjusted individually.
[0107] In some embodiments, the spectrogram processing device 500 further includes a track display unit for displaying at least two segmented spectrograms in a same track so as to continuously play the candidate sound segments corresponding to the at least two segmented spectrograms.
[0108] In some embodiments, the audio track display unit is further configured to: splice and display at least two segmented spectrograms on the same audio track in a time-ordered order based on the time position information corresponding to the at least two segmented spectrograms.
[0109] In some embodiments, the spectrogram processing apparatus 500 further includes a marking display unit for marking and displaying segmented spectrograms in the spectrogram of the target audio.
[0110] In some embodiments, the spectrogram processing apparatus 500 further includes: a detection unit and a canvas display unit. The detection unit is configured to detect a selection operation triggered on the marked segmented spectrogram for display. The canvas display unit is configured to display each selected segmented spectrogram in a display window on the canvas.
[0111] In some embodiments, the spectrogram processing device 500 further includes a speech recognition unit for performing speech recognition on the target audio to obtain a speech recognition result of the target audio.
[0112] In some embodiments, the speech recognition unit further includes an active speech detection unit for performing active speech detection on the target audio to determine the active speech in the target audio. The speech recognition unit is further configured to perform speech recognition on the active speech in the target audio to obtain a speech recognition result for the target audio.
[0113] In some embodiments, the spectrogram processing device also includes a sorting module for sorting the at least two segmented spectrograms according to preset rules to determine the sorting of the segmented spectrograms; in this embodiment, the display module is further configured to: sort the segmented spectrograms and sequentially display the at least two segmented spectrograms in multiple display windows in the same canvas.
[0114] In some embodiments, the sorting module includes a sorting unit for sorting the at least two segmented spectrograms in descending order of speech stability to determine the segmented spectrogram sorting.
[0115] According to one aspect of the present application, an electronic device is also provided, which includes: a processor; a memory, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the method in any of the above embodiments is implemented.
[0116] Figure 6 The following is a schematic diagram showing the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application. Figure 6 The computer system 600 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0117] like Figure 6 As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage part 608 into the random access memory (RAM) 603, such as executing the method in the above embodiment. Various programs and data required for system operation are also stored in the RAM 603. The CPU 601, ROM 602 and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0118] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, and the like; an output section 607 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. Removable media 611, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 610 as needed, so that computer programs read from the removable media can be installed in the storage section 608 as needed.
[0119] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from a removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the various functions defined in the system of the present application are executed.
[0120] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0122] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0123] According to one aspect of an embodiment of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of any of the above embodiments.
[0124] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0125] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0126] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed herein.
[0127] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
[0128] The above content is only a preferred exemplary embodiment of the present invention and is not intended to limit the implementation scheme of the present invention. Ordinary technicians in this field can easily make corresponding changes or modifications based on the main concept and spirit of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for processing a spectrogram, characterized in that: The method comprises: Searching for specified content in the speech recognition results of the target audio, and determining time position information of a sound segment corresponding to the specified content in the target audio; the specified content includes a specified phoneme or a specified text; Determining, in the spectrogram of the target audio according to the time position information, at least two segmented spectrograms corresponding to candidate segments, the candidate segments being segments corresponding to the specified content in the target audio; Sorting the at least two segmented spectrograms in descending order of speech stability to determine a segmented spectrogram sorting order; Sorting the segmented spectrograms, displaying the at least two segmented spectrograms in the same canvas, wherein the canvas is provided with a plurality of display windows, each display window being used to display one of the segmented spectrograms, and display parameters corresponding to each of the display windows being independently adjustable; The step of sorting the at least two segmented spectrograms in descending order of speech stability to determine the segmented spectrogram sorting comprises: For each of the segmented spectrograms, determining a reference region corresponding to each of the four formant frequencies in the segmented spectrogram based on the positions of the four formant frequencies in the segmented spectrogram; the reference region corresponding to the formant frequency is a spectrogram region in the segmented spectrogram whose distance from the position of the formant frequency is less than a first length, wherein the first length is less than a distance between two adjacent formant frequencies in the segmented spectrogram; Determining the reference area corresponding to each of the formant frequencies as the reference area corresponding to the segmented spectrogram, and calculating the average grayscale value of each reference area corresponding to the segmented spectrogram; Performing weighted processing on the average grayscale values of the four reference areas corresponding to the same segmented spectrogram to obtain a target average grayscale value corresponding to each segmented spectrogram; The at least two segmented spectrograms are sorted in order from small to large according to the target average grayscale value corresponding to each segmented spectrogram to obtain a segmented spectrogram sorting; wherein the target average grayscale value corresponding to the segmented spectrogram is negatively correlated with the speech stability of the segmented spectrogram.
2. The method according to claim 1, characterized in that After determining at least two segmented spectrograms corresponding to the candidate segments in the spectrogram of the target audio according to the time position information, the method further includes: The at least two segmented spectrograms are spliced and displayed in the same audio track, so that the candidate sound segments corresponding to the at least two segmented spectrograms are played continuously.
3. The method according to claim 2, characterized in that The step of splicing and displaying the at least two segmented spectrograms in the same audio track includes: According to the time position information corresponding to the at least two segmented spectrograms, the at least two segmented spectrograms are spliced and displayed on the same audio track in a time-ordered order.
4. The method according to claim 1, wherein After determining at least two segmented spectrograms corresponding to the candidate segments in the spectrogram of the target audio according to the time position information, the method further includes: The segmented spectrogram is marked and displayed in the spectrogram of the target audio.
5. The method according to claim 4, characterized in that After marking and displaying the segmented spectrogram in the spectrogram of the target audio, the method further includes: Detecting a selection operation triggered on the marked displayed segmented spectrogram; Each selected segmented spectrogram is displayed in a display window in the canvas.
6. The method according to claim 1, wherein Before searching for the designated content in the speech recognition results of the target audio and determining the time position information of the sound segment corresponding to the designated content in the target audio, the method further includes: Perform speech recognition on the target audio to obtain a speech recognition result of the target audio.
7. The method according to claim 6, characterized in that Before performing speech recognition on the target audio to obtain a speech recognition result of the target audio, the method further includes: Performing active voice detection on the target audio to determine the active voice in the target audio; The performing speech recognition on the target audio to obtain a speech recognition result of the target audio includes: Perform speech recognition on the active speech in the target audio to obtain a speech recognition result of the target audio.
8. A spectrogram processing device, characterized in that: include: A search module is configured to search for a specified content in the speech recognition results of the target audio and determine the time position information of the sound segment corresponding to the specified content in the target audio; the specified content includes a specified phoneme or a specified text; a determination module, configured to determine, in the spectrogram of the target audio according to the time position information, at least two segmented spectrograms corresponding to candidate segments, the candidate segments being segments corresponding to the specified content in the target audio; The sorting module is used to sort the at least two segmented spectrograms in descending order of speech stability to determine the sorting of the segmented spectrograms; specifically, for each segmented spectrogram, based on the positions of the four formant frequencies in the segmented spectrogram, determine a reference area corresponding to each formant frequency; the reference area corresponding to the formant frequency refers to a spectrogram area in the segmented spectrogram whose distance from the position of the formant frequency is less than a first length, and the first length is less than the distance between two adjacent formant frequencies in the segmented spectrogram; each of the formant frequencies is sorted into a plurality of sub-areas. The corresponding reference area is determined as the reference area corresponding to the segmented spectrogram, and the average grayscale value of each reference area corresponding to the segmented spectrogram is calculated; the average grayscale values of the four reference areas corresponding to the same segmented spectrogram are weighted to obtain the target average grayscale value corresponding to each segmented spectrogram; the at least two segmented spectrograms are sorted in order from small to large according to the target average grayscale values corresponding to each segmented spectrogram to obtain a segmented spectrogram sorting; wherein the target average grayscale value corresponding to the segmented spectrogram is negatively correlated with the speech stability of the segmented spectrogram; A display module is used to sort the segmented spectrograms and display the at least two segmented spectrograms in the same canvas, wherein a plurality of display windows are provided in the canvas, each display window is used to display a segmented spectrogram, and the display parameters corresponding to each display window can be adjusted separately.
9. An electronic device, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Audio marking method, device, equipment and readable storage medium
CN111639157A