An audio recognition result visual display method and system based on an acoustic spectrogram
By creating a dialogue region on the acoustic waveform and combining it with speech recognition results, the problem of difficulty in locating information in traditional audio recognition and display technologies has been solved, enabling rapid location and editing of audio and dialogue content and improving case-handling efficiency.
Patent Information
- Application Number
- CN202310699228.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-06-13
AI Technical Summary
Traditional audio recognition and display technologies cannot quickly and accurately locate important information in audio, nor can they intuitively display the content of dialogue between multiple characters.
By creating regions on the waveform to represent each dialogue sentence, and combining them with the speech recognition results, the system enables the visualization of audio and speech recognition results. It allows for the quick location of corresponding information through the waveform, audio playback points, or dialogue content, and supports editing and saving dialogue content.
It enables rapid location of sound waves, audio, and dialogue content, improving the work efficiency of investigators. It can quickly locate key roles and dialogues, and supports the editing and saving of dialogue content.
Smart Images

Figure CN116705050B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of police audio processing, and more particularly to an audio recognition result visualization display method and system based on an audio spectrogram. BACKGROUND
[0002] With the rapid development of audio spectrogram technology and speech recognition technology, the collected audio in the case handling process can be visualized and recognized by using computer technology. The traditional display method only stays in the association of audio and the whole speech recognition result, and it is difficult for case handling personnel to quickly locate the important information in the audio. It may need to click multiple times to accurately locate the information, which is a waste of time. In an audio, there may be multiple characters talking, and case handling personnel need to remember different words spoken by different characters, which cannot intuitively display the words spoken by each person. SUMMARY
[0003] In view of the above defects or improvement needs of the prior art, the present application provides an audio recognition result visualization display method and system based on an audio spectrogram, which aims to solve the technical problems that the existing audio recognition display technology cannot accurately and quickly locate the required information and display the recognition results by character.
[0004] To achieve the above purpose, on the one hand, the present application provides an audio recognition result visualization display method based on an audio spectrogram, which comprises:
[0005] Obtaining the audio spectrogram and speech recognition result of the audio, and obtaining the total length L and total time totalT of the audio spectrogram based on the audio spectrogram, and obtaining the start time beginT, end time endT, character and conversation content of each conversation based on the speech recognition result;
[0006] Creating different areas on the audio spectrogram corresponding to each conversation in the audio, and the start position offsetX of the area relative to the audio spectrogram is:
[0007]
[0008] The width width is:
[0009]
[0010] After the area is selected, the audio jumps to the start position of the selected area for playing;
[0011] The result of speech recognition is displayed in the form of each dialogue with the role and dialogue content in the result display area. When the audio is played, the current playing time is obtained. If the playing time is between the start time beginT and the end time endT of a dialogue, the dialogue is highlighted.
[0012] If a dialogue in the display is selected, the start time beginT of the dialogue is obtained, the audio is played from beginT, and the corresponding area of the dialogue in the sound wave graph is highlighted.
[0013] Preferably, if a role in the display is selected, the dialogue content spoken by the role in the display is highlighted, and the corresponding area of the dialogue content in the sound wave graph is highlighted.
[0014] Preferably, the area is displayed in segments in the sound wave graph. If the area is selected, the sound wave graph of the area is highlighted, and the corresponding dialogue content of the area is highlighted in the speech recognition result display area.
[0015] Preferably, if a dialogue in the display is selected, the editing function of the dialogue content is entered, and the edited dialogue content can be saved in the speech recognition result.
[0016] According to another aspect of the present application, the present application provides a sound wave graph-based audio recognition result visualization display system, which comprises:
[0017] A parameter acquisition module is configured to obtain the sound wave graph of the audio and the speech recognition result, obtain the total length L and the total time totalT of the sound wave graph based on the sound wave graph, and obtain the start time beginT, the end time endT, the role, and the dialogue content of each dialogue based on the speech recognition result.
[0018] A region selection module is configured to create different regions corresponding to each dialogue in the audio on the sound wave graph. The start position offsetX of the region relative to the sound wave graph is:
[0019]
[0020] The width width is:
[0021]
[0022] After the region is selected, the audio jumps to the start position of the selected region for playing.
[0023] The content display module is configured to display each dialogue in the voice recognition result in the form of dialogue content according to roles in a result display area, and when the audio is played, the current playing time is obtained, and if the playing time is between the start time beginT and the end time endT of a dialogue, the dialogue is highlighted;
[0024] The playing selection module is configured to determine if a dialogue in the display is selected, and if so, the start time beginT of the dialogue is obtained, the audio is played from beginT, and the corresponding area of the dialogue on the sonogram is highlighted.
[0025] Preferably, if a role in the display is selected, the dialogue content spoken by the role in the display is highlighted, and the corresponding area of the dialogue content on the sonogram is highlighted.
[0026] Preferably, the area is displayed in segments on the sonogram, and if the area is selected, the sonogram of the area is highlighted, and the corresponding dialogue content of the area in the recognition result display area is highlighted.
[0027] Preferably, if a dialogue in the display is selected, the editing function of the dialogue content is entered, and the edited dialogue content can be saved in the voice recognition result.
[0028] Overall, the above technical solutions conceived by the present application have the following beneficial effects compared with the prior art:
[0029] (1) The present application establishes the association and positioning between the sonogram, the audio and the voice recognition result, and any one of the sonogram area, the audio playing point and the dialogue content can be associated and positioned to the other two, thus the case handling personnel can quickly position the sonogram and the audio of the dialogue according to the dialogue content, or quickly position the dialogue content and the sonogram according to the audio, realizing the quick positioning of the sonogram, the audio and the dialogue content, and improving the case handling efficiency of the case handling personnel;
[0030] (2) The present application can also quickly associate all dialogue content and corresponding audio of a role according to the role information in the voice recognition result, thus the case handling personnel can quickly position the key dialogue of the key role, i.e. the voice;
[0031] (3) The present application also has a voice recognition result editing function, which can quickly find the corresponding dialogue content according to the audio editing, and edit and save the dialogue content. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 is a flowchart of creating different areas on the sonogram corresponding to each dialogue in the audio;
[0033] Figure 2 is a flow chart of synchronously displaying audio playing content in a speech recognition result;
[0034] Figure 3 is a flow chart of quickly locating each sentence in a speech recognition result in an audio wave chart and audio. DETAILED DESCRIPTION
[0035] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0036] The present application realizes an audio recognition result visualization display method based on an audio wave chart, which comprises:
[0037] The audio wave chart and the speech recognition result of the audio are obtained, and the total length L and the total time totalT of the audio wave chart are obtained based on the audio wave chart, and the start time beginT, the end time endT, the character and the dialogue content of each dialogue in the audio are obtained based on the speech recognition result.
[0038] Different areas corresponding to each dialogue in the audio are created on the audio wave chart, and the start position offsetX of the area relative to the audio wave chart is:
[0039]
[0040] The width width is:
[0041]
[0042] After the area is selected, such as after the created area click event is activated, the audio wave chart of the area is highlighted and displayed, and the audio jumps to the start position of the selected area for playing; the specific flow is as shown in Figure 1
[0043] Each dialogue in the speech recognition result is displayed in the form of character and dialogue content, and when the audio is played, the current playing time is obtained, if the playing time is located between the start time beginT and the end time endT of a dialogue, the dialogue is highlighted, such as highlighted; the specific flow is as shown in Figure 2
[0044] If a dialogue in the display is selected, such as a created dialogue click event is activated, the start time beginT of the dialogue is obtained, the audio is played from the beginT, and the corresponding area of the dialogue on the sound wave chart is highlighted, such as highlighted. The specific process is as shown in Figure 3 If a character in the display is selected, such as a created character click event is activated, the dialogue content spoken by the character in the display is highlighted, and the corresponding area of the dialogue on the sound wave chart is highlighted.
[0045] If a dialogue in the display is selected, such as a created dialogue double-click event is activated, the editing function of the dialogue content is entered, and the edited dialogue content can be saved in the speech recognition result.
[0046] The above is easily understood by those skilled in the art, and the above is only a preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A method for visualizing audio recognition results based on sonograms, characterized in that, The method comprises: obtaining an audio wave chart and a speech recognition result of audio, and obtaining a total length L and a total time totalT of the audio wave chart based on the audio wave chart, and obtaining a start time beginT, an end time endT, a character and a dialogue content of each dialogue in the audio based on the speech recognition result; creating different regions on the audio wave chart corresponding to each dialogue in the audio, the regions having a start position offsetX relative to the audio wave chart as: a width width as: after the regions are selected, the audio jumps to the start position of the selected region for playing; displaying each dialogue in the speech recognition result in the form of character plus dialogue content in a result display area, and when the audio is played, obtaining a current playing time, and if the playing time is between the start time beginT and the end time endT of a dialogue, highlighting the dialogue; if a dialogue in the display is selected, obtaining the start time beginT of the dialogue, jumping to the beginT of the audio for playing, and highlighting the region on the audio wave chart corresponding to the dialogue.
2. The method of claim 1, wherein, if a character in the display is selected, the dialogue content spoken by the character in the display is highlighted, and the region on the audio wave chart corresponding to the dialogue content is highlighted.
3. The method of claim 1, wherein, the regions are displayed in segments on the audio wave chart, if the regions are selected, the audio wave chart of the regions is highlighted, and the dialogue content corresponding to the regions in the recognition result display area is highlighted.
4. The method of claim 1, wherein, if a dialogue in the display is selected, entering the editing function of the dialogue content, and being able to save the edited dialogue content in the speech recognition result.
5. An audio recognition result visualization display system based on an audio wave chart, the system comprising: a parameter obtaining module for obtaining an audio wave chart and a speech recognition result of audio, and obtaining a total length L and a total time totalT of the audio wave chart based on the audio wave chart, and obtaining a start time beginT, an end time endT, a character and a dialogue content of each dialogue in the audio based on the speech recognition result; a region selecting module for creating different regions on the audio wave chart corresponding to each dialogue in the audio, the regions having a start position offsetX relative to the audio wave chart as: a width width as: after the regions are selected, the audio jumps to the start position of the selected region for playing; a content displaying module for displaying each dialogue in the speech recognition result in the form of character plus dialogue content in a result display area, and when the audio is played, obtaining a current playing time, and if the playing time is between the start time beginT and the end time endT of a dialogue, highlighting the dialogue; a playing selecting module for judging, if a dialogue in the display is selected, obtaining the start time beginT of the dialogue, jumping to the beginT of the audio for playing, and highlighting the region on the audio wave chart corresponding to the dialogue.
6. The system of claim 5, wherein, if a character in the display is selected, the dialogue content spoken by the character in the display is highlighted, and the region on the audio wave chart corresponding to the dialogue content is highlighted.
7. The system of claim 5, wherein, The region is displayed in segments on the sonogram, and if the region is selected, the sonogram of the region is highlighted, and the corresponding dialogue content of the region is highlighted in the identification result display area.
8. The system of claim 5, wherein, If a displayed dialogue is selected, the editing function of the dialogue content is entered, and the edited dialogue content can be saved in the speech recognition result.
Citation Information
Patent Citations
Wetland ecological monitoring system with audio separation voiceprint recognition function and audio separation method thereof
CN112735442A
Audio display method and device
CN114579017A