An audio recognition result visual display method and system based on an acoustic spectrogram

By creating a dialogue region on the acoustic waveform and combining it with speech recognition results, the problem of difficulty in locating information in traditional audio recognition and display technologies has been solved, enabling rapid location and editing of audio and dialogue content and improving case-handling efficiency.

CN116705050BActive Publication Date: 2025-12-12709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310699228.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2025-12-12
Estimated Expiration
2043-06-13

AI Technical Summary

Technical Problem

Traditional audio recognition and display technologies cannot quickly and accurately locate important information in audio, nor can they intuitively display the content of dialogue between multiple characters.

Method used

By creating regions on the waveform to represent each dialogue sentence, and combining them with the speech recognition results, the system enables the visualization of audio and speech recognition results. It allows for the quick location of corresponding information through the waveform, audio playback points, or dialogue content, and supports editing and saving dialogue content.

Benefits of technology

It enables rapid location of sound waves, audio, and dialogue content, improving the work efficiency of investigators. It can quickly locate key roles and dialogues, and supports the editing and saving of dialogue content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116705050B_ABST
    Figure CN116705050B_ABST
Patent Text Reader

Abstract

The application discloses an audio recognition result visual display method based on an audio graph, and belongs to the technical field of police audio processing. The application provides a perfect audio recognition result display idea and an interaction mode with an audio graph, realizes the mutual correspondence and dynamic interaction between each sentence in the audio recognition result and a segment of the audio graph, clearly displays the position of each sentence in the audio recognition result on the audio graph and the corresponding audio picture segment, forms the visual display capability of the recognition result, and highlights the dialogue content corresponding to the current playing position in the voice recognition result display area during the audio graph playing, so as to realize the synchronous display of the recognition result and the audio graph. Clicking each sentence in the voice recognition result area controls the audio graph to jump to the corresponding position, so as to realize the quick positioning of the recognition result. The application realizes the quick positioning of the audio graph, the audio and the dialogue content, and improves the case handling efficiency of case handling personnel.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of police audio processing, and more particularly to an audio recognition result visualization display method and system based on an audio spectrogram. BACKGROUND

[0002] With the rapid development of audio spectrogram technology and speech recognition technology, the collected audio in the case handling process can be visualized and recognized by using computer technology. The traditional display method only stays in the association of audio and the whole speech recognition result, and it is difficult for case handling personnel to quickly locate the important information in the audio. It may need to click multiple times to accurately locate the information, which is a waste of time. In an audio, there may be multiple characters talking, and case handling personnel need to remember different words spoken by different characters, which cannot intuitively display the words spoken by each person. SUMMARY

[0003] In view of the above defects or improvement needs of the prior art, the present application provides an audio recognition result visualization display method and system based on an audio spectrogram, which aims to solve the technical problems that the existing audio recognition display technology cannot accurately and quickly locate the required information and display the recognition results by character.

[0004] To achieve the above purpose, on the one hand, the present application provides an audio recognition result visualization display method based on an audio spectrogram, which comprises:

[0005] Obtaining the audio spectrogram and speech recognition result of the audio, and obtaining the total length L and total time totalT of the audio spectrogram based on the audio spectrogram, and obtaining the start time beginT, end time endT, character and conversation content of each conversation based on the speech recognition result;

[0006] Creating different areas on the audio spectrogram corresponding to each conversation in the audio, and the start position offsetX of the area relative to the audio spectrogram is:

[0007]

[0008] The width width is:

[0009]

[0010] After the area is selected, the audio jumps to the start position of the selected area for playing;

[0011] The result of speech recognition is displayed in the form of each dialogue with the role and dialogue content in the result display area. When the audio is played, the current playing time is obtained. If the playing time is between the start time beginT and the end time endT of a dialogue, the dialogue is highlighted.

[0012] If a dialogue in the display is selected, the start time beginT of the dialogue is obtained, the audio is played from beginT, and the corresponding area of the dialogue in the sound wave graph is highlighted.

[0013] Preferably, if a role in the display is selected, the dialogue content spoken by the role in the display is highlighted, and the corresponding area of the dialogue content in the sound wave graph is highlighted.

[0014] Preferably, the area is displayed in segments in the sound wave graph. If the area is selected, the sound wave graph of the area is highlighted, and the corresponding dialogue content of the area is highlighted in the speech recognition result display area.

[0015] Preferably, if a dialogue in the display is selected, the editing function of the dialogue content is entered, and the edited dialogue content can be saved in the speech recognition result.

[0016] According to another aspect of the present application, the present application provides a sound wave graph-based audio recognition result visualization display system, which comprises:

[0017] A parameter acquisition module is configured to obtain the sound wave graph of the audio and the speech recognition result, obtain the total length L and the total time totalT of the sound wave graph based on the sound wave graph, and obtain the start time beginT, the end time endT, the role, and the dialogue content of each dialogue based on the speech recognition result.

[0018] A region selection module is configured to create different regions corresponding to each dialogue in the audio on the sound wave graph. The start position offsetX of the region relative to the sound wave graph is:

[0019]

[0020] The width width is:

[0021]

[0022] After the region is selected, the audio jumps to the start position of the selected region for playing.

[0023] The content display module is configured to display each dialogue in the voice recognition result in the form of dialogue content according to roles in a result display area, and when the audio is played, the current playing time is obtained, and if the playing time is between the start time beginT and the end time endT of a dialogue, the dialogue is highlighted;

[0024] The playing selection module is configured to determine if a dialogue in the display is selected, and if so, the start time beginT of the dialogue is obtained, the audio is played from beginT, and the corresponding area of the dialogue on the sonogram is highlighted.

[0025] Preferably, if a role in the display is selected, the dialogue content spoken by the role in the display is highlighted, and the corresponding area of the dialogue content on the sonogram is highlighted.

[0026] Preferably, the area is displayed in segments on the sonogram, and if the area is selected, the sonogram of the area is highlighted, and the corresponding dialogue content of the area in the recognition result display area is highlighted.

[0027] Preferably, if a dialogue in the display is selected, the editing function of the dialogue content is entered, and the edited dialogue content can be saved in the voice recognition result.

[0028] Overall, the above technical solutions conceived by the present application have the following beneficial effects compared with the prior art:

[0029] (1) The present application establishes the association and positioning between the sonogram, the audio and the voice recognition result, and any one of the sonogram area, the audio playing point and the dialogue content can be associated and positioned to the other two, thus the case handling personnel can quickly position the sonogram and the audio of the dialogue according to the dialogue content, or quickly position the dialogue content and the sonogram according to the audio, realizing the quick positioning of the sonogram, the audio and the dialogue content, and improving the case handling efficiency of the case handling personnel;

[0030] (2) The present application can also quickly associate all dialogue content and corresponding audio of a role according to the role information in the voice recognition result, thus the case handling personnel can quickly position the key dialogue of the key role, i.e. the voice;

[0031] (3) The present application also has a voice recognition result editing function, which can quickly find the corresponding dialogue content according to the audio editing, and edit and save the dialogue content. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 is a flowchart of creating different areas on the sonogram corresponding to each dialogue in the audio;

[0033] Figure 2 is a flow chart of synchronously displaying audio playing content in a speech recognition result;

[0034] Figure 3 is a flow chart of quickly locating each sentence in a speech recognition result in an audio wave chart and audio. DETAILED DESCRIPTION

[0035] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0036] The present application realizes an audio recognition result visualization display method based on an audio wave chart, which comprises:

[0037] The audio wave chart and the speech recognition result of the audio are obtained, and the total length L and the total time totalT of the audio wave chart are obtained based on the audio wave chart, and the start time beginT, the end time endT, the character and the dialogue content of each dialogue in the audio are obtained based on the speech recognition result.

[0038] Different areas corresponding to each dialogue in the audio are created on the audio wave chart, and the start position offsetX of the area relative to the audio wave chart is:

[0039]

[0040] The width width is:

[0041]

[0042] After the area is selected, such as after the created area click event is activated, the audio wave chart of the area is highlighted and displayed, and the audio jumps to the start position of the selected area for playing; the specific flow is as shown in Figure 1

[0043] Each dialogue in the speech recognition result is displayed in the form of character and dialogue content, and when the audio is played, the current playing time is obtained, if the playing time is located between the start time beginT and the end time endT of a dialogue, the dialogue is highlighted, such as highlighted; the specific flow is as shown in Figure 2

[0044] ​​If a dialogue in the display is selected, such as a created dialogue click event is activated, the start time beginT of the dialogue is obtained, the audio is played from the beginT, and the corresponding area of the dialogue on the sound wave chart is highlighted, such as highlighted. The specific process is as shown in Figure 3 If a character in the display is selected, such as a created character click event is activated, the dialogue content spoken by the character in the display is highlighted, and the corresponding area of the dialogue on the sound wave chart is highlighted.

[0045] If a dialogue in the display is selected, such as a created dialogue double-click event is activated, the editing function of the dialogue content is entered, and the edited dialogue content can be saved in the speech recognition result.

[0046] The above is easily understood by those skilled in the art, and the above is only a preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A method for visualizing audio recognition results based on sonograms, characterized in that, The method comprises: obtaining an audio wave chart and a speech recognition result of audio, and obtaining a total length L and a total time totalT of the audio wave chart based on the audio wave chart, and obtaining a start time beginT, an end time endT, a character and a dialogue content of each dialogue in the audio based on the speech recognition result; creating different regions on the audio wave chart corresponding to each dialogue in the audio, the regions having a start position offsetX relative to the audio wave chart as: a width width as: after the regions are selected, the audio jumps to the start position of the selected region for playing; displaying each dialogue in the speech recognition result in the form of character plus dialogue content in a result display area, and when the audio is played, obtaining a current playing time, and if the playing time is between the start time beginT and the end time endT of a dialogue, highlighting the dialogue; if a dialogue in the display is selected, obtaining the start time beginT of the dialogue, jumping to the beginT of the audio for playing, and highlighting the region on the audio wave chart corresponding to the dialogue.

2. The method of claim 1, wherein, if a character in the display is selected, the dialogue content spoken by the character in the display is highlighted, and the region on the audio wave chart corresponding to the dialogue content is highlighted.

3. The method of claim 1, wherein, the regions are displayed in segments on the audio wave chart, if the regions are selected, the audio wave chart of the regions is highlighted, and the dialogue content corresponding to the regions in the recognition result display area is highlighted.

4. The method of claim 1, wherein, if a dialogue in the display is selected, entering the editing function of the dialogue content, and being able to save the edited dialogue content in the speech recognition result.

5. An audio recognition result visualization display system based on an audio wave chart, the system comprising: a parameter obtaining module for obtaining an audio wave chart and a speech recognition result of audio, and obtaining a total length L and a total time totalT of the audio wave chart based on the audio wave chart, and obtaining a start time beginT, an end time endT, a character and a dialogue content of each dialogue in the audio based on the speech recognition result; a region selecting module for creating different regions on the audio wave chart corresponding to each dialogue in the audio, the regions having a start position offsetX relative to the audio wave chart as: a width width as: after the regions are selected, the audio jumps to the start position of the selected region for playing; a content displaying module for displaying each dialogue in the speech recognition result in the form of character plus dialogue content in a result display area, and when the audio is played, obtaining a current playing time, and if the playing time is between the start time beginT and the end time endT of a dialogue, highlighting the dialogue; a playing selecting module for judging, if a dialogue in the display is selected, obtaining the start time beginT of the dialogue, jumping to the beginT of the audio for playing, and highlighting the region on the audio wave chart corresponding to the dialogue.

6. The system of claim 5, wherein, if a character in the display is selected, the dialogue content spoken by the character in the display is highlighted, and the region on the audio wave chart corresponding to the dialogue content is highlighted.

7. The system of claim 5, wherein, The region is displayed in segments on the sonogram, and if the region is selected, the sonogram of the region is highlighted, and the corresponding dialogue content of the region is highlighted in the identification result display area.

8. The system of claim 5, wherein, If a displayed dialogue is selected, the editing function of the dialogue content is entered, and the edited dialogue content can be saved in the speech recognition result.

Citation Information

Patent Citations

  • Wetland ecological monitoring system with audio separation voiceprint recognition function and audio separation method thereof

    CN112735442A

  • Audio display method and device

    CN114579017A