Automatic rail alignment system and method based on AdobeAudition (AU) extension program
By implementing an automated tracking system in the Adobe Audition extension, the time-consuming and labor-intensive problem of audiobook production was solved, production efficiency and output were improved, and the needs of different production scenarios were met.
Patent Information
- Application Number
- CN202510788836.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-16
AI Technical Summary
In the prior art, in the audiobook production process, the track-matching operation based on Adobe Audition consumes a lot of manpower and time, resulting in low production output and waste of human resources.
By developing an automatic tracking system based on the Adobe Audition extension program, including a script parsing module, an audio matching loading module, a silence detection and marking module, a dialogue-audio verification module and an automatic execution module, automatic dialogue parsing, audio matching, silence detection and splicing are achieved, and fully automatic and semi-automatic modes are provided to adapt to different production needs.
It significantly shortens the time for tracking, reduces the mismatch rate, improves production efficiency and creative experience, enables production staff to focus on artistic review and emotional expression, and increases the production output of audiobooks.
Smart Images

Figure CN120656458A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to an automatic track alignment system and method based on an Adobe Audition (AU) extension program. Background Art
[0002] Tracking is a process in audiobook production, aiming to stitch together character voices recorded by multiple audio teachers at appropriate intervals to create the final product. Currently, most audiobook production processes are based on Adobe Audition. Tracking within this software is labor-intensive and time-consuming due to repetitive cutting and alignment, which reduces audiobook production output and drains audio teachers' energy on repetitive work.
[0003] Therefore, an automatic track alignment system and method based on Adobe Audition (AU) extension program is needed to solve the above problems. Summary of the Invention
[0004] In order to solve the problems of the prior art, the present invention provides an automatic track alignment system and method based on Adobe Audition (AU) extension program.
[0005] In order to solve the above technical problems, the present invention is implemented through the following technical solutions: In a first aspect, an automatic tracking system based on the Adobe Audition (AU) extension program comprises:
[0006] Script parsing module: used to receive the script data uploaded or imported by the user, and parse the content and order of each character's lines in the script according to the preset recognition rules;
[0007] Audio matching loading module: Based on the parsed role information, it automatically scans the project directory and matches the audio files of the corresponding roles, and loads the audio into the specified initial track of the Adobe Audition project;
[0008] Silence Detection and Marking Module: This module identifies silence segments in the loaded audio file and adds markers to the start of the sound segments in the audio file based on the user-configured silence detection threshold decibel value and minimum silence duration.
[0009] Lines-audio verification module: automatically compares the number of added markers with the number of lines in the script;
[0010] When the numbers are inconsistent, a visual comparison interface is triggered, allowing users to click on a line to play the corresponding audio clip;
[0011] Convert audio clips into text through speech recognition, calculate the matching degree between audio clips and dialogues by combining glyph and pinyin similarity, and provide manual calibration entry for low-matching clips;
[0012] Automatic execution module:
[0013] Based on the order of the lines and the intervals between sentences configured by the user, the audio clips at each marker are cut and then spliced together in sequence to generate the finished track (fully automatic mode);
[0014] Or retain the marked fragments after cutting, and the user manually triggers the forward splicing of fragments through the interface operation button (semi-automatic mode).
[0015] In order to completely solve the problem of tedious, time-consuming and labor-intensive repetitive operations in traditional manual tracking, this application realizes the intelligent transformation of the tracking process by deeply integrating the whole process of picture book content analysis, audio intelligent matching, silent precise positioning, multi-dimensional calibration and automatic splicing. Specifically, the character line sequence is automatically analyzed through the picture book, and the corresponding audio files are accurately matched and loaded, eliminating the redundant steps of manual search and arrangement of audio. The starting position of the human voice segment is automatically marked based on the threshold silence detection algorithm, replacing the traditional manual repeated monitoring and positioning operation, and the cutting accuracy is significantly improved. Through the automatic comparison of the number of marks and the amount of lines, and the voice-to-text + glyph pinyin two-way matching analysis, the audio is exposed in advance. It can solve the problem of deviation between audio frequency and lines, support clicking on lines to instantly play related audio, and intuitively display the voice recognition text and similarity score, greatly reducing the rework rate and troubleshooting time, and automatically generate spliced finished products with one click, or semi-actively click to advance sentence by sentence, meeting the dual requirements of detailed review and fast production, and completely liberating manpower from mechanical cutting and alignment operations, allowing production staff to focus on artistic expression and quality control, significantly improving the output and creative experience of audiobooks; this technical solution is seamlessly embedded in the industry-wide Adobe Audition platform, realizing intelligent upgrades to existing processes in the form of native extensions, and providing the audiobook industry with a standardized tracking solution that takes into account speed, accuracy and flexibility.
[0016] In a specific implementation of the first aspect, the script parsing module extracts line order, role allocation, and sentence spacing configuration parameters through regular expression matching and role tag recognition.
[0017] In a specific implementation of the first aspect, the audio matching loading module matches the audio file according to a file name rule (role name+timestamp) or a metadata tag and associates it with the dialogue role.
[0018] In a specific implementation of the first aspect, the silence detection and marking module determines, through frame energy calculation, a segment whose decibel value is continuously lower than a threshold value and lasts for more than a set duration as a silence interval, and adds a mark at the position of the first non-silent frame thereafter.
[0019] In a specific implementation of the first aspect, in the visual comparison interface of the dialogue-audio verification module, the audio waveform and the dialogue text are displayed aligned with each other in time and space, and the speech recognition conversion results and matching degree values are annotated in real time.
[0020] In a specific implementation of the first aspect, the automatic execution module provides a "single sentence splicing" control in the semi-automatic mode, which, when clicked, automatically splices the currently selected segment with the previous segment according to the interval configuration.
[0021] In a second aspect, an automatic tracking method based on an Adobe Audition (AU) extension program includes the following steps:
[0022] Step S1: importing the script through the extension interface, parsing the character's line sequence and interval parameters;
[0023] Step S2: Scan the project directory, match the character audio file and load it into the initial track;
[0024] Step S3: Detect the audio silence interval and add a time marker at the start position of the non-silence;
[0025] Step S4: Compare the number of marks and the number of lines. If they are inconsistent:
[0026] Display a visual verification interface with voice recognition transcription text;
[0027] Calculate matching degree based on glyph and phonetic similarity, and support manual calibration;
[0028] Step S5: User selects the mode to execute:
[0029] Fully automatic mode: Cut and splice the marked segments in the order of the lines and intervals to generate the finished track;
[0030] Semi-automatic mode: Only cuts the audio segments, and the user clicks to trigger the forward splicing of sentences one by one.
[0031] In a specific implementation of the second aspect, the similarity calculation in step S4 uses an edit distance algorithm to compare the original dialogue text, the pinyin sequence and the speech recognition result respectively.
[0032] In a specific implementation of the second aspect, in step S5 , in a fully automatic mode, preset audio crossfade processing parameters are automatically applied during splicing.
[0033] In a specific implementation of the second aspect, a copy of the original audio file is retained after the splicing operation is performed, and a finished file is generated on a new track for review and export.
[0034] The beneficial effects of the present invention are:
[0035] 1. This invention eliminates the repetitive labor of traditional manual cutting and splicing through automated dialogue analysis, audio matching, and silence detection, effectively shortening the time required for track alignment. The dual-mode (fully automatic / semi-automatic) design adapts to different production scenarios, meeting the dual needs of fine-tuning and batch processing.
[0036] 2. The intelligent verification mechanism significantly reduces the mismatch rate between dialogue and audio through multi-dimensional similarity calculation (voice / text / pinyin). Cross-fading processing and buffer cutting technology avoid popping or voice clipping caused by traditional manual operations.
[0037] 3. Freeing production staff from mechanical operations, allowing them to focus on artistic review and emotional expression optimization, the visual calibration interface provides intuitive creative feedback and supports rapid iterative modifications. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic flow chart of the method of the present invention. DETAILED DESCRIPTION
[0039] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0040] like Figure 1 An automatic tracking system and method based on the Adobe Audition (AU) extension program is shown.
[0041] 1. Picture book analysis module
[0042] Input: User-uploaded script text file (.docx format), or log in to the cooperation platform website through the extended interface to import data.
[0043] Parsing rules:
[0044] Use regular expressions to match character tags ([character name]) and dialogue content (character A: "dialogue content");
[0045] Extract line order, role assignment, and sentence interval configuration (default interval: 300ms±100ms, user-adjustable range 100-1000ms).
[0046] Output: structured data table, including fields: character name, line text, line number, and interval time.
[0047] 2. Audio matching loading module
[0048] Audio positioning rules:
[0049] File name matching: Scans the same directory as the Adobe Audition project and identifies files in the role name_timestamp.wav format (Zhang San_20230915.wav);
[0050] Metadata matching: Read the author metadata tag of the audio file (supports WAV / MP3 format) and associate it with the character name.
[0051] Loading logic:
[0052] Create an independent audio track for each character, and name it with the character name_track (Li Si_Track1);
[0053] Sort the matching audio files in ascending order by file name timestamps and load them to the start position of the corresponding track (0:00 on the timeline by default).
[0054] 3.Silence detection and marking module
[0055] Detection parameters:
[0056] Mute threshold: -35dB±5dB (user adjustable range -50dB to -20dB);
[0057] Minimum silence duration: 250ms±50ms (adjustable range 100-500ms).
[0058] Algorithm flow:
[0059] Frame processing: divide the audio into 10ms frames and calculate the energy value of each frame;
[0060] Silence determination: If the energy of 25 consecutive frames (i.e. 250ms) is lower than the threshold, it is determined to be a silence interval;
[0061] Marker position: Add a "Speech_Start" marker at the first frame position (±5ms accuracy) where the energy exceeds the threshold after the end of the silence segment.
[0062] 4. Lines-audio verification module
[0063] Verification process:
[0064] Quantity comparison: Count the total number of tags and the number of lines. If the difference is ≥1, a warning window will pop up.
[0065] Speech recognition: Call the localized speech recognition engine (Windows SR engine) to convert the marked interval audio into text;
[0066] Similarity calculation:
[0067] Glyph similarity: Calculates the difference between the original lines and the recognized text based on the edit distance algorithm;
[0068] Pinyin similarity: convert both sides into pinyin and calculate the edit distance (independent comparison of initials and finals);
[0069] Comprehensive matching degree = 0.6 × glyph similarity + 0.4 × pinyin similarity.
[0070] Visual calibration interface:
[0071] Display the dialogue text, recognition text, and matching degree side by side (red warning if below 80%);
[0072] Click on the lines to automatically play the audio of the corresponding marked interval, and support manual dragging to correct the marked position.
[0073] 5. Automatic execution module
[0074] Fully automatic mode process:
[0075] Traverse the markers in the order of the lines and cut the audio into independent segments (retaining a 50ms buffer before and after to avoid truncating the vocals);
[0076] Splice to a new track (named Finished Track) at the configured interval (300ms by default);
[0077] Add 20ms crossfades between clips to avoid popping sounds.
[0078] Semi-automatic mode process:
[0079] Cut audio into segments and keep them on the original track;
[0080] After the user selects a segment, click the single sentence splicing button:
[0081] Automatically align with the end of the previous segment and apply the configured value at intervals;
[0082] Automatically generate splicing guides (blue dotted lines) on the timeline to assist in positioning.
[0083] Backup mechanism:
[0084] The original audio track is set to read-only;
[0085] The finished track is saved as an independent audio file (in the same format as the source file).
[0086] 6. Key parameter configuration table
[0087] Functional modules Parameter name default value Adjustable range Silence detection Silence Threshold -35dB -50dB~-20dB Minimum silence duration 250ms 100ms~500ms Audio Cutting Cutting buffer time 50ms 30ms~100ms Fragment splicing Intersentence spacing 300ms 100ms~1000ms Crossfade duration 20ms 10ms~50ms Verification and calibration Low match threshold 80% 70%~95%
[0088] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An automatic tracking system based on Adobe Audition (AU) extension program, characterized in that: include: Script parsing module: used to receive the script data uploaded or imported by the user, and parse the content and order of each character's lines in the script according to the preset recognition rules; Audio matching loading module: Based on the parsed role information, it automatically scans the project directory and matches the audio files of the corresponding roles, and loads the audio into the specified initial track of the Adobe Audition project; Silence Detection and Marking Module: This module identifies silence segments in the loaded audio file and adds markers to the start of the sound segments in the audio file based on the user-configured silence detection threshold decibel value and minimum silence duration. Lines-audio verification module: automatically compares the number of added markers with the number of lines in the script; When the numbers are inconsistent, a visual comparison interface is triggered, allowing users to click on a line to play the corresponding audio clip; Convert audio clips into text through speech recognition, calculate the matching degree between audio clips and dialogues by combining glyph and pinyin similarity, and provide manual calibration entry for low-matching clips; Automatic execution module: Based on the order of the lines and the intervals between sentences configured by the user, the audio clips at each marker are cut and then spliced together in sequence to generate the finished track (fully automatic mode); Or retain the marked fragments after cutting, and the user manually triggers the forward splicing of fragments through the interface operation button (semi-automatic mode).
2. The automatic track alignment system based on the Adobe Audition (AU) extension program according to claim 1, characterized in that: The script parsing module extracts dialogue sequence, role assignment, and sentence spacing configuration parameters through regular expression matching and character tag recognition.
3. The automatic track alignment system based on the Adobe Audition (AU) extension program according to claim 1, characterized in that: The audio matching loading module matches audio files according to file name rules (role name + timestamp) or metadata tags and is associated with the dialogue role.
4. The automatic track alignment system based on the Adobe Audition (AU) extension program according to claim 1, characterized in that: The silence detection and marking module determines the segments with continuously lower decibel values than the threshold and lasting longer than the set duration as silence intervals through frame energy calculation, and adds a mark at the position of the first non-silent frame thereafter.
5. The automatic tracking system and method based on the Adobe Audition (AU) extension program according to claim 1, characterized in that: In the visual comparison interface of the dialogue-audio verification module, the audio waveform and the dialogue text are displayed aligned with each other in time and space, and the speech recognition conversion results and matching values are annotated in real time.
6. The automatic track alignment system based on the Adobe Audition (AU) extension program according to claim 1, characterized in that: The automatic execution module provides a "single sentence splicing" control in semi-automatic mode. When clicked, the currently selected segment and the previous segment are automatically spliced according to the interval configuration.
7. An automatic tracking method based on Adobe Audition (AU) extension program, characterized in that: The following steps are involved: Step S1: importing the script through the extension interface, parsing the character's line sequence and interval parameters; Step S2: Scan the project directory, match the character audio file and load it into the initial track; Step S3: Detect the audio silence interval and add a time marker at the start position of the non-silence; Step S4: Compare the number of marks and the number of lines. If they are inconsistent: Display a visual verification interface with voice recognition transcription text; Calculate matching degree based on glyph and phonetic similarity, and support manual calibration; Step S5: User selects the mode to execute: Fully automatic mode: Cut and splice the marked segments in the order of the lines and intervals to generate the finished track; Semi-automatic mode: Only cuts the audio segments, and the user clicks to trigger the forward splicing of sentences one by one.
8. The automatic tracking method based on the Adobe Audition (AU) extension program according to claim 1, characterized in that: The similarity calculation in step S4 uses the edit distance algorithm to compare the original dialogue text, the pinyin sequence and the speech recognition result respectively.
9. The automatic tracking method based on the Adobe Audition (AU) extension program according to claim 1, characterized in that: In the fully automatic mode of step S5, preset audio crossfade processing parameters are automatically applied during splicing.
10. The automatic tracking method based on the Adobe Audition (AU) extension program according to claim 1, characterized in that: After performing the splicing operation, a copy of the original audio file is retained and the finished file is generated on a new track for review and export.
Citation Information
Cited By
Visual configuration method and system for voice social product activity page
CN121187564A