Text and audio presentation processing method and system

Through the text and audio presentation processing method and system, the synchronous presentation of book listening and reading is achieved, which solves the problem of seamless switching between text version and audio version, provides an efficient post-editing solution, and supports the updating of recording materials and sound effect processing.

CN114595356BActive Publication Date: 2025-09-30DALIAN INSTANT INTELLIGENCE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210089590.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-26
Publication Date
2025-09-30
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

In the existing technology, the text version and audio version of the book cannot be switched seamlessly, resulting in an inability to smoothly proceed during reading and listening. When the host's recording cannot be completed, the entire production process is hindered, and post-editing cannot intuitively use the script text for editing.

Method used

Scripts are generated through script editing equipment, and sound effects and mixing are performed by post-production equipment. Combined with audio presentation equipment, synchronous presentation of text and audio is achieved, providing a visual multi-track audio non-linear editing interface, and efficient production is achieved using recorded materials and sound effects materials.

Benefits of technology

It realizes the integrated text and audio presentation of books for listening and reading. Editing users can conveniently carry out post-production of audio books and complete editing quickly and accurately. It can even be completed efficiently when there is a lack of recording materials, reducing the burden of manual adjustment of materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114595356B_ABST
    Figure CN114595356B_ABST
Patent Text Reader

Abstract

The present invention discloses a text and audio presentation processing method, comprising: a script editing device generates a script; a post-production device obtains the script from the script editing device and creates a post-production project; the post-production device performs post-production processing on the sound clip through the post-production project, generates a post-production result and outputs it; and an audio presentation device presents the post-production result. In addition, the present invention also discloses a text and audio presentation processing system. The present invention can realize the text and audio presentation of books in an integrated listening and reading mode, structure the audio data through the script and establish a connection between the audio and the text; enable editing users to conveniently perform post-production of audio books and efficiently complete production using materials; associate the materials with the script, provide a visual multi-track non-linear editing interface, and complete post-production quickly and accurately even if there is a lack of materials; and eliminate the burden of manually adjusting the material time when the lines in the script are updated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of recording and mixing, and in particular to a text and audio presentation processing method and system. Background Art

[0002] Currently, due to the popularization and promotion of computer and Internet technologies, the dissemination form of books is no longer limited to traditional paper reading materials. A large number of books, especially novels, have electronic text and audio versions, among which the audio version is an audio book. However, in the existing technology, reading software can usually only present the text version, and audio book software can usually only present the audio version.

[0003] With the increasing popularity of audiobooks, users are demanding a new mode that allows them to seamlessly switch between reading and listening. For example, in a daily scenario, a user might read the text version of a book before bed at night, continue listening to the audio version of the book from where they left off the previous night while commuting to work the next morning, and then resume reading the text version of the book in the afternoon from where they had listened in the morning.

[0004] However, the inventors discovered through research that the problem with the existing technology is that in traditional audiobook production, the production process is a serial process, starting with script editing, continuing with host recording, and finally completing post-editing. The host recording process is a key and core link in the production process. If the host recording cannot be completed, it will hinder the entire production process. In post-editing, the editing object only includes audio, and it is impossible to gain a better understanding of the editing object intuitively from the text of the script. To solve this problem, optimize the production process and improve production efficiency, a production solution is needed that can fully utilize the relationship between audio and text, thereby achieving a text and audio presentation that is integrated with listening and reading. Summary of the Invention

[0005] Based on this, in order to solve the technical problems in the existing technology, a text and audio presentation processing method is proposed, including:

[0006] The script editing device generates a script; the script includes one or more paragraphs;

[0007] The post-production device obtains the script from the script editing device connected thereto and creates a post-production project based on the script; the post-production project includes sound clips corresponding to paragraphs in the script;

[0008] The post-production device performs post-production processing on the sound clip through a post-production project, generates a post-production result and outputs it to the audio presentation device connected thereto;

[0009] The audio presentation device presents the post-production result.

[0010] In one embodiment, the script includes recording materials, audio materials, sound effect processing methods, and paragraph presentation order corresponding to the paragraphs;

[0011] The post-production equipment includes a sound effect processor, a mixing processor, a material manager, and an editing display;

[0012] The material manager includes a material library, which stores materials locally; the material manager organizes and manages the materials in the material library in a hierarchical structure, and the material types in the material library include audio materials and effect materials;

[0013] The sound effect processor obtains audio material through the material manager connected thereto, applies sound effect processing to the audio material according to the sound effect processing method set by the script, and outputs the audio material after sound effect processing to the mixing processor connected thereto;

[0014] The mixing processor includes a main track and an auxiliary track, wherein the main track and the auxiliary track are used to carry sound clips; the mixing processor performs a mixing processing operation on the sound clips in the main track and the auxiliary track according to the script to obtain a mixing processing result;

[0015] The script paragraphs include text paragraphs and audio paragraphs; the sound clips include sound clips in the main track corresponding to the text paragraphs and sound clips in the auxiliary track corresponding to the audio paragraphs;

[0016] The position of the sound clips in the main track is determined by the paragraph presentation order set by the script, and the position of the sound clips in the auxiliary track is set by the editing user;

[0017] The editing display displays an editing view during the mixing process of the mixing processor, wherein the editing view includes a text editing view and a multi-track editing view; the text editing view is bound to the cursor position in the multi-track editing view;

[0018] In the editing display, the text editing view and the multi-track editing view are presented simultaneously, or the editing user selects one of the views to be presented and can switch between them.

[0019] In one embodiment, the clip content of the sound clip in the main track includes an association relationship between the sound clip in the main track and a text paragraph in the script, text content of the associated text paragraph, an association relationship between the sound clip in the main track and the recording material, the associated recording material, and editing information of the sound clip in the main track;

[0020] When a sound clip in the main track is associated with an audio recording, the clip content of the sound clip in the main track also includes text alignment information between the text content and the audio recording. For a script that has completed the recording of the audio recording, the sound clips associated with all text paragraphs are arranged on the main track in the paragraph presentation order set by the script, and are separated by leading silence.

[0021] wherein the multi-track editing view presents a main track and a secondary track, and the timeline of the multi-track editing view is associated with the text editing view via text alignment information of the sound clip in the main track;

[0022] The editing information of the sound clip in the main track includes the leading silence duration, the starting position of the material playback, the duration of the sound clip in the main track, and the effect setting information;

[0023] The segment content of the sound clip in the main track includes one or more anchor points; wherein the anchor points are set at positions in the text alignment information; or, the anchor points are set at specific positions of the sound clip in the main track, and the specific positions include the start and end of the sound clip in the main track; or, the anchor points are set at relative positions based on the recording material set by the editing user; wherein the anchor points include semantic information, and the semantic information includes words or phrases in the text, audio event descriptions, start times, and end times.

[0024] In one embodiment, the segment content of the sound segment in the auxiliary track includes an association relationship between the sound segment and the audio material in the auxiliary track, the associated audio material, and editing information of the sound segment in the auxiliary track;

[0025] The editing information of the sound clip in the auxiliary track includes the material playback start position, the duration of the sound clip in the auxiliary track, the sound clip start position in the auxiliary track, the loop playback setting information, and the track description information;

[0026] The positioning methods for the start position of the sound clip in the auxiliary track include absolute time positioning and relative binding positioning. In the absolute time positioning method, the start position of the sound clip in the auxiliary track is located at a specific time defined by the global time axis; in the relative binding positioning method, the start position of the sound clip in the auxiliary track is located at an offset time relative to the anchor point, and the anchor point is provided by the sound clip in the main track.

[0027] Wherein, when loop play is effective, the loop play setting information includes the loop end position; the loop end position is set by the absolute duration after the loop, or by the anchor point provided by the sound clip in the main track;

[0028] During the mixing process, the mixing processor obtains the audio material processed by the sound effect processor and adds the obtained audio material to the editing view of the editing display. The added audio material is placed at the position set by the editing user.

[0029] In one embodiment, a post-production project includes project default configuration information; the project default configuration information includes a default interval value between sound clips in a master track, a target gain setting value for a mixing process, and an estimated speech rate of a sound recordist; wherein the default interval value between sound clips in a master track includes a default leading silence duration, and the target gain setting value for a mixing process includes target volume values ​​for vocals, music, and sound effects;

[0030] During the mixing process, if a text paragraph has not yet been recorded and the main audio clip lacks associated recording material, the duration of the main audio clip on the main track is determined by the recording engineer's estimated speaking speed and the number of words in the paragraph in the project's default configuration information, or the editor selects the speech synthesis result as the recording material for temporary use.

[0031] In addition, in order to solve the technical problems in the prior art, a text and audio presentation processing system is proposed, comprising a script editing device, a post-production device, and an audio presentation device which are sequentially connected to each other;

[0032] Wherein, the script editing device generates a script; the script includes one or more paragraphs;

[0033] The post-production device obtains the script and creates a post-production project according to the script; the post-production project includes sound clips, and the sound clips correspond to paragraphs in the script;

[0034] The post-production device performs post-production processing on the sound clip through a post-production project, generates a post-production result and outputs it to the audio presentation device;

[0035] The audio presentation device presents the post-production result.

[0036] In one embodiment, the script includes recording materials, audio materials, sound effect processing methods, and paragraph presentation order corresponding to the paragraphs;

[0037] The post-production equipment includes a sound effect processor, a sound mixing processor, a material manager, and an editing display; the sound effect processor is connected to the material manager; the sound effect processor is connected to the sound mixing processor; the sound mixing processor is connected to the editing display;

[0038] The material manager includes a material library, which stores materials locally; the material manager organizes and manages the materials in the material library in a hierarchical structure, and the material types in the material library include audio materials and effect materials;

[0039] The sound effect processor obtains the audio material through the material manager, applies a sound effect processing operation to the audio material according to the sound effect processing method set by the script, and outputs the audio material after the sound effect processing to the mixing processor;

[0040] The mixing processor includes a main track and an auxiliary track, wherein the main track and the auxiliary track are used to carry sound clips; the mixing processor performs a mixing processing operation on the sound clips in the main track and the auxiliary track according to the script to obtain a mixing processing result;

[0041] The script paragraphs include text paragraphs and audio paragraphs; the sound clips include sound clips in the main track corresponding to the text paragraphs and sound clips in the auxiliary track corresponding to the audio paragraphs;

[0042] The position of the sound clips in the main track is determined by the paragraph presentation order set by the script, and the position of the sound clips in the auxiliary track is set by the editing user;

[0043] The editing display displays an editing view during the mixing process of the mixing processor, wherein the editing view includes a text editing view and a multi-track editing view; the text editing view is bound to the cursor position in the multi-track editing view;

[0044] In the editing display, the text editing view and the multi-track editing view are presented simultaneously, or the editing user selects one of the views to be presented and can switch between them.

[0045] In one embodiment, the clip content of the sound clip in the main track includes an association relationship between the sound clip in the main track and a text paragraph in the script, text content of the associated text paragraph, an association relationship between the sound clip in the main track and the recording material, the associated recording material, and editing information of the sound clip in the main track;

[0046] When a sound clip in the main track is associated with an audio recording, the clip content of the sound clip in the main track also includes text alignment information between the text content and the audio recording. For a script that has completed the recording of the audio recording, the sound clips associated with all text paragraphs are arranged on the main track in the paragraph presentation order set by the script, and are separated by leading silence.

[0047] wherein the multi-track editing view presents a main track and a secondary track, and the timeline of the multi-track editing view is associated with the text editing view via text alignment information of the sound clip in the main track;

[0048] The editing information of the sound clip in the main track includes the leading silence duration, the starting position of the material playback, the duration of the sound clip in the main track, and the effect setting information;

[0049] The segment content of the sound clip in the main track includes one or more anchor points; wherein the anchor points are set at positions in the text alignment information; or, the anchor points are set at specific positions of the sound clip in the main track, and the specific positions include the start and end of the sound clip in the main track; or, the anchor points are set at relative positions based on the recording material set by the editing user; wherein the anchor points include semantic information, and the semantic information includes words or phrases in the text, audio event descriptions, start time, and end time.

[0050] In one embodiment, the segment content of the sound segment in the auxiliary track includes an association relationship between the sound segment and the audio material in the auxiliary track, the associated audio material, and editing information of the sound segment in the auxiliary track;

[0051] The editing information of the sound clip in the auxiliary track includes the material playback start position, the duration of the sound clip in the auxiliary track, the sound clip start position in the auxiliary track, the loop playback setting information, and the track description information;

[0052] The positioning methods for the start position of the sound clip in the auxiliary track include absolute time positioning and relative binding positioning. In the absolute time positioning method, the start position of the sound clip in the auxiliary track is located at a specific time defined by the global time axis; in the relative binding positioning method, the start position of the sound clip in the auxiliary track is located at an offset time relative to the anchor point, and the anchor point is provided by the sound clip in the main track.

[0053] Wherein, when loop playback is effective, the loop playback setting information includes the loop end position; the loop end position is set by the absolute duration after the loop, or by the anchor point provided by the sound clip in the main track;

[0054] During the mixing process, the mixing processor obtains the audio material processed by the sound effect processor and adds the obtained audio material to the editing view of the editing display. The added audio material is placed at the position set by the editing user.

[0055] In one embodiment, a post-production project includes project default configuration information; the project default configuration information includes a default interval value between sound clips in a master track, a target gain setting value for a mixing process, and an estimated speech rate of a sound recordist; wherein the default interval value between sound clips in a master track includes a default leading silence duration, and the target gain setting value for a mixing process includes target volume values ​​for vocals, music, and sound effects;

[0056] During the mixing process, if a text paragraph has not yet been recorded and the main audio clip lacks associated recording material, the duration of the main audio clip on the main track is determined by the recording engineer's estimated speaking speed and the number of words in the paragraph in the project's default configuration information, or the editor selects the speech synthesis result as the recording material for temporary use.

[0057] The implementation of the present invention will have the following beneficial effects:

[0058] The present invention can realize the integrated text and audio presentation of books for listening and reading, structure the audio data through the mixing script, and establish the connection between audio and text; the present invention enables editing users to easily carry out post-production of audio books according to the script, and efficiently complete the production and output by using the recording materials recorded by the sound recordist and audio materials such as music and sound effects; the present invention associates the recording materials, audio materials and scripts, and provides a visual multi-track audio non-linear editing interface, which can complete fast and accurate post-editing even in the absence of recording materials; when the recording materials of the lines in the script are updated, the user can be relieved of the burden of manually adjusting the material time. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0060] in:

[0061] Figure 1 is a schematic diagram of a text and audio presentation processing system according to the present invention;

[0062] Figure 2 Schematic diagram of the flow of the text and audio presentation processing method of the present invention. DETAILED DESCRIPTION

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0064] The present invention discloses a text and audio presentation processing system, comprising a script editing device, a post-production device, and an audio presentation device which are sequentially connected to each other;

[0065] Wherein, the script editing device generates a script; the script includes one or more paragraphs;

[0066] The script includes the recording materials, audio materials, sound effect processing methods, and the order of paragraph presentation corresponding to the paragraphs;

[0067] The post-production device obtains the script and creates a post-production project according to the script; the post-production project includes sound clips, and the sound clips correspond to paragraphs in the script;

[0068] In particular, the post-production project includes default project configuration information; the default project configuration information includes a default interval value between sound clips in the main track, a target gain setting value for the mixing process, and an estimated speech rate of the sound recordist; wherein the default interval value between sound clips in the main track is a default leading silence duration, and the target gain setting value for the mixing process includes target volume values ​​for vocals, music, and sound effects;

[0069] In particular, the post-production equipment includes a sound effect processor, a sound mixing processor, a material manager, and an editing display; the sound effect processor is connected to the material manager; the sound effect processor is connected to the sound mixing processor; the sound mixing processor is connected to the editing display;

[0070] The material manager includes a material library, which stores materials locally; the material manager organizes and manages the materials in the material library in a hierarchical structure; the material types in the material library include audio materials and effect materials;

[0071] Among them, audio materials include but are not limited to sound effect materials and music materials; effect materials include effector setting parameters, such as cave environment effects, broadcast effects, etc.

[0072] The sound effect processor obtains the audio material through the material manager, applies a sound effect processing operation to the audio material according to the sound effect processing method set by the script, and outputs the audio material after the sound effect processing to the mixing processor;

[0073] Among them, sound effects include but are not limited to gain, channel mixing, fade in and fade out, equalization, environment, noise reduction, compression, time stretching, and three-dimensional effects;

[0074] In the sound effect processing method, multiple sound effects are connected in series and the order of the sound effects can be adjusted. The sound effects include one or more adjustable operating parameters, and the adjustable operating parameters of the sound effects are set to change according to a specific curve over the playing time.

[0075] The mixing processor includes a main track and an auxiliary track, wherein the main track and the auxiliary track are used to carry sound clips; the mixing processor performs a mixing processing operation on the sound clips in the main track and the auxiliary track according to the script to obtain a mixing processing result;

[0076] The script paragraphs include text paragraphs and audio paragraphs; the sound clips include sound clips in the main track corresponding to the text paragraphs and sound clips in the auxiliary track corresponding to the audio paragraphs;

[0077] The position of the sound clips in the main track is determined by the paragraph presentation order set by the script, and the position of the sound clips in the auxiliary track is set by the editing user;

[0078] The main track is positioned according to the order of presentation of the sound clips it carries, the duration of the paragraphs, and the leading silence;

[0079] In particular, the clip content of the sound clip in the main track includes the association relationship between the sound clip in the main track and the text paragraph in the script, the text content of the associated text paragraph, the association relationship between the sound clip in the main track and the recording material, the associated recording material, and the editing information of the sound clip in the main track;

[0080] In particular, for text paragraphs that have not yet been recorded, since there is no corresponding recording material, the corresponding sound clip content in the main track does not include the association relationship between the sound clip in the main track and the recording material;

[0081] When the sound clip in the main track is associated with the recording material, that is, when the clip content of the sound clip in the main track includes the association relationship between the sound clip in the main track and the recording material, the clip content of the sound clip in the main track also includes text alignment information between the text content and the recording material;

[0082] For a completed script, all the sound clips associated with the text paragraphs are arranged on the main track in the order of paragraph presentation set in the script, and separated from each other by leading silence.

[0083] The editing information of the sound clip in the main track includes the leading silence duration, the starting position of the material playback, the duration of the sound clip in the main track, and the effect setting information;

[0084] In particular, the segment content of the sound segment in the main track includes one or more anchor points;

[0085] The anchor point is set at a position in the text alignment information; or, the anchor point is set at a specific position of the sound clip in the main track, the specific position including but not limited to the beginning and end of the sound clip in the main track; or, the anchor point is set at a relative position based on the recording material set by the editing user;

[0086] The anchor includes semantic information, including a word or phrase in the text, an audio event description, a start time, and an end time;

[0087] The anchor point is used to locate the text content based on semantics; the corresponding text paragraph is located from the audio moment according to the anchor point, or the corresponding audio moment is located from the text paragraph;

[0088] The sound clips in the main track are the main presentation objects of the script and the recording materials corresponding to the text paragraphs in the script during the post-production of the audiobook. That is, the main track usually only contains the recording materials corresponding to the text paragraphs in the script, and the recording materials are recorded by the sound recordist or obtained through speech synthesis.

[0089] The sound clips in the master track can also be used to present audio materials, that is, to place audio materials in the master track;

[0090] Inserting an audio material into the middle of the main audio clip of the master track will convert the audio material into a main audio clip and position it according to the positioning method of the master track.

[0091] In particular, the clip content of the sound clip in the auxiliary track includes the association relationship between the sound clip and the audio material in the auxiliary track, the associated audio material, and the editing information of the sound clip in the auxiliary track;

[0092] In particular, the editing information of the sound clip in the auxiliary track includes the material playback start position, the duration of the sound clip in the auxiliary track, the start position of the sound clip in the auxiliary track, loop playback setting information, and track description information;

[0093] The positioning methods for the start position of the sound clip in the auxiliary track include absolute time positioning and relative binding positioning. In the absolute time positioning method, the start position of the sound clip in the auxiliary track is located at a specific time defined by the global time axis; in the relative binding positioning method, the start position of the sound clip in the auxiliary track is located at an offset time relative to the anchor point, and the anchor point is provided by the sound clip in the main track.

[0094] Wherein, when loop play is effective, the loop play setting information includes the loop end position; the loop end position is set by the absolute duration after the loop, or by the anchor point provided by the sound clip in the main track;

[0095] The editing display displays an editing view during the mixing process of the mixing processor, wherein the editing view includes a text editing view and a multi-track editing view; the text editing view is bound to a cursor position in the multi-track editing view;

[0096] Wherein, in the editing display, the text editing view and the multi-track editing view are presented simultaneously, or the editing user selects one of the views to be presented and can switch between them;

[0097] The multi-track editing view provides editing users with a visual multi-track audio non-linear editing interface;

[0098] wherein the multi-track editing view presents a primary track and a secondary track, and the timeline of the multi-track editing view is associated with the text editing view via text alignment information of the sound clip in the primary track;

[0099] During the mixing process, the mixing processor obtains the audio material processed by the sound effect processor and adds the obtained audio material to the editing view of the editing display. The added audio material is placed at the position set by the editing user.

[0100] Specifically, when the mixing processor adds the acquired audio material to the multi-track editing view, the added audio material is placed at the position set by the editing user in an absolute time positioning manner;

[0101] After adding an audio material to the multi-track editing view, you can choose to change its positioning method to relative binding positioning;

[0102] Specifically, when the audio mixing processor adds the acquired audio material to the text editing view, the added audio material is placed at the position set by the editing user in a relative binding positioning manner;

[0103] The audio material positioned by absolute time positioning is displayed in the sidebar of the text editing view in a floating form;

[0104] In particular, in the text editing view or the multi-track editing view, the editing user can edit the properties of the added audio material;

[0105] In particular, during the mixing process, if a text segment has not yet been recorded, resulting in a lack of associated recording material for the main audio segment, the duration of the main audio segment on the main track will be determined by the recording engineer's estimated speaking rate and the number of words in the segment in the project's default configuration information, or the editor may select the speech synthesis result as the recording material for temporary use;

[0106] The mixing processor performs a mixing operation on the sound clips according to the script to obtain a mixing processing result; the post-production device generates a post-production result according to the mixing processing result;

[0107] The post-production device generates a post-production result through the created post-production project and outputs the result to the audio presentation device;

[0108] The post-production results include audio files, scripts and anchor points of the mixing processing results;

[0109] The audio presentation device presents the post-production result;

[0110] Specifically, the audio presentation device plays the recording materials and audio materials in the main track and the auxiliary track according to the paragraph presentation order set by the script.

[0111] The present invention discloses a text and audio presentation processing method, comprising:

[0112] The script editing device generates a script; the script includes one or more paragraphs;

[0113] In particular, the script includes recording materials, audio materials, sound effect processing methods, and the order of presentation of paragraphs corresponding to the paragraphs;

[0114] The post-production device obtains the script from the script editing device connected thereto and creates a post-production project based on the script; the post-production project includes sound clips corresponding to paragraphs in the script;

[0115] In particular, the post-production project includes default project configuration information; the default project configuration information includes a default interval value between sound clips in the main track, a target gain setting value for the mixing process, and an estimated speech rate of the sound recordist; wherein the default interval value between sound clips in the main track is a default leading silence duration, and the target gain setting value for the mixing process includes target volume values ​​for vocals, music, and sound effects;

[0116] In particular, the post-production equipment includes an audio effects processor, a mixing processor, a material manager, and an editing display;

[0117] The sound effect processor is connected to the material manager; the sound effect processor is connected to the sound mixing processor; the sound mixing processor is connected to the editing display;

[0118] Wherein, the material manager includes a material library, and the material library stores materials locally;

[0119] The material manager organizes and manages the materials in the material library in a hierarchical structure. The material types in the material library include audio materials and effect materials. Among them, audio materials include but are not limited to sound effect materials and music materials; effect materials include effector setting parameters, such as cave environment effects, broadcast effects, etc.

[0120] The sound effect processor obtains audio material through the material manager connected thereto, applies sound effect processing to the audio material according to the sound effect processing method set by the script, and outputs the audio material after sound effect processing to the mixing processor connected thereto;

[0121] Among them, sound effects include but are not limited to gain, channel mixing, fade in and fade out, equalization, environment, noise reduction, compression, time stretching, and three-dimensional effects;

[0122] In the sound effect processing method, multiple sound effects are connected in series and the order of the sound effects can be adjusted. The sound effects include one or more adjustable operating parameters, and the adjustable operating parameters of the sound effects are set to change according to a specific curve over the playing time.

[0123] The mixing processor includes a main track and an auxiliary track, wherein the main track and the auxiliary track are used to carry sound clips; the mixing processor performs a mixing processing operation on the sound clips in the main track and the auxiliary track according to the script to obtain a mixing processing result;

[0124] The script paragraphs include text paragraphs and audio paragraphs; the sound clips include sound clips in the main track corresponding to the text paragraphs and sound clips in the auxiliary track corresponding to the audio paragraphs;

[0125] The position of the sound clips in the main track is determined by the paragraph presentation order set by the script, and the position of the sound clips in the auxiliary track is set by the editing user;

[0126] The main track is positioned according to the order of presentation of the sound clips it carries, the duration of the paragraphs, and the leading silence;

[0127] In particular, the clip content of the sound clip in the main track includes the association relationship between the sound clip in the main track and the text paragraph in the script, the text content of the associated text paragraph, the association relationship between the sound clip in the main track and the recording material, the associated recording material, and the editing information of the sound clip in the main track;

[0128] In particular, for text paragraphs that have not yet been recorded, since there is no corresponding recording material, the corresponding sound clip content in the main track does not include the association relationship between the sound clip in the main track and the recording material;

[0129] When the sound clip in the main track is associated with the recording material, that is, when the clip content of the sound clip in the main track includes the association relationship between the sound clip in the main track and the recording material, the clip content of the sound clip in the main track also includes text alignment information between the text content and the recording material;

[0130] For a completed script, all the sound clips associated with the text paragraphs are arranged on the main track in the order of paragraph presentation set in the script, and separated from each other by leading silence.

[0131] The editing information of the sound clip in the main track includes the leading silence duration, the starting position of the material playback, the duration of the sound clip in the main track, and the effect setting information;

[0132] In particular, the segment content of the sound segment in the master track includes one or more anchor points;

[0133] The anchor point is set at a position in the text alignment information; or, the anchor point is set at a specific position of the sound clip in the main track, the specific position including but not limited to the beginning and end of the sound clip in the main track; or, the anchor point is set at a relative position based on the recording material set by the editing user;

[0134] The anchor includes semantic information; wherein the semantic information includes a word or phrase in the text, an audio event description, a start time, and an end time;

[0135] The anchor point is used to locate the text content based on semantics; the corresponding text paragraph is located from the audio moment according to the anchor point, or the corresponding audio moment is located from the text paragraph;

[0136] The sound clips in the main track are the main presentation objects of the script and the recording materials corresponding to the text paragraphs in the script during the post-production of the audiobook. That is, the main track usually only contains the recording materials corresponding to the text paragraphs in the script, and the recording materials are recorded by the sound recordist or obtained through speech synthesis.

[0137] The sound clips in the master track can also be used to present audio materials, that is, to place audio materials in the master track;

[0138] Inserting an audio material into the middle of the main audio clip of the master track will convert the audio material into a main audio clip and position it according to the positioning method of the master track.

[0139] In particular, the clip content of the sound clip in the auxiliary track includes the association relationship between the sound clip and the audio material in the auxiliary track, the associated audio material, and the editing information of the sound clip in the auxiliary track;

[0140] In particular, the editing information of the sound clip in the auxiliary track includes the material playback start position, the duration of the sound clip in the auxiliary track, the start position of the sound clip in the auxiliary track, loop playback setting information, and track description information;

[0141] The positioning methods for the start position of the sound clip in the auxiliary track include absolute time positioning and relative binding positioning. In the absolute time positioning method, the start position of the sound clip in the auxiliary track is located at a specific time defined by the global time axis; in the relative binding positioning method, the start position of the sound clip in the auxiliary track is located at an offset time relative to the anchor point, and the anchor point is provided by the sound clip in the main track.

[0142] Wherein, when loop play is effective, the loop play setting information includes the loop end position; the loop end position is set by the absolute duration after the loop, or by the anchor point provided by the sound clip in the main track;

[0143] The editing display displays an editing view during the mixing process of the mixing processor, wherein the editing view includes a text editing view and a multi-track editing view; the text editing view is bound to a cursor position in the multi-track editing view;

[0144] Wherein, in the editing display, the text editing view and the multi-track editing view are presented simultaneously, or the editing user selects one of the views to be presented and can switch between them;

[0145] The multi-track editing view provides editing users with a visual multi-track audio non-linear editing interface;

[0146] wherein the multi-track editing view presents a primary track and a secondary track, and the timeline of the multi-track editing view is associated with the text editing view via text alignment information of the sound clip in the primary track;

[0147] During the mixing process, the mixing processor obtains the audio material processed by the sound effect processor and adds the obtained audio material to the editing view of the editing display. The added audio material is placed at the position set by the editing user.

[0148] Specifically, when the mixing processor adds the acquired audio material to the multi-track editing view, the added audio material is placed at the position set by the editing user in an absolute time positioning manner;

[0149] After adding an audio material to the multi-track editing view, you can choose to change its positioning method to relative binding positioning;

[0150] Specifically, when the audio mixing processor adds the acquired audio material to the text editing view, the added audio material is placed at the position set by the editing user in a relative binding positioning manner;

[0151] The audio material positioned by absolute time positioning is displayed in the sidebar of the text editing view in a floating form;

[0152] In particular, in the text editing view or the multi-track editing view, the editing user can edit the properties of the added audio material;

[0153] In particular, during the mixing process, if a text paragraph has not yet been recorded, resulting in a lack of associated recording material for the main audio clip, the duration of the main audio clip on the main track will be determined by the recording engineer's estimated speaking rate and the number of words in the paragraph in the project's default configuration information, or the editor may select the speech synthesis result as the recording material for temporary use;

[0154] The mixing processor performs a mixing operation on the sound clips according to the script to obtain a mixing processing result; the post-production device generates a post-production result according to the mixing processing result;

[0155] The post-production device performs post-production processing on the sound clip through a post-production project, generates a post-production result and outputs it to the audio presentation device connected thereto;

[0156] The post-production results include audio files, scripts and anchor points of the mixing processing results;

[0157] The audio presentation device presents the post-production result;

[0158] Specifically, the audio presentation device plays the recording materials and audio materials in the main track and the auxiliary track according to the paragraph presentation order set by the script.

[0159] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A text and audio presentation processing method, characterized in that: include: The script editing device generates a script; the script includes one or more paragraphs; The post-production device obtains a script from a script editing device connected thereto and creates a post-production project based on the script; the post-production project includes sound clips corresponding to paragraphs in the script; the post-production device performs post-production processing on the sound clips through the post-production project, generates a post-production result, and outputs it to an audio presentation device connected thereto; The audio presentation device presents the post-production result; The script includes the recording materials, audio materials, sound effect processing methods, and the order of paragraph presentation corresponding to the paragraphs; The post-production equipment includes a sound effect processor, a mixing processor, a material manager, and an editing display; The material manager includes a material library, which stores materials locally; the material manager organizes and manages the materials in the material library in a hierarchical structure, and the material types in the material library include audio materials and effect materials; The sound effect processor obtains audio material through the material manager connected thereto, applies sound effect processing to the audio material according to the sound effect processing method set by the script, and outputs the audio material after sound effect processing to the mixing processor connected thereto; The mixing processor includes a main track and an auxiliary track, wherein the main track and the auxiliary track are used to carry sound clips; the mixing processor performs a mixing processing operation on the sound clips in the main track and the auxiliary track according to the script to obtain a mixing processing result; The script paragraphs include text paragraphs and audio paragraphs; the sound clips include sound clips in the main track corresponding to the text paragraphs and sound clips in the auxiliary track corresponding to the audio paragraphs; The position of the sound clips in the main track is determined by the paragraph presentation order set by the script, and the position of the sound clips in the auxiliary track is set by the editing user; The editing display displays an editing view during the mixing process of the mixing processor, wherein the editing view includes a text editing view and a multi-track editing view; the text editing view is bound to a cursor position in the multi-track editing view; In the editing display, the text editing view and the multi-track editing view are presented simultaneously, or the editing user selects one of the views to be presented and can switch between them.

2. The text and audio presentation processing method according to claim 1, characterized in that: in, The content of the sound clip in the main track includes the association between the sound clip in the main track and the text paragraph in the script, the text content of the associated text paragraph, the association between the sound clip in the main track and the recording material, the associated recording material, and editing information of the sound clip in the main track; When a sound clip in the main track is associated with an audio recording, the clip content of the sound clip in the main track also includes text alignment information between the text content and the audio recording. For a script that has completed the recording of the audio recording, the sound clips associated with all text paragraphs are arranged on the main track in the paragraph presentation order set by the script, and are separated by leading silence. wherein the multi-track editing view presents a primary track and a secondary track, and the timeline of the multi-track editing view is associated with the text editing view via text alignment information of the sound clip in the primary track; The editing information of the sound clip in the main track includes the leading silence duration, the starting position of the material playback, the duration of the sound clip in the main track, and the effect setting information; The segment content of the sound clip in the main track includes one or more anchor points; wherein the anchor points are set at positions in the text alignment information; or, the anchor points are set at specific positions of the sound clip in the main track, and the specific positions include the start and end of the sound clip in the main track; or, the anchor points are set at relative positions based on the recording material set by the editing user; wherein the anchor points include semantic information, and the semantic information includes words or phrases in the text, audio event descriptions, start times, and end times.

3. The text and audio presentation processing method according to claim 2, characterized in that: in, The clip content of the sound clip in the auxiliary track includes the association relationship between the sound clip and the audio material in the auxiliary track, the associated audio material, and the editing information of the sound clip in the auxiliary track; The editing information of the sound clip in the auxiliary track includes the material playback start position, the duration of the sound clip in the auxiliary track, the sound clip start position in the auxiliary track, the loop playback setting information, and the track description information; The positioning methods for the start position of the sound clip in the auxiliary track include absolute time positioning and relative binding positioning. In the absolute time positioning method, the start position of the sound clip in the auxiliary track is located at a specific time defined by the global time axis; in the relative binding positioning method, the start position of the sound clip in the auxiliary track is located at an offset time relative to the anchor point, and the anchor point is provided by the sound clip in the main track. Wherein, when loop play is effective, the loop play setting information includes the loop end position; the loop end position is set by the absolute duration after the loop, or by the anchor point provided by the sound clip in the main track; During the mixing process, the mixing processor obtains the audio material processed by the sound effect processor and adds the obtained audio material to the editing view of the editing display. The added audio material is placed at the position set by the editing user.

4. The text and audio presentation processing method according to claim 1, characterized in that: in, A post-production project includes default project configuration information; this includes the default interval between sound clips in the main track, the target gain setting for the mixing process, and the estimated speech rate of the sound recordist. The default interval between sound clips in the main track includes the default leading silence duration, and the target gain setting for the mixing process includes the target volume values ​​for vocals, music, and sound effects. During the mixing process, if a text paragraph has not yet been recorded and the main audio clip lacks associated recording material, the duration of the main audio clip on the main track is determined by the recording engineer's estimated speaking speed and the number of words in the paragraph in the project's default configuration information, or the editor selects the speech synthesis result as the recording material for temporary use.

5. A text and audio presentation processing system, characterized in that: It includes script editing equipment, post-production equipment, and audio presentation equipment that are interconnected in sequence; Wherein, the script editing device generates a script; the script includes one or more paragraphs; The post-production device obtains the script and creates a post-production project according to the script; the post-production project includes sound clips, and the sound clips correspond to paragraphs in the script; The post-production device performs post-production processing on the sound clip through a post-production project, generates a post-production result and outputs it to the audio presentation device; The audio presentation device presents the post-production result; The script includes the recording materials, audio materials, sound effect processing methods, and the order of paragraph presentation corresponding to the paragraphs; The post-production equipment includes a sound effect processor, a sound mixing processor, a material manager, and an editing display; the sound effect processor is connected to the material manager; the sound effect processor is connected to the sound mixing processor; the sound mixing processor is connected to the editing display; The material manager includes a material library, which stores materials locally; the material manager organizes and manages the materials in the material library in a hierarchical structure, and the material types in the material library include audio materials and effect materials; The sound effect processor obtains the audio material through the material manager, applies a sound effect processing operation to the audio material according to the sound effect processing method set by the script, and outputs the audio material after the sound effect processing to the mixing processor; The mixing processor includes a main track and an auxiliary track, wherein the main track and the auxiliary track are used to carry sound clips; the mixing processor performs a mixing processing operation on the sound clips in the main track and the auxiliary track according to the script to obtain a mixing processing result; The script paragraphs include text paragraphs and audio paragraphs; the sound clips include sound clips in the main track corresponding to the text paragraphs and sound clips in the auxiliary track corresponding to the audio paragraphs; The position of the sound clips in the main track is determined by the paragraph presentation order set by the script, and the position of the sound clips in the auxiliary track is set by the editing user; The editing display displays an editing view during the mixing process of the mixing processor, wherein the editing view includes a text editing view and a multi-track editing view; the text editing view is bound to a cursor position in the multi-track editing view; In the editing display, the text editing view and the multi-track editing view are presented simultaneously, or the editing user selects one of the views to be presented and can switch between them.

6. The text and audio presentation processing system according to claim 5, characterized in that: in, The content of the sound clip in the main track includes the association between the sound clip in the main track and the text paragraph in the script, the text content of the associated text paragraph, the association between the sound clip in the main track and the recording material, the associated recording material, and editing information of the sound clip in the main track; When a sound clip in the main track is associated with an audio recording, the clip content of the sound clip in the main track also includes text alignment information between the text content and the audio recording. For a script that has completed the recording of the audio recording, the sound clips associated with all text paragraphs are arranged on the main track in the paragraph presentation order set by the script, and are separated by leading silence. wherein the multi-track editing view presents a primary track and a secondary track, and the timeline of the multi-track editing view is associated with the text editing view via text alignment information of the sound clip in the primary track; The editing information of the sound clip in the main track includes the leading silence duration, the starting position of the material playback, the duration of the sound clip in the main track, and the effect setting information; The segment content of the sound clip in the main track includes one or more anchor points; wherein the anchor points are set at positions in the text alignment information; or, the anchor points are set at specific positions of the sound clip in the main track, and the specific positions include the start and end of the sound clip in the main track; or, the anchor points are set at relative positions based on the recording material set by the editing user; wherein the anchor points include semantic information, and the semantic information includes words or phrases in the text, audio event descriptions, start time, and end time.

7. The text and audio presentation processing system according to claim 6, characterized in that: in, The clip content of the sound clip in the auxiliary track includes the association relationship between the sound clip and the audio material in the auxiliary track, the associated audio material, and the editing information of the sound clip in the auxiliary track; The editing information of the sound clip in the auxiliary track includes the material playback start position, the duration of the sound clip in the auxiliary track, the sound clip start position in the auxiliary track, the loop playback setting information, and the track description information; The positioning methods for the start position of the sound clip in the auxiliary track include absolute time positioning and relative binding positioning. In the absolute time positioning method, the start position of the sound clip in the auxiliary track is located at a specific time defined by the global time axis; in the relative binding positioning method, the start position of the sound clip in the auxiliary track is located at an offset time relative to the anchor point, and the anchor point is provided by the sound clip in the main track. Wherein, when loop play is effective, the loop play setting information includes the loop end position; the loop end position is set by the absolute duration after the loop, or by the anchor point provided by the sound clip in the main track; During the mixing process, the mixing processor obtains the audio material processed by the sound effect processor and adds the obtained audio material to the editing view of the editing display. The added audio material is placed at the position set by the editing user.

8. The text and audio presentation processing system according to claim 5, characterized in that: in, A post-production project includes default project configuration information; this includes the default interval between sound clips in the main track, the target gain setting for the mixing process, and the estimated speech rate of the sound recordist. The default interval between sound clips in the main track includes the default leading silence duration, and the target gain setting for the mixing process includes the target volume values ​​for vocals, music, and sound effects. During the mixing process, if a text paragraph has not yet been recorded and the main audio clip lacks associated recording material, the duration of the main audio clip on the main track is determined by the recording engineer's estimated speaking speed and the number of words in the paragraph in the project's default configuration information, or the editor selects the speech synthesis result as the recording material for temporary use.

Citation Information

Patent Citations

  • Recording editing management method and system

    CN111508468A