Audio acquisition system and method
Patent Information
- Application Number
- CN202211723488.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2042-12-30
AI Technical Summary
[0003]发明人经研究发现,现有技术中有声读物的制作不是系统化、流程化的,尤其是将现有的文本例如小说或者剧本转换为配音剧本的过程需要花费大量人工和时间,而根据配音剧本进行配音所得到的最终声音作品的质量也由于不是系统化、流程化的产出而变得极不可控,不同的配音部分之间的质量参差不齐,导致读者的用户体验大大下降,可见现有技术中将文本转化为配音剧本并由配音演员根据配音剧本进行配音来制作有声读物的这种技术方案的生产效率低、灵活性差
本发明通过剧本编辑子系统完成原始文本向配音剧本的编辑转换,通过配音管理子系统实现对配音任务的分配和管理,通过配音工作子系统实现配音演员的配音演绎,通过音频分析子系统自动对配音演绎的内容进行质量检验及反馈,从而高效率且高质量地实现了在有声读物制作中由文本向音频转换的制作生成流程。
Smart Images

Figure CN116092468B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to an audio acquisition system and method. Background Technology
[0002] With the development of technology and the diversification of information access channels, digital media is constantly impacting traditional media, and the public's reading habits are constantly changing, giving rise to audiobooks as a reading format. Audiobooks are electronic publications that use sound as the medium, including audio news, audio novels, and so on. Audiobooks overlap with and differ from both digital and traditional media, and their unique advantages can meet the needs of various users.
[0003] The inventors discovered through research that the production of audiobooks in existing technologies is not systematic or procedural. In particular, the process of converting existing texts, such as novels or scripts, into dubbing scripts requires a lot of manpower and time. Furthermore, the quality of the final audio works obtained by dubbing based on the dubbing scripts is extremely uncontrollable due to the lack of a systematic and procedural output. The quality of different dubbing parts varies, resulting in a significant decline in the user experience. It is evident that the existing technology of converting texts into dubbing scripts and having voice actors dub based on the scripts to produce audiobooks is inefficient and lacks flexibility. Summary of the Invention
[0004] Based on this, and to solve the technical problems in the existing technology, an audio acquisition method is proposed, including: The script editing subsystem receives the input text to be edited. The script editor edits the content of the text through the script editing subsystem to obtain the dubbing script, and sends the edited dubbing script to the dubbing management subsystem and the dubbing work subsystem. The dubbing management subsystem receives the dubbing scripts from the editors, and the dubbing director allocates and manages dubbing tasks according to the dubbing scripts in the dubbing management subsystem; The dubbing subsystem receives the edited dubbing script and includes a recording device. Voice actors perform dubbing in the dubbing subsystem according to the dubbing script and record their voices using the recording device. The recording device converts the voice actors' dubbing performances into an audio stream. The dubbing work subsystem sends the audio stream to the dubbing management subsystem; the dubbing director reviews the recording content in the audio stream submitted by the dubbing actor through the dubbing management subsystem and decides whether to accept the submitted recording; when the dubbing director fails to approve the recording, the dubbing management subsystem returns the task and corresponding rework guidance to the dubbing work subsystem, and the corresponding dubbing actor carries out the rework. Voice actors view rework tasks in the voice-over work subsystem, re-enact and record their performances according to the rework guidelines, and then resubmit the re-enacted performances to the voice-over management subsystem for further review and a decision on whether to adopt them.
[0005] In one embodiment, the script editor is a screenwriter or a voice director; The organizational structure of the text to be edited, from largest to smallest, is: book, chapter, paragraph; A voice-over script includes a set of characters and a set of voice actors; within the set of characters, there is a one-to-many relationship between characters and segments; within the set of voice actors, there is a one-to-many relationship between voice actors and characters; and there is a one-to-many relationship between voice actors and segments. The script editor uses the script editing subsystem to configure the correspondence between characters and voice actors, and between paragraphs and characters in the dubbing script, and provides corresponding performance guidance in the dubbing script for the paragraphs that the voice actors need to perform; The organizational structure of a dubbing script, from largest to smallest, is as follows: book, chapter. For each chapter of the dubbing script, once the correspondence between the assigned paragraphs, characters, and voice actors is determined, the dubbing script content for that chapter can be edited. The dubbing script content for each chapter includes the content of each paragraph in that chapter, the character corresponding to the paragraph, the voice actor corresponding to the character, and performance guidance for the paragraph and character.
[0006] In one embodiment, a character played by a voice actor and a corresponding segment constitute a task; the task includes task status and task information. Task status includes Incomplete, Rejected, Completed, and Pending Rework; Incomplete status means the voice actor has not completed the voice acting and recording; Rejected status means the voice actor has found an error in the script and refused to perform the voice acting; Completed status means the voice actor has completed the voice acting and recording; Pending Rework status means the voice actor has completed the voice acting and recording, but the submitted recording result has not passed the voice director's review. The task information includes the paragraph content, the performance guidance given by the dubbing director, and the rework guidance given by the dubbing director; A voice actor's portrayal of a character in a chapter corresponds to multiple segments. The collection of multiple segments corresponding to multiple tasks constitutes a batch. The voice acting script of a chapter consists of multiple batches. Each batch includes batch information, which includes the character attributes corresponding to the batch and the completion status of the tasks in the batch. The character attributes corresponding to a batch include character name, character type, character gender, character age, and character characteristic description; the completion status of tasks in a batch includes the number of tasks completed, the number of tasks rejected, the number of tasks awaiting rework, and the total number of tasks; The organizational structure of a dubbing task, from largest to smallest, is: book, batch, task; The dubbing director reviews tasks rejected by voice actors through the dubbing management subsystem; if the rejection is reasonable, the dubbing director re-edits the dubbing script through the script editing subsystem and then reassigns the voice actors' tasks; if the rejection is unreasonable, the dubbing director directly reassigns the voice actors' tasks in the dubbing management subsystem.
[0007] In one embodiment, each voice actor is configured with a corresponding book list in the voice acting work subsystem, and the book list contains all the books in which the voice actor has voice acting tasks. The books for which voice actors participate in dubbing are configured with a batch list in the dubbing work subsystem; Voice actors can index batches and voice-over tasks in the batch list by chapter or character in the voice-over work subsystem; Voice actors sort the batches in the batch list by chapter or character. Voice actors select batches and execute tasks within those batches in the voice-over work subsystem; voice actors perform the corresponding segments by task and record them using a recording device; the voice-over work subsystem submits the audio streams to the audio analysis subsystem by task. When a voice actor discovers an error in the script that prevents them from completing the performance, they should refuse to perform the task and submit the reason for refusal to the voice acting management subsystem.
[0008] In one embodiment, before submitting the audio stream to the dubbing management subsystem, the dubbing work subsystem submits the audio stream to the audio analysis subsystem; the audio analysis subsystem analyzes and performs quality checks on the audio stream, and feeds back the quality check results to the dubbing work subsystem. Voice actors receive and refer to the quality inspection results fed back by the audio analysis subsystem through the voice acting work subsystem. If the voice actor believes that the voice acting performance does not need adjustment, he / she will directly submit the quality inspection results fed back by the recording and audio analysis subsystem to the voice acting management subsystem. If the voice actor believes that the voice acting performance needs adjustment, he / she will adjust the recording environment or performance method according to the quality inspection information and perform the voice acting performance again. The voice acting work subsystem will submit the audio stream generated by the adjusted recording to the audio analysis subsystem. The quality inspection and adjustment process is repeated until the voice actor is satisfied. Then, the quality inspection results fed back by the recording and audio analysis subsystem will be submitted to the voice acting management subsystem. The dubbing director reviews the recordings submitted by the voice actors through the dubbing management subsystem and, with reference to the feedback quality inspection results, decides whether to adopt the submitted recordings. The audio analysis subsystem includes a signal analysis device, a content analysis device, and a quality detection device. The signal analysis device analyzes the audio stream and obtains the human voice attributes and environmental attributes of the audio stream; among which, the human voice attributes include human voice energy and human voice duration, and the environmental attributes include noise energy and spatial reverberation parameters; The content analysis device analyzes the audio stream and obtains the human voice content in the signal; wherein, the human voice content includes speech human voice content described by text symbols and non-speech human voice content described by sound symbols; the text symbols and sound symbols are accompanied by start and end timestamp information, and the duration of occupancy of the text symbols or sound symbols is obtained based on the timestamp information; The quality inspection device performs quality inspection based on the attributes of human voice, environmental attributes, and human voice content, and feeds back the quality inspection results to the dubbing subsystem. The quality inspection results include whether the quality inspection is passed, the quality inspection score, and the corresponding quality inspection error type, error content, and error correction method prompts.
[0009] In addition, to address the problems in existing technologies, an audio acquisition system is proposed, comprising a script editing subsystem, a dubbing management subsystem, and a dubbing work subsystem; The script editing subsystem receives the input text to be edited. The script editor edits the content of the text through the script editing subsystem to obtain the dubbing script, and sends the edited dubbing script to the dubbing management subsystem and the dubbing work subsystem. The dubbing management subsystem receives the dubbing scripts edited by the dubbing director, who then allocates and manages dubbing tasks based on the scripts within the dubbing management subsystem. The dubbing subsystem receives the edited dubbing script and includes a recording device. Voice actors perform dubbing according to the script and record their voices using the recording device. The recording device converts the voice actors' voices into an audio stream. The dubbing subsystem then sends the audio stream to the dubbing management subsystem. The dubbing director reviews the audio recordings submitted by the voice actors through the dubbing management subsystem and decides whether to accept the submitted recordings. When the dubbing director fails to approve the recordings, the dubbing management subsystem returns the task and corresponding rework instructions to the dubbing work subsystem, whereby the corresponding voice actor will carry out the rework. Voice actors view rework tasks in the voice-over work subsystem, re-enact and record their performances according to the rework guidelines, and then resubmit the re-enacted performances to the voice-over management subsystem for further review and a decision on whether to adopt them.
[0010] In one embodiment, the script editor is a screenwriter or a voice director; The organizational structure of the text to be edited, from largest to smallest, is: book, chapter, paragraph; A voice-over script includes a set of characters and a set of voice actors; within the set of characters, there is a one-to-many relationship between characters and segments; within the set of voice actors, there is a one-to-many relationship between voice actors and characters; and there is a one-to-many relationship between voice actors and segments. The script editor uses the script editing subsystem to configure the correspondence between characters and voice actors, and between paragraphs and characters in the dubbing script, and provides corresponding performance guidance in the dubbing script for the paragraphs that the voice actors need to perform; The organizational structure of a dubbing script, from largest to smallest, is as follows: book, chapter. For each chapter of the dubbing script, once the correspondence between the assigned paragraphs, characters, and voice actors is determined, the dubbing script content for that chapter can be edited. The dubbing script content for each chapter includes the content of each paragraph in that chapter, the character corresponding to the paragraph, the voice actor corresponding to the character, and performance guidance for the paragraph and character.
[0011] In one embodiment, a character played by a voice actor and a corresponding segment constitute a task; the task includes task status and task information. Task status includes Incomplete, Rejected, Completed, and Pending Rework; Incomplete status means the voice actor has not completed the voice acting and recording; Rejected status means the voice actor has found an error in the script and refused to perform the voice acting; Completed status means the voice actor has completed the voice acting and recording; Pending Rework status means the voice actor has completed the voice acting and recording, but the submitted recording result has not passed the voice director's review. The task information includes the paragraph content, the performance guidance given by the dubbing director, and the rework guidance given by the dubbing director; A voice actor's portrayal of a character in a chapter corresponds to multiple segments. The collection of multiple segments corresponding to multiple tasks constitutes a batch. The voice acting script of a chapter consists of multiple batches. Each batch includes batch information, which includes the character attributes corresponding to the batch and the completion status of the tasks in the batch. The character attributes corresponding to a batch include character name, character type, character gender, character age, and character characteristic description; the completion status of tasks in a batch includes the number of tasks completed, the number of tasks rejected, the number of tasks awaiting rework, and the total number of tasks; The organizational structure of a dubbing task, from largest to smallest, is: book, batch, task; The dubbing director reviews tasks rejected by voice actors through the dubbing management subsystem; if the rejection is reasonable, the dubbing director re-edits the dubbing script through the script editing subsystem and then reassigns the voice actors' tasks; if the rejection is unreasonable, the dubbing director directly reassigns the voice actors' tasks in the dubbing management subsystem.
[0012] In one embodiment, each voice actor is configured with a corresponding book list in the voice acting work subsystem, and the book list contains all the books in which the voice actor has voice acting tasks. The books for which voice actors participate in dubbing are configured with a batch list in the dubbing work subsystem; Voice actors can index batches and voice-over tasks in the batch list by chapter or character in the voice-over work subsystem; Voice actors sort the batches in the batch list by chapter or character. Voice actors select batches and execute tasks within those batches in the voice-over work subsystem; voice actors perform the corresponding segments by task and record them using a recording device; the voice-over work subsystem submits the audio streams to the audio analysis subsystem by task. When a voice actor discovers an error in the script that prevents them from completing the performance, they should refuse to perform the task and submit the reason for refusal to the voice acting management subsystem.
[0013] In one embodiment, the audio acquisition system includes an audio analysis subsystem; before submitting the audio stream to the dubbing management subsystem, the dubbing work subsystem submits the audio stream to the audio analysis subsystem; the audio analysis subsystem analyzes and performs quality checks on the audio stream, and feeds back the quality check results to the dubbing work subsystem; Voice actors receive and refer to the quality inspection results fed back by the audio analysis subsystem through the voice acting work subsystem. If the voice actor believes that the voice acting performance does not need adjustment, he / she will directly submit the quality inspection results fed back by the recording and audio analysis subsystem to the voice acting management subsystem. If the voice actor believes that the voice acting performance needs adjustment, he / she will adjust the recording environment or performance method according to the quality inspection information and perform the voice acting performance again. The voice acting work subsystem will submit the audio stream generated by the adjusted recording to the audio analysis subsystem. The quality inspection and adjustment process is repeated until the voice actor is satisfied. Then, the quality inspection results fed back by the recording and audio analysis subsystem will be submitted to the voice acting management subsystem. The dubbing director reviews the recordings submitted by the voice actors through the dubbing management subsystem and, with reference to the feedback quality inspection results, decides whether to adopt the submitted recordings. The audio analysis subsystem includes a signal analysis device, a content analysis device, and a quality detection device. The signal analysis device analyzes the audio stream and obtains the human voice attributes and environmental attributes of the audio stream; among which, the human voice attributes include human voice energy and human voice duration, and the environmental attributes include noise energy and spatial reverberation parameters; The content analysis device analyzes the audio stream and obtains the human voice content in the signal; wherein, the human voice content includes speech human voice content described by text symbols and non-speech human voice content described by sound symbols; the text symbols and sound symbols are accompanied by start and end timestamp information, and the duration of occupancy of the text symbols or sound symbols is obtained based on the timestamp information; The quality inspection device performs quality inspection based on the attributes of human voice, environmental attributes, and human voice content, and feeds back the quality inspection results to the dubbing subsystem. The quality inspection results include whether the quality inspection is passed, the quality inspection score, and the corresponding quality inspection error type, error content, and error correction method prompts.
[0014] Implementing the embodiments of the present invention will have the following beneficial effects: This invention completes the editing and conversion of the original text into a dubbing script through a script editing subsystem, realizes the allocation and management of dubbing tasks through a dubbing management subsystem, realizes the dubbing performance of dubbing actors through a dubbing work subsystem, and automatically performs quality inspection and feedback on the dubbing performance through an audio analysis subsystem. Thus, it realizes a production process of converting text into audio in audiobook production with high efficiency and high quality. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] in: Figure 1 This is a schematic diagram of the audio acquisition system in this invention; Figure 2 This is a flowchart illustrating the audio acquisition method of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] This invention discloses an audio acquisition system, including a script editing subsystem, a dubbing management subsystem, and a dubbing work subsystem; The script editing subsystem receives the input text to be edited. The script editor edits the content of the text through the script editing subsystem to obtain the dubbing script, and sends the edited dubbing script to the dubbing management subsystem and the dubbing work subsystem. A script editor is either a screenwriter or a voice director; The organizational structure of the text to be edited, from largest to smallest, is: book, chapter, paragraph; A voice-over script includes a set of characters and a set of voice actors; within the set of characters, there is a one-to-many relationship between characters and segments; within the set of voice actors, there is a one-to-many relationship between voice actors and characters; and there is a one-to-many relationship between voice actors and segments. The script editor uses the script editing subsystem to configure the correspondence between characters and voice actors, and between paragraphs and characters in the dubbing script, and provides corresponding performance guidance in the dubbing script for the paragraphs that the voice actors need to perform; The organizational structure of a dubbing script, from largest to smallest, is as follows: book, chapter. For each chapter of the dubbing script, once the correspondence between the assigned paragraphs, characters, and voice actors is determined, the dubbing script content for that chapter can be edited. The dubbing script content for each chapter includes the content of each paragraph in that chapter, the character corresponding to the paragraph, the voice actor corresponding to the character, and performance guidance for the paragraph and character. The dubbing management subsystem receives the dubbing scripts edited by the dubbing director, who then allocates and manages dubbing tasks based on the scripts within the dubbing management subsystem. A task consists of a character portrayed by a voice actor and a corresponding segment; a task includes task status and task information. Task status includes Incomplete, Rejected, Completed, and Pending Rework; Incomplete status means the voice actor has not completed the voice acting and recording; Rejected status means the voice actor has found an error in the script and refused to perform the voice acting; Completed status means the voice actor has completed the voice acting and recording; Pending Rework status means the voice actor has completed the voice acting and recording, but the submitted recording result has not passed the voice director's review. The task information includes the paragraph content, the performance guidance given by the dubbing director, and the rework guidance given by the dubbing director; A voice actor's portrayal of a character in a chapter corresponds to multiple segments. The collection of multiple segments corresponding to multiple tasks constitutes a batch. The voice acting script of a chapter consists of multiple batches. Each batch includes batch information, which includes the character attributes corresponding to the batch and the completion status of the tasks in the batch. The character attributes corresponding to a batch include character name, character type, character gender, character age, and character characteristic description; the completion status of tasks in a batch includes the number of tasks completed, the number of tasks rejected, the number of tasks awaiting rework, and the total number of tasks; The organizational structure of a dubbing task, from largest to smallest, is: book, batch, task; The dubbing director reviews tasks rejected by voice actors through the dubbing management subsystem; if the rejection is reasonable, the dubbing director re-edits the dubbing script through the script editing subsystem and then reassigns the voice actors' tasks; if the rejection is unreasonable, the dubbing director directly reassigns the voice actors' tasks in the dubbing management subsystem. The dubbing subsystem receives the edited dubbing script and includes a recording device. The dubbing subsystem performs dubbing according to the dubbing script and records the performance using the recording device. The recording device converts the voices of the dubbing actors into an audio stream. Each voice actor has a corresponding book list configured in the voice acting work subsystem, which contains all the books in which the voice actor has voice acting assignments. The books for which voice actors participate in dubbing are configured with a batch list in the dubbing work subsystem; Voice actors can index batches and voice-over tasks in the batch list by chapter or character in the voice-over work subsystem; Voice actors sort the batches in the batch list by chapter or character. Voice actors select a batch and execute tasks within that batch in the voice-over work subsystem. For a specific batch, voice actors can view the batch information and the dubbing scripts for the corresponding chapters through the dubbing work subsystem; voice actors can also view the task information belonging to this batch. Voice actors perform corresponding segments in units of tasks and record their performances using a recording device, which converts the voice actors' performances into an audio stream. The dubbing work subsystem sends the audio stream to the dubbing management subsystem; the dubbing director reviews the recording content in the audio stream submitted by the dubbing actors through the dubbing management subsystem and decides whether to adopt the submitted recording; Specifically, the audio acquisition system includes an audio analysis subsystem; before submitting the audio stream to the dubbing management subsystem, the dubbing work subsystem submits the audio stream to the audio analysis subsystem; Specifically, the dubbing subsystem submits audio streams to the audio analysis subsystem on a task-by-task basis; Specifically, if a voice actor discovers an error in the script that prevents them from completing the performance, they can refuse to perform the task and submit the reason for refusal to the voice acting management subsystem. Among them, the audio analysis subsystem analyzes and performs quality inspection on the audio stream, and feeds back the quality inspection results to the voice actors on the dubbing work subsystem side; The audio analysis subsystem includes a signal analysis device, a content analysis device, and a quality detection device; The signal analysis device analyzes the audio stream and obtains the human voice attributes and environmental attributes of the audio stream; among which, the human voice attributes include human voice energy and human voice duration, and the environmental attributes include noise energy and spatial reverberation parameters; Specifically, the signal analysis device uses a human voice activation detection method to divide the audio stream into a human voice part and a non-human voice part, and calculates the energy and duration of the human voice part, while also calculating the energy of the non-human voice part; wherein, the non-human voice part includes noise; Among them, the signal analysis device calculates the energy using the root mean square of the signal, or calculates the energy using a loudness model based on a psychoacoustic model; The signal analysis device uses machine learning regression algorithms to calculate the spatial reverberation parameters in the audio stream; Alternatively, the signal analysis device can use a synthetic approach to analyze the audio stream and use a generative model to iteratively approximate the spatial reverberation effect, thereby calculating the spatial reverberation parameters. Among them, spatial reverberation parameters include reverberation time; The content analysis device analyzes the audio stream and extracts the human voice content from the signal; Human voice content includes spoken human voice content described by written symbols and non-speech human voice content described by sound symbols; Non-voice human voice content includes breathing, laughter, crying, and vocalization sounds; Text symbols and sound symbols are accompanied by start and end timestamp information, and the duration of occupancy of text symbols or sound symbols can be obtained based on the timestamp information; Specifically, the content analysis device analyzes the audio stream and obtains human voice content through speech recognition; The quality inspection device performs quality inspection based on human voice attributes, environmental attributes, and human voice content, and feeds back the quality inspection results to the dubbing work subsystem. The quality inspection results include whether the quality inspection is passed, the quality inspection score, and the corresponding quality inspection error type, error content, and error correction method prompts. Among them, a rule-based method is used to determine whether the sound meets the quality inspection requirements; Alternatively, a machine learning-based method can be used to determine whether the sound meets the quality inspection requirements using a classification model. Alternatively, a machine learning-based approach can be used to obtain a sound quality score using a regression model. Among them, voice actors receive and refer to the quality inspection results fed back by the quality inspection device through the voice acting work subsystem; If the voice actor believes that the dubbing performance does not need adjustment, the quality inspection results fed back by the recording and audio analysis subsystem will be directly submitted to the dubbing management subsystem; if the voice actor believes that the dubbing performance needs adjustment, the recording environment or performance method will be adjusted with reference to the quality inspection information and the dubbing performance will be performed again. The dubbing work subsystem will submit the audio stream generated by the adjusted recording to the audio analysis subsystem. The quality inspection and adjustment process will be repeated until the voice actor is satisfied. Then the quality inspection results fed back by the recording and audio analysis subsystem will be submitted to the dubbing management subsystem. The dubbing director reviews the recordings submitted by the voice actors through the dubbing management subsystem and, with reference to the feedback quality inspection results, decides whether to adopt the submitted recordings. When the dubbing director fails to approve the dubbing, the dubbing management subsystem will return the task and corresponding rework instructions to the dubbing work subsystem, and the corresponding dubbing actor will carry out the rework. Voice actors view rework tasks in the voice-over work subsystem, and re-enact and record their performances according to the rework guidelines. The voice-over work subsystem then submits the re-enacted recordings to the voice-over management subsystem and the audio analysis subsystem, where the voice-over director reviews them and decides whether to adopt them.
[0019] In addition, the present invention also discloses an audio acquisition method, comprising: The script editing subsystem receives the input text to be edited. The dubbing director edits the content of the text through the script editing subsystem to obtain the dubbing script, and sends the edited dubbing script to the dubbing management subsystem and the dubbing work subsystem. A script editor is either a screenwriter or a voice director; The organizational structure of the text to be edited, from largest to smallest, is: book, chapter, paragraph; The dubbing script includes a set of characters and a set of voice actors; the relationship between characters and segments in the character set is one-to-many; the relationship between voice actors and characters in the voice actor set is one-to-many; and the relationship between voice actors and segments is one-to-many. The script editor uses the script editing subsystem to configure the correspondence between characters and voice actors, and between paragraphs and characters in the dubbing script, and provides corresponding performance guidance in the dubbing script for the paragraphs that the voice actors need to perform; The dubbing script is organized from largest to smallest as a book and then a chapter. Once the correspondence between the paragraphs, characters, and voice actors is determined, the dubbing script content for that chapter can be edited. The dubbing script content for each chapter includes the content of each paragraph in that chapter, the character corresponding to the paragraph, the voice actor corresponding to the character, and performance guidance for the paragraph and character. The dubbing management subsystem receives the dubbing scripts from the editors, and the dubbing director allocates and manages dubbing tasks according to the dubbing scripts in the dubbing management subsystem; A task consists of a character played by a voice actor and a corresponding segment; a task includes task status and task information. Task status includes Incomplete, Rejected, Completed, and Pending Rework; Incomplete status means the voice actor has not completed the voice acting and recording; Rejected status means the voice actor has found an error in the script and refused to perform the voice acting; Completed status means the voice actor has completed the voice acting and recording; Pending Rework status means the voice actor has completed the voice acting and recording, but the submitted recording result has not passed the voice director's review. The task information includes the paragraph content, the performance guidance given by the dubbing director, and the rework guidance given by the dubbing director; In this context, a voice actor's portrayal of a character in a chapter corresponds to multiple segments, and the collection of multiple segments corresponding to multiple tasks constitutes a batch. The voice acting script for a chapter consists of multiple batches. Each batch includes batch information, which includes the character attributes corresponding to the batch and the completion status of the tasks within the batch. The character attributes corresponding to a batch include character name, character type, character gender, character age, and character characteristic description; the completion status of tasks in a batch includes the number of tasks completed, the number of tasks rejected, the number of tasks awaiting rework, and the total number of tasks; The organizational structure of a dubbing task, from largest to smallest, is: book, batch, task; The dubbing director reviews tasks rejected by voice actors through the dubbing management subsystem; if the rejection is reasonable, the dubbing director re-edits the dubbing script through the script editing subsystem and then reassigns the voice actors' tasks; if the rejection is unreasonable, the dubbing director directly reassigns the voice actors' tasks in the dubbing management subsystem. The dubbing subsystem receives the edited dubbing script and includes a recording device. Voice actors perform dubbing according to the dubbing script in the dubbing subsystem and record their voices using the recording device. The recording device converts the voice actors' voices into an audio stream. Each voice actor has a corresponding book list configured in the voice acting work subsystem, which contains all the books in which the voice actor has voice acting assignments. The books for which voice actors participate in dubbing are configured with a batch list in the dubbing work subsystem; Voice actors can index batches and voice-over tasks in the batch list by chapter or character in the voice-over work subsystem; Voice actors sort the batches in the batch list by chapter or character. Voice actors select a batch and execute tasks within that batch in the voice-over work subsystem. For a specific batch, voice actors can view the batch information and the dubbing scripts for the corresponding chapters through the dubbing work subsystem; voice actors can also view the task information belonging to this batch. Voice actors perform corresponding segments in units of tasks and record their performances using a recording device, which converts the voice actors' performances into an audio stream. The dubbing work subsystem sends the audio stream to the dubbing management subsystem; the dubbing director reviews the recording content in the audio stream submitted by the dubbing actors through the dubbing management subsystem and decides whether to adopt the submitted recording; Specifically, before submitting the audio stream to the dubbing management subsystem, the dubbing work subsystem submits the audio stream to the audio analysis subsystem; Specifically, the dubbing subsystem submits audio streams to the audio analysis subsystem on a task-by-task basis; Specifically, if a voice actor discovers an error in the script that prevents them from completing the performance, they can refuse to perform the task and submit the reason for refusal to the voice acting management subsystem. The audio analysis subsystem includes a signal analysis device, a content analysis device, and a quality detection device; Among them, the audio analysis subsystem analyzes and performs quality inspection on the audio stream, and feeds back the quality inspection results to the voice actors on the dubbing work subsystem side; The signal analysis device analyzes the audio stream and obtains the human voice attributes and environmental attributes of the audio stream; among which, the human voice attributes include human voice energy and human voice duration, and the environmental attributes include noise energy and spatial reverberation parameters; Specifically, the signal analysis device uses a human voice activation detection method to segment the audio stream into a human voice part and a non-human voice part, and calculates the energy and duration of the human voice part, while simultaneously calculating the energy of the non-human voice part; wherein, the non-human voice part includes noise; Among them, the signal analysis device calculates the energy using the root mean square of the signal, or calculates the energy using a loudness model based on a psychoacoustic model; Specifically, the signal analysis device uses machine learning regression algorithms to calculate spatial reverberation parameters in the audio stream; or, the signal analysis device uses synthesis to analyze the audio stream and uses generative models to iteratively approximate the spatial reverberation effect, thereby calculating the spatial reverberation parameters. Among them, spatial reverberation parameters include reverberation time; The content analysis device analyzes the audio stream and extracts the human voice content from the signal; The voice content includes spoken voice content described by written symbols and non-spoken voice content described by sound symbols. Non-voice human voice content includes breathing, laughter, crying, and vocalization sounds; Text symbols and sound symbols are accompanied by start and end timestamp information, and the duration of occupancy of text symbols or sound symbols can be obtained based on the timestamp information; Specifically, the content analysis device analyzes the audio stream and obtains human voice content through speech recognition; The quality inspection device performs quality inspection processing based on the human voice attributes, environmental attributes, and human voice content in the signal, and feeds back the quality inspection results to the dubbing subsystem. The quality inspection results include whether the quality inspection is passed, the quality inspection score, and the corresponding quality inspection error type, error content, and error correction method prompts. Among them, a rule-based method is used to determine whether the sound meets the quality inspection requirements; Alternatively, a machine learning-based method can be used to determine whether the sound meets the quality inspection requirements using a classification model. Alternatively, a machine learning-based approach can be used to obtain a sound quality score using a regression model. Voice actors receive and refer to the quality inspection results fed back by the quality inspection device through the voice acting work subsystem; If the voice actor believes that the dubbing performance does not need adjustment, the quality inspection results fed back by the recording and audio analysis subsystem will be directly submitted to the dubbing management subsystem; if the voice actor believes that the dubbing performance needs adjustment, the recording environment or performance method will be adjusted with reference to the quality inspection information and the dubbing performance will be performed again. The dubbing work subsystem will submit the audio stream generated by the adjusted recording to the audio analysis subsystem. The quality inspection and adjustment process will be repeated until the voice actor is satisfied. Then the quality inspection results fed back by the recording and audio analysis subsystem will be submitted to the dubbing management subsystem. The dubbing director reviews the recordings submitted by the voice actors through the dubbing management subsystem and, with reference to the feedback quality inspection results, decides whether to adopt the submitted recordings. When the dubbing director fails to approve the dubbing, the dubbing management subsystem will return the task and corresponding rework instructions to the dubbing work subsystem, and the corresponding dubbing actor will carry out the rework. Voice actors view rework tasks in the voice-over work subsystem, and re-enact and record their performances according to the rework guidelines. The voice-over work subsystem then resubmits the re-enacted recordings to the voice-over management subsystem and audio analysis subsystem, where the voice-over director reviews them and decides whether to adopt them.
[0020] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio acquisition method, characterized in that, include: The script editing subsystem receives the input text to be edited. The script editor edits the content of the text through the script editing subsystem to obtain the dubbing script, and sends the edited dubbing script to the dubbing management subsystem and the dubbing work subsystem. The dubbing management subsystem receives the dubbing scripts from the editors, and the dubbing director allocates and manages dubbing tasks according to the dubbing scripts in the dubbing management subsystem; When a voice actor discovers an error in the script that prevents them from completing the performance, they must refuse to perform the task and submit the reason for refusal to the voice acting management subsystem. The dubbing director reviews tasks rejected by voice actors through the dubbing management subsystem; if the rejection is reasonable, the dubbing director re-edits the dubbing script through the script editing subsystem and then reassigns the voice actors' tasks; if the rejection is unreasonable, the dubbing director directly reassigns the voice actors' tasks in the dubbing management subsystem. The dubbing subsystem receives the edited dubbing script and includes a recording device. Voice actors perform dubbing in the dubbing subsystem according to the dubbing script and record their voices using the recording device. The recording device converts the voice actors' dubbing performances into an audio stream. Before submitting the audio stream to the dubbing management subsystem, the dubbing work subsystem submits the audio stream to the audio analysis subsystem; The audio analysis subsystem analyzes and performs quality checks on the audio stream, and then feeds back the quality check results to the dubbing subsystem. Voice actors receive and refer to the quality inspection results fed back by the audio analysis subsystem through the voice acting work subsystem; The dubbing subsystem submits the audio stream and the quality inspection results fed back by the audio analysis subsystem to the dubbing management subsystem; The dubbing director reviews the audio recordings in the audio streams submitted by the voice actors through the dubbing management subsystem, and decides whether to adopt the submitted recordings based on the feedback quality inspection results. When the dubbing director fails to approve the dubbing, the dubbing management subsystem will return the task and corresponding rework instructions to the dubbing work subsystem, and the corresponding dubbing actor will carry out the rework. Voice actors view rework tasks in the voice-over work subsystem, and re-enact and record their performances according to the rework guidelines. The dubbing work subsystem will resubmit the re-dubbed recordings to the dubbing management subsystem, where the dubbing director will review them and decide whether to adopt them.
2. The audio acquisition method according to claim 1, Its features are, Among them, the script editor is either the screenwriter or the voice director; The organizational structure of the text to be edited, from largest to smallest, is: book, chapter, paragraph; A voice-over script includes a set of characters and a set of voice actors; within the set of characters, there is a one-to-many relationship between characters and segments; within the set of voice actors, there is a one-to-many relationship between voice actors and characters; and there is a one-to-many relationship between voice actors and segments. The script editor uses the script editing subsystem to configure the correspondence between characters and voice actors, and between paragraphs and characters in the dubbing script, and provides corresponding performance guidance in the dubbing script for the paragraphs that the voice actors need to perform; The organizational structure of a dubbing script, from largest to smallest, is as follows: book, chapter. For each chapter of the dubbing script, once the correspondence between the assigned paragraphs, characters, and voice actors is determined, the dubbing script content for that chapter can be edited. The dubbing script content for each chapter includes the content of each paragraph in that chapter, the character corresponding to the paragraph, the voice actor corresponding to the character, and performance guidance for the paragraph and character.
3. The audio acquisition method according to claim 1, characterized in that, in, A task consists of a character portrayed by a voice actor and a corresponding segment; a task includes task status and task information. Task status includes Incomplete, Rejected, Completed, and Pending Rework; Incomplete status means the voice actor has not completed the voice acting and recording; Rejected status means the voice actor has found an error in the script and refused to perform the voice acting; Completed status means the voice actor has completed the voice acting and recording; Pending Rework status means the voice actor has completed the voice acting and recording, but the submitted recording result has not passed the voice director's review. The task information includes the paragraph content, the performance guidance given by the dubbing director, and the rework guidance given by the dubbing director; A voice actor's portrayal of a character in a chapter corresponds to multiple segments. The collection of multiple segments corresponding to multiple tasks constitutes a batch. The voice acting script of a chapter consists of multiple batches. Each batch includes batch information, which includes the character attributes corresponding to the batch and the completion status of the tasks in the batch. The character attributes corresponding to a batch include character name, character type, character gender, character age, and character characteristic description; the completion status of tasks in a batch includes the number of tasks completed, the number of tasks rejected, the number of tasks awaiting rework, and the total number of tasks; The organizational structure of a dubbing task, from largest to smallest, is: book, batch, task.
4. The audio acquisition method according to claim 1, Its features are, Each voice actor has a corresponding book list configured in the voice acting work subsystem, which contains all the books in which the voice actor has voice acting tasks; The books for which voice actors participate in dubbing are configured with a batch list in the dubbing work subsystem; Voice actors can index batches and voice-over tasks in the batch list by chapter or character in the voice-over work subsystem; Voice actors sort the batches in the batch list by chapter or character. Voice actors select batches and execute tasks within those batches in the voice-over work subsystem; voice actors perform the corresponding segments by task and record their voices using recording devices. The dubbing subsystem submits audio streams to the audio analysis subsystem on a task-by-task basis.
5. The audio acquisition method according to claim 1, Its features are, in, If the voice actor believes that the dubbing performance does not need adjustment, the quality inspection results fed back by the recording and audio analysis subsystem will be directly submitted to the dubbing management subsystem; if the voice actor believes that the dubbing performance needs adjustment, the recording environment or performance method will be adjusted with reference to the quality inspection information and the dubbing performance will be performed again. The dubbing work subsystem will submit the audio stream generated by the adjusted recording to the audio analysis subsystem. The quality inspection and adjustment process will be repeated until the voice actor is satisfied. Then the quality inspection results fed back by the recording and audio analysis subsystem will be submitted to the dubbing management subsystem. The audio analysis subsystem includes a signal analysis device, a content analysis device, and a quality detection device. The signal analysis device analyzes the audio stream and obtains the human voice attributes and environmental attributes of the audio stream; among which, the human voice attributes include human voice energy and human voice duration, and the environmental attributes include noise energy and spatial reverberation parameters; The content analysis device analyzes the audio stream and obtains the human voice content in the signal; wherein, the human voice content includes speech human voice content described by text symbols and non-speech human voice content described by sound symbols; the text symbols and sound symbols are accompanied by start and end timestamp information, and the duration of occupancy of the text symbols or sound symbols is obtained based on the timestamp information; The quality inspection device performs quality inspection based on human voice attributes, environmental attributes, and human voice content, and feeds back the quality inspection results to the dubbing subsystem. The quality inspection results include whether the quality inspection is passed, the quality inspection score, and the corresponding quality inspection error type, error content, and error correction method prompts.
6. An audio acquisition system, characterized in that, It includes a script editing subsystem, a voice-over management subsystem, and a voice-over work subsystem; The script editing subsystem receives the input text to be edited. The script editor edits the content of the text through the script editing subsystem to obtain the dubbing script, and sends the edited dubbing script to the dubbing management subsystem and the dubbing work subsystem. The dubbing management subsystem receives the dubbing scripts edited by the dubbing director, who then allocates and manages dubbing tasks based on the scripts within the dubbing management subsystem. When a voice actor discovers an error in the script that prevents them from completing the performance, they must refuse to perform the task and submit the reason for refusal to the voice acting management subsystem. The dubbing director reviews tasks rejected by voice actors through the dubbing management subsystem; when the rejection is reasonable, the dubbing director re-edits the dubbing script through the script editing subsystem and then reassigns the voice actors' tasks; when the rejection is unreasonable, the dubbing director directly reassigns the voice actors' tasks in the dubbing management subsystem. The dubbing subsystem receives the edited dubbing script and includes a recording device. The dubbing subsystem performs dubbing according to the dubbing script and records the performance using the recording device. The recording device converts the voices of the dubbing actors into an audio stream. Before submitting the audio stream to the dubbing management subsystem, the dubbing work subsystem submits the audio stream to the audio analysis subsystem; the audio analysis subsystem analyzes and performs quality checks on the audio stream and feeds back the quality check results to the dubbing work subsystem; the dubbing actors receive and refer to the quality check results fed back by the audio analysis subsystem through the dubbing work subsystem; The dubbing work subsystem submits the audio stream and the quality inspection results fed back by the audio analysis subsystem to the dubbing management subsystem; the dubbing director reviews the recording content in the audio stream submitted by the dubbing actor through the dubbing management subsystem, and decides whether to adopt the submitted recording based on the feedback quality inspection results; when the dubbing director's review fails, the dubbing management subsystem returns the task and corresponding rework guidance to the dubbing work subsystem, and the corresponding dubbing actor carries out the rework. Voice actors view rework tasks in the voice-over work subsystem, re-enact and record their performances according to the rework guidelines, and then resubmit the re-enacted performances to the voice-over management subsystem for further review and a decision on whether to adopt them.
7. The audio acquisition system according to claim 6, Its features are, Among them, the script editor is either the screenwriter or the voice director; The organizational structure of the text to be edited, from largest to smallest, is: book, chapter, paragraph; A voice-over script includes a set of characters and a set of voice actors; within the set of characters, there is a one-to-many relationship between characters and segments; within the set of voice actors, there is a one-to-many relationship between voice actors and characters; and there is a one-to-many relationship between voice actors and segments. The script editor uses the script editing subsystem to configure the correspondence between characters and voice actors, and between paragraphs and characters in the dubbing script, and provides corresponding performance guidance in the dubbing script for the paragraphs that the voice actors need to perform; The organizational structure of a dubbing script, from largest to smallest, is as follows: book, chapter. For each chapter of the dubbing script, once the correspondence between the assigned paragraphs, characters, and voice actors is determined, the dubbing script content for that chapter can be edited. The dubbing script content for each chapter includes the content of each paragraph in that chapter, the character corresponding to the paragraph, the voice actor corresponding to the character, and performance guidance for the paragraph and character.
8. The audio acquisition system according to claim 6, characterized in that, in, A task consists of a character portrayed by a voice actor and a corresponding segment; a task includes task status and task information. Task status includes Incomplete, Rejected, Completed, and Pending Rework; Incomplete status means the voice actor has not completed the voice acting and recording; Rejected status means the voice actor has found an error in the script and refused to perform the voice acting; Completed status means the voice actor has completed the voice acting and recording; Pending Rework status means the voice actor has completed the voice acting and recording, but the submitted recording result has not passed the voice director's review. The task information includes the paragraph content, the performance guidance given by the dubbing director, and the rework guidance given by the dubbing director; A voice actor's portrayal of a character in a chapter corresponds to multiple segments. The collection of multiple segments corresponding to multiple tasks constitutes a batch. The voice acting script of a chapter consists of multiple batches. Each batch includes batch information, which includes the character attributes corresponding to the batch and the completion status of the tasks in the batch. The character attributes corresponding to a batch include character name, character type, character gender, character age, and character characteristic description; the completion status of tasks in a batch includes the number of tasks completed, the number of tasks rejected, the number of tasks awaiting rework, and the total number of tasks; The organizational structure of a dubbing task, from largest to smallest, is: book, batch, task.
9. The audio acquisition system according to claim 6, Its features are, in, Each voice actor has a corresponding book list configured in the voice acting work subsystem, which contains all the books in which the voice actor has voice acting assignments. The books for which voice actors participate in dubbing are configured with a batch list in the dubbing work subsystem; Voice actors can index batches and voice-over tasks in the batch list by chapter or character in the voice-over work subsystem; Voice actors sort the batches in the batch list by chapter or character. Voice actors select batches and execute tasks within those batches in the voice-over work subsystem; voice actors perform the corresponding segments by task and record their voices using recording devices. The dubbing subsystem submits audio streams to the audio analysis subsystem on a task-by-task basis.
10. The audio acquisition system according to claim 6, Its features are, in, If the voice actor believes that the dubbing performance does not need adjustment, the quality inspection results fed back by the recording and audio analysis subsystem will be directly submitted to the dubbing management subsystem; if the voice actor believes that the dubbing performance needs adjustment, the recording environment or performance method will be adjusted with reference to the quality inspection information and the dubbing performance will be performed again. The dubbing work subsystem will submit the audio stream generated by the adjusted recording to the audio analysis subsystem. The quality inspection and adjustment process will be repeated until the voice actor is satisfied. Then the quality inspection results fed back by the recording and audio analysis subsystem will be submitted to the dubbing management subsystem. The audio analysis subsystem includes a signal analysis device, a content analysis device, and a quality detection device. The signal analysis device analyzes the audio stream and obtains the human voice attributes and environmental attributes of the audio stream; among which, the human voice attributes include human voice energy and human voice duration, and the environmental attributes include noise energy and spatial reverberation parameters; The content analysis device analyzes the audio stream and obtains the human voice content in the signal; wherein, the human voice content includes speech human voice content described by text symbols and non-speech human voice content described by sound symbols; the text symbols and sound symbols are accompanied by start and end timestamp information, and the duration of occupancy of the text symbols or sound symbols is obtained based on the timestamp information; The quality inspection device performs quality inspection based on human voice attributes, environmental attributes, and human voice content, and feeds back the quality inspection results to the dubbing subsystem. The quality inspection results include whether the quality inspection is passed, the quality inspection score, and the corresponding quality inspection error type, error content, and error correction method prompts.
Citation Information
Patent Citations
Voice sample collection method based on network dubbing game
CN107293286A
Recording editing management method and system
CN111508468A
Office task management system and management method
CN112465468A
Audio presentation method of text
CN114579798A