A dubbing interaction method and device, computer equipment and a storage medium

By displaying information about multiple characters and generating aggregated voice-over audio in book reading applications, the lack of interactivity in existing technologies is solved, enabling multi-person voice-over and joint voice-over for characters, thus improving the user's voice-over experience.

CN117075839BActive Publication Date: 2026-04-14BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-15
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing user voiceover features lack interactivity in book reading applications, failing to create a good voiceover experience.

Method used

This paper provides a voice-over interaction method that displays information about multiple characters, obtains and associates voice-over audio generated by users and the system, generates aggregated voice-over audio, and displays audio identifiers. It supports multi-person voice-over and joint voice-over by characters.

Benefits of technology

It enables multi-person voice acting and joint voice acting for different characters, enriching the voice acting methods and enhancing the user's voice acting and reading experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117075839B_ABST
    Figure CN117075839B_ABST
Patent Text Reader

Abstract

The present disclosure provides a dubbing interaction method and device, computer equipment and a storage medium, wherein the method comprises: displaying a plurality of character information associated with a text to be dubbed; in response to a selection operation on the displayed first character information, obtaining a first dubbing audio of a first user, and associating the first dubbing audio with the first character information; the first dubbing audio is a dubbing performed on a text segment in the text to be dubbed that is associated with the first character information; based on the first dubbing audio associated with the first character information and a second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, an aggregated dubbing audio corresponding to the text to be dubbed is obtained, and a first audio identifier corresponding to the aggregated dubbing audio is displayed; the first audio identifier indicates the character information corresponding to each dubbing audio in the aggregated dubbing audio. The present embodiment realizes the dubbing interaction function and improves the dubbing and reading experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of information technology, and more specifically, to a voice-over interaction method, apparatus, computer device, and storage medium. Background Technology

[0002] In book reading apps, if users are interested in certain content, they may want to narrate it and share it. In this case, such apps can offer a user-generated voiceover feature, allowing users to add their own voiceovers to the content they are reading, thereby enhancing the user's interactive experience.

[0003] Typical user voiceover features only allow users to voice the characters they select, lacking interaction and failing to create a good voiceover experience. Summary of the Invention

[0004] This disclosure provides at least one voice-over interaction method, apparatus, computer device, and storage medium.

[0005] In a first aspect, embodiments of this disclosure provide a voice-over interaction method, including:

[0006] Displays information about multiple characters associated with the text to be dubbed;

[0007] In response to a selection operation on the first character information displayed, the system obtains the first voice-over audio of the first user and associates the first voice-over audio with the first character information; the first voice-over audio is a voice-over of a text segment in the text to be voiced that is associated with the first character information.

[0008] Based on the first dubbing audio associated with the first character information, and the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, an aggregated dubbing audio corresponding to the text to be dubbed is obtained, and a first audio identifier corresponding to the aggregated dubbing audio is displayed; the first audio identifier indicates the character information corresponding to each dubbing audio in the aggregated dubbing audio.

[0009] In one optional implementation, the first audio identifier includes audio playback identifiers for dubbing audio corresponding to multiple text segments; the order in which the audio playback identifiers are arranged in the first audio identifier is related to the contextual order of the text segments in the text to be dubbed.

[0010] In one optional implementation, the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed is obtained according to the following steps:

[0011] The voice-over dynamic information of the first user is published; the voice-over dynamic information includes the first voice-over audio associated with the text to be voiced;

[0012] Obtain the second voice-over audio corresponding to the second character information fed back by the second user based on the voice-over dynamic information.

[0013] In one optional implementation, the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed is obtained according to the following steps:

[0014] In response to an intelligent dubbing request, a second dubbing audio, generated in an artificial intelligence manner and corresponding to at least one second character information associated with the text to be dubbed, is obtained.

[0015] In one optional implementation, the voice-over audio corresponding to each of the multiple character information is obtained according to the following method:

[0016] In response to a dubbing trigger operation on the text to be dubbed, the chat interface of the target virtual room is displayed; the chat interface displays information about multiple characters associated with the text to be dubbed.

[0017] The system acquires the first voice-over audio of the first user in the chat interface for the first character information, and the second voice-over audio of the second user in the chat interface for the second character information.

[0018] In one optional implementation, after obtaining the first dubbing audio, the method further includes:

[0019] In response to obtaining the target dialect type selected by the first user, the first dubbing audio is converted into dubbing audio in the target dialect type; or...

[0020] Based on the authorized geographical location information of the first user, or the geographical location information contained in the text to be dubbed, a recommended dialect type is displayed; in response to a confirmation operation for the recommended dialect type, the first dubbing audio is converted into dubbing audio under the recommended dialect type.

[0021] In one optional implementation, the method further includes:

[0022] For at least one audio distribution scenario, based on the plot or character information associated with the target theme information in the audio distribution scenario, determine the aggregated dubbing audio associated with the target theme information;

[0023] Under the displayed target topic information, the first audio identifier corresponding to the aggregated dubbing audio is displayed in association; the first audio identifier is used to respond to the first trigger operation to display the text information and audio playback identifier corresponding to each dubbing audio of the aggregated dubbing audio, and the corresponding dubbing audio is played after any audio playback identifier is triggered.

[0024] In one optional implementation, the text to be dubbed is determined according to the following steps:

[0025] Display multiple fragment dimensions associated with the target book; the fragment dimensions are used to indicate text fragments in the target book that match preset attribute features;

[0026] In response to the target segment dimension selected by the first user, the text in the target book that matches the target segment dimension is used as the text to be dubbed.

[0027] In one optional implementation, the method further includes:

[0028] In response to a request to view the audio recording of a target book, obtain and display the aggregated audio recording information associated with the target book;

[0029] The voice-over aggregation information includes multiple segment dimensions, and a first audio identifier of the aggregated voice-over audio associated with each segment dimension; the segment dimensions are used to indicate text segments in the target book that match preset attribute features; the multiple segment dimensions include several of the following: popularity dimension, target character dimension, and target plot dimension;

[0030] The first audio identifier is used to respond to the second trigger operation by sequentially playing each dubbing audio of the aggregated dubbing audio under the segment dimension, or, after any audio playback identifier in the first audio identifier is triggered, playing the dubbing audio of the text segment corresponding to the audio playback identifier.

[0031] Secondly, embodiments of this disclosure also provide a voice-over interactive device, comprising:

[0032] The first display module is used to display information about multiple characters associated with the text to be dubbed;

[0033] The first acquisition module is configured to, in response to a selection operation on the displayed first role information, acquire the first voice-over audio of the first user and associate the first voice-over audio with the first role information; the first voice-over audio is a voice-over of a text segment in the text to be voiced that is associated with the first role information.

[0034] The second display module is used to obtain the aggregated dubbing audio corresponding to the text to be dubbed based on the first dubbing audio associated with the first character information and the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, and to display the first audio identifier corresponding to the aggregated dubbing audio; the first audio identifier indicates the character information corresponding to each dubbing audio in the aggregated dubbing audio.

[0035] Thirdly, embodiments of this disclosure also provide a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any optional implementation of the first aspect, are performed.

[0036] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the first aspect or any optional implementation thereof.

[0037] The voice-over interaction method provided in this embodiment can obtain aggregated voice-over audio based on the first voice-over audio of a first user dubbing a first character and the second voice-over audio of a second character dubbing a second character. The above-mentioned voice-over interaction method can not only realize multi-person voice-over, but also realize joint voice-over of multiple characters, realizing voice-over interaction function. While enriching the voice-over methods for users, it can also improve the user's voice-over and reading experience.

[0038] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0039] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0040] Figure 1 A flowchart of a voice-over interaction method provided by an embodiment of this disclosure is shown;

[0041] Figure 2 A schematic diagram illustrating a display of text to be dubbed, provided by an embodiment of this disclosure, is shown.

[0042] Figure 3 This illustration shows another schematic diagram of displaying text to be dubbed, provided by an embodiment of this disclosure;

[0043] Figure 4 A schematic diagram of the second user feedback second dubbing audio provided in an embodiment of this disclosure is shown;

[0044] Figure 5 This illustration shows a diagram of multiple users performing voiceovers in a chat interface of a target virtual room, as provided in an embodiment of this disclosure.

[0045] Figure 6 A schematic diagram illustrating aggregated dubbing audio provided in an embodiment of this disclosure is shown;

[0046] Figure 7 A schematic diagram illustrating the various dubbing audios under aggregated dubbing audio provided in an embodiment of this disclosure is shown;

[0047] Figure 8 A schematic diagram of the structure of a voice-over interactive device provided in an embodiment of this disclosure is shown;

[0048] Figure 9 A schematic diagram of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0050] In book reading apps, users can use the app's voice-over feature if they want to add voiceover to content they're interested in. Generally, these voice-over features only allow users to voice their chosen characters, lacking interactivity and failing to create a good voice-over and interactive experience.

[0051] Based on this, this disclosure provides a voice-over interaction method, including: displaying multiple character information associated with a text to be voiced; in response to a selection operation on the displayed first character information, obtaining a first voice-over audio of a first user, and associating the first voice-over audio with the first character information; the first voice-over audio is a voice-over of a text segment in the text to be voiced that is associated with the first character information; based on the first voice-over audio associated with the first character information, and the second voice-over audios corresponding to at least one second character information associated with the text to be voiced, obtaining an aggregated voice-over audio corresponding to the text to be voiced, and displaying a first audio identifier corresponding to the aggregated voice-over audio; the first audio identifier indicates the character information corresponding to each voice-over audio in the aggregated voice-over audio.

[0052] The voice-over interaction method provided in this embodiment can obtain aggregated voice-over audio based on a first voice-over audio of a first user dubbing a first character and a second voice-over audio of a second character. The above-mentioned voice-over interaction method can not only realize multi-person voice-over, but also realize joint voice-over of multiple characters, realizing voice-over interaction function. While enriching the voice-over methods for users, it can also improve the user's voice-over and reading experience.

[0053] The deficiencies of the above solutions and the proposed solutions are the result of the inventor's practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below should be considered as the inventor's contribution to this disclosure.

[0054] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0055] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0056] To facilitate understanding of this embodiment, a detailed description of a voice-over interaction method disclosed in this disclosure embodiment will be provided first. The execution subject of the voice-over interaction method provided in this disclosure embodiment is generally a computer device with a certain computing power.

[0057] The voice-over interaction method provided in the embodiments of this disclosure will be described below.

[0058] See Figure 1 The diagram shows a flowchart of a voice-over interaction method provided in an embodiment of this disclosure. The method includes steps S101 to S103, wherein:

[0059] S101: Displays information about multiple characters associated with the text to be dubbed.

[0060] S102: In response to a selection operation on the first character information to be displayed, obtain the first voice-over audio of the first user and associate the first voice-over audio with the first character information; the first voice-over audio is a voice-over of a text segment in the text to be voiced that is associated with the first character information.

[0061] S103: Based on the first dubbing audio associated with the first character information and the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, obtain the aggregated dubbing audio corresponding to the text to be dubbed, and display the first audio identifier corresponding to the aggregated dubbing audio; the first audio identifier indicates the character information corresponding to each dubbing audio in the aggregated dubbing audio.

[0062] In this embodiment of the disclosure, the text to be dubbed may include the complete text content to be dubbed, such as an entire book or a script, or it may include a portion of the complete text content, such as a paragraph or chapter in a book or script. When the text to be dubbed includes a portion of the complete text content, the text to be dubbed can be any part of the text content, or it can be the text content at the target segment level.

[0063] In one implementation, the text to be dubbed can be determined according to the following steps:

[0064] Display multiple segment dimensions associated with the target book; the segment dimensions are used to indicate text segments in the target book that match preset attribute features; in response to the target segment dimension selected by the first user, the text in the target book that matches the target segment dimension is used as the text to be dubbed.

[0065] In the above embodiments, the segment dimension can include a popularity dimension, a target character dimension, a target plot dimension, a target chapter dimension, etc. The popularity dimension can be a segment whose popularity value exceeds a first preset threshold. The target character dimension can be a dimension of a target character in the target book. A target character can be any character in the target book, or a character whose popularity exceeds a second preset threshold. The target plot dimension can be a dimension of a target plot in the target book. A target plot can be any plot in the target book, or a plot whose popularity exceeds a third preset threshold. The target chapter dimension can be a dimension of a target chapter in the target book. A target chapter can be any chapter in the target book, or a chapter whose popularity exceeds a fourth preset threshold.

[0066] like Figure 2As shown, multiple segment dimensions can be displayed on the book's introduction page, such as popular characters, popular plot points, and popular chapters. Responding to the first user's selected target segment dimension (e.g., popular characters), the book's content page can be displayed, showing text segments under that target segment dimension. For example... Figure 3 As shown, multiple segment dimensions can also be displayed on the book content page. Responding to the first user's selected target segment dimension (such as a trending character dimension), text segments under that target segment dimension are displayed on the book content page. These text segments under the target segment dimension can be used as the text to be dubbed.

[0067] The text to be dubbed can be associated with multiple characters, and each character can have corresponding text fragments, such as the monologue of each character, the dialogue between each character and other characters, etc.

[0068] In one implementation, in response to a dubbing trigger operation on the text to be dubbed, multiple character information associated with the text to be dubbed can be displayed.

[0069] Dubbing triggering operations can include dubbing triggering operations for the text to be dubbed displayed on the content display page while browsing the text to be dubbed; or dubbing triggering operations for the text to be dubbed displayed on the content discussion page while participating in a discussion about the text to be dubbed.

[0070] Different role information can refer to different roles. The displayed role information may include, for example, the role's name, identifier, and avatar.

[0071] In this embodiment of the disclosure, the first dubbing audio may be a dubbing of a text segment associated with the first character information after triggering a selection operation for the first character information, or it may be a dubbing of a text segment associated with the first character information that the first user has already dubbed in advance.

[0072] As mentioned earlier, the text to be dubbed can be associated with multiple character information. These multiple character information may include first character information and at least one second character information. In this embodiment of the disclosure, second dubbing audio corresponding to at least one second character information associated with the text to be dubbed can also be obtained. The character indicated by the second character information may be different from the character indicated by the first character information.

[0073] The second voice-over audio can be a voice-over of a text segment associated with information about a second character. This second voice-over audio can be user-generated or generated using artificial intelligence. The following sections describe these two voice-over methods for the second voice-over audio.

[0074] Regarding the method of providing a second voice-over audio for a user, in one implementation, the second voice-over audio can be obtained according to the following steps: publishing the voice-over dynamic information of a first user; the voice-over dynamic information includes a first voice-over audio associated with the text to be voiced; and obtaining the second voice-over audio corresponding to the second character information fed back by the second user based on the voice-over dynamic information.

[0075] Here, "dubbing dynamic information" can refer to the posting of a voice-over, indicating that a first user has dubbed a text segment associated with a first character. This dubbing dynamic information can be used by a second user to provide feedback on a text segment associated with a second character. For example, after a first user posts a voice-over containing the first dubbing audio, a second user can provide feedback with a second dubbing audio in the form of a comment.

[0076] like Figure 4 As shown, the first user's activity display page shows the voice-over activity information posted by the first user, namely, the voice-over activity information posted by user "Nickname 1". The voice-over activity information includes the text to be voiced and the voice-over audio of character 1 in the text to be voiced by user "Nickname 1". The second user can provide feedback in the comments based on the voice-over activity information, providing a second voice-over audio for character 2. After the second user, for example, the user "Nickname 4", triggers the voice-over button on the activity display page, they can post their own voice-over audio for character 2. The voice-over audio posted by the second user can be displayed in the comment area of ​​the activity display page. Figure 4 After each user publishes their voice-over audio for their chosen character, the voice-over audio for each character can be arranged in chronological order of publication. For example, the voice-over audio published by user "Nickname 2" for character 3, the voice-over audio published by user "Nickname 3" for character 4, and the voice-over audio published by user "Nickname 4" for character 2 can be arranged from top to bottom in chronological order of publication.

[0077] At least one second character's information corresponds to a second voice-over audio, which can be performed by the same second user or by different second users. When performing for different second characters, the resulting second voice-over audio can match the vocal characteristics of each second character.

[0078] In one implementation of a method where the second voice-over audio is generated using artificial intelligence, the second voice-over audio may be obtained by following the steps of: in response to an intelligent voice-over request, obtaining second voice-over audio generated using artificial intelligence and corresponding to at least one second character information associated with the text to be voiced.

[0079] Here, pre-trained models or Text-to-Speech (TTS) technology can be used to generate second dubbing audio corresponding to at least one second character information associated with the text to be dubbed. For different second character information, the second dubbing audio generated using artificial intelligence can also conform to the voice characteristics of that character. For example, for the role of a little girl, the generated second dubbing audio could be a high-pitched, pure female voice; for the role of an elderly woman, the generated second dubbing audio could be a deep, soft female voice.

[0080] The intelligent dubbing request may include a first intelligent dubbing request issued separately for each piece of second character information, or a second intelligent request issued for at least one piece of second character information. Specifically, in one approach, the first intelligent dubbing request issued separately for each piece of second character information may acquire, respectively, a second dubbing audio generated based on artificial intelligence and corresponding to each piece of second character information. In another approach, the second intelligent request issued for at least one piece of second character information may simultaneously acquire, based on artificial intelligence, a second dubbing audio corresponding to at least one piece of second character information associated with the text to be dubbed.

[0081] In the aforementioned embodiments, the first and second voice-over audio can be generated independently. For example, the first and second voice-over audio can be obtained in different voice-over scenarios or at different time periods. Another embodiment can be provided below, in which the first and second voice-over audio can be obtained from a group chat voice-over performed by a first user and a second user.

[0082] Specifically, the dubbing audio corresponding to multiple character information can be obtained in the following way: in response to the dubbing trigger operation for the text to be dubbed, the chat interface of the target virtual room is displayed; multiple character information associated with the text to be dubbed is displayed in the chat interface; the first dubbing audio of the first user for the first character information in the chat interface, and the second dubbing audio of the second user for the second character information in the chat interface are obtained.

[0083] In the above embodiments, the dubbing trigger operation for the text to be dubbed may include creating a virtual room or entering a virtual room. For example, controls for creating a virtual room and / or entering a virtual room may be displayed on the display page or discussion page of the text to be dubbed. After a user triggers the control to create a virtual room, they can create a new virtual room; or, after triggering the control to enter a virtual room, they can enter a virtual room created by another user.

[0084] In cases where the voice-over triggering operation includes the creation of a virtual room, in one approach, in response to the operation of creating a virtual room for the text to be voiced, a virtual room creation prompt message can be displayed; in response to receiving confirmation feedback from the user who created the virtual room, and other users selected by the user who created the virtual room, the chat interface of the created target virtual room can be displayed.

[0085] In cases where voice-over triggering operations include entering a virtual room, in one approach, in response to an operation to enter a virtual room for the text to be voiced, the chat interface of the target virtual room can be displayed after receiving acceptance information from the user who created the virtual room.

[0086] The chat interface of the target virtual room can display information about multiple characters associated with the text to be dubbed. Users entering the target virtual room can dub for at least one selected character in the chat interface.

[0087] Figure 5 This diagram illustrates multiple users providing voiceovers within a target virtual room's chat interface. After multiple users enter the same target virtual room, their individual information (such as avatars and names) is displayed in the chat interface, with each user corresponding to at least one role. The chat page displays the voiceover content for each role. Each user can then provide voiceovers for their chosen role in a chat dialogue format.

[0088] Based on the voiceovers of each user in the chat interface, the first voiceover audio of the first user for the first character information and the second voiceover audio of the second user for the second character information can be obtained.

[0089] To enrich the dubbing effect, in this embodiment of the disclosure, the obtained first dubbing audio or second dubbing audio can also be converted into dubbing audio in the target dialect type, making the language types of dubbing more diverse and interesting.

[0090] The following describes the dialect type conversion method for dubbing audio, taking the first dubbing audio as an example. In one implementation, after obtaining the first dubbing audio, in response to obtaining the target dialect type selected by the first user, the first dubbing audio is converted into a dubbing type under the target dialect type; or, based on the authorized geographical location information of the first user, or the geographical location information contained in the text to be dubbed, a recommended dialect type is displayed; in response to a confirmation operation for the recommended dialect type, the first dubbing audio is converted into a dubbing audio under the recommended dialect type.

[0091] In the above implementation, multiple dialect types can be displayed for the first user to choose from. After the first user selects a target dialect type from the multiple dialect types, the first dubbing audio can be converted into a dubbing type under the target dialect type.

[0092] If the text to be dubbed contains geographical location information, it can first be determined whether the geographical location information is real. If so, the dialect type matching the real geographical location information can be displayed. If the real geographical location information contains multiple dialect types, it can either display all the dialect types included in the real geographical location information for the first user to choose from, or display the dialect type whose number of users ranks higher than a preset ranking among the multiple dialect types included in the real geographical location information for the first user to choose from.

[0093] The second voice-over audio can also be converted into voice-over audio in the target dialect type according to the above implementation method. When the second voice-over audio is performed by a user, the second user can select the target dialect type and convert the second voice-over audio into voice-over audio in the target dialect type, or another user (e.g., the first user) can select the target dialect type and convert the second voice-over audio into voice-over audio in the target dialect type. When the second voice-over audio is performed using artificial intelligence, a user (e.g., the first user) can select the target dialect type and convert the second voice-over audio into voice-over audio in the target dialect type.

[0094] In this embodiment of the disclosure, for example, the aggregated dubbing audio can be obtained by aggregating the dubbing audio corresponding to multiple character information according to the contextual order of the text segments corresponding to each character information. Another example is that the aggregated dubbing audio can also be obtained by sorting the dubbing audio corresponding to each character information according to a preset arrangement order (such as appearance order, popularity ranking, etc.).

[0095] In this embodiment of the disclosure, the first audio identifier of the aggregated dubbing audio can be displayed at an associated position in the text to be dubbed, for example, it can be displayed at the end of the text to be dubbed.

[0096] In one implementation, the first audio identifier of the aggregated dubbing audio may include audio playback identifiers for the dubbing audio corresponding to multiple text segments. The order of the audio playback identifiers in the first audio identifier may be related to the contextual order of each text segment in the text to be dubbed.

[0097] The audio playback identifier can also include information such as the corresponding character name, audio identifier number, and play button.

[0098] The audio identifier number can be determined according to the contextual order of each text segment within the text to be dubbed. The play button can be used to play the dubbing audio corresponding to that audio playback identifier upon triggering the play.

[0099] After obtaining the aggregated dubbing audio, it can be distributed. In one implementation, for at least one audio distribution scenario, the aggregated dubbing audio associated with the target theme information can be determined based on the plot or character information associated with the target theme information in the audio distribution scenario; under the displayed target theme information, a first audio identifier corresponding to the aggregated dubbing audio is displayed; the first audio identifier is used to respond to a first trigger operation to display the text information and audio playback identifier corresponding to each dubbing audio in the aggregated dubbing audio, and the corresponding dubbing audio is played after any audio playback identifier is triggered.

[0100] In the above implementation, the aggregated audio recordings can be distributed to either the content display page or the content discussion page.

[0101] Target topic information can include topics discussed by users, comments posted, book chapters, and other topic information. Topics discussed by users, comments posted, and book chapters browsed by users often involve plot, characters, and other information from the target book or script. Therefore, in audio distribution scenarios, aggregated dubbing audio related to the target topic information can be distributed based on the plot or character information associated with the topics discussed by users, comments, book chapters, etc.

[0102] Aggregated voice-over audio can be obtained by combining the voice-over audio corresponding to each character's information. After triggering the first audio identifier corresponding to the aggregated voice-over audio, the text information and audio playback identifier corresponding to each voice-over audio in the aggregated voice-over audio can be displayed.

[0103] Figure 6 This displays aggregated audio recordings associated with multiple book chapters under the target book. These aggregated audio recordings can be displayed on the target book's content display page. Each book chapter displays the text content of that chapter, as well as the first audio identifier corresponding to the aggregated audio recording. After a user triggers the first audio identifier, the text information and audio playback identifier corresponding to each audio recording will be displayed on the audio recording display page of that book chapter, such as... Figure 7 As shown, the corresponding dubbing audio will be played when any audio playback flag is triggered.

[0104] By displaying the text information corresponding to each audio recording, users can understand which part of the text each audio recording is for. An audio playback indicator can be used to play the corresponding audio recording when triggered.

[0105] In this embodiment of the disclosure, viewing of voice-over aggregation information can also be supported. In one implementation, in response to a request to view the voice-over for a target book, voice-over aggregation information associated with the target book can be obtained and displayed.

[0106] Here, the voice-over aggregation information can include multiple segment dimensions, as well as the first audio identifier of the aggregated voice-over audio associated with each segment dimension; the segment dimension is used to indicate the text segment in the target book that matches the preset attribute features; the multiple segment dimensions include various ones such as popularity dimension, target character dimension, and target plot dimension.

[0107] In response to the second triggering operation, the first audio identifier sequentially plays each of the aggregated dubbing audios under the segment dimension, or, after any audio playback identifier in the first audio identifier is triggered, the dubbing audio of the text segment corresponding to the audio playback identifier is played.

[0108] Multiple fragment dimensions can be preset. A description of multiple fragment dimensions can be found above and will not be repeated here.

[0109] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0110] Based on the same inventive concept, this disclosure also provides a dubbing interaction device corresponding to the dubbing interaction method. Since the principle of the device in this disclosure for solving the problem is similar to the dubbing interaction method described above in this disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0111] Reference Figure 8 The diagram shown is an architectural schematic of a voice-over interactive device provided in an embodiment of this disclosure. The device includes:

[0112] The first display module 801 is used to display information about multiple characters associated with the text to be dubbed;

[0113] The first acquisition module 802 is used to acquire the first voice-over audio of the first user in response to a selection operation on the displayed first role information, and associate the first voice-over audio with the first role information; the first voice-over audio is a voice-over of a text segment in the text to be voiced that is associated with the first role information.

[0114] The second display module 803 is used to obtain the aggregated dubbing audio corresponding to the text to be dubbed based on the first dubbing audio associated with the first character information and the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, and to display the first audio identifier corresponding to the aggregated dubbing audio; the first audio identifier indicates the character information corresponding to each dubbing audio in the aggregated dubbing audio.

[0115] In one optional implementation, the first audio identifier includes audio playback identifiers for dubbing audio corresponding to multiple text segments; the order in which the audio playback identifiers are arranged in the first audio identifier is related to the contextual order of the text segments in the text to be dubbed.

[0116] In an optional embodiment, the apparatus further includes a second acquisition module for acquiring second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, the second acquisition module being specifically used for:

[0117] The voice-over dynamic information of the first user is published; the voice-over dynamic information includes the first voice-over audio associated with the text to be voiced;

[0118] Obtain the second voice-over audio corresponding to the second character information fed back by the second user based on the voice-over dynamic information.

[0119] In an optional embodiment, the apparatus further includes a third acquisition module for acquiring second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, the third acquisition module being specifically used for:

[0120] In response to an intelligent dubbing request, a second dubbing audio, generated in an artificial intelligence manner and corresponding to at least one second character information associated with the text to be dubbed, is obtained.

[0121] In an optional embodiment, the device further includes a fourth acquisition module for acquiring the dubbing audio corresponding to the plurality of character information, the fourth acquisition module being specifically used for:

[0122] In response to a dubbing trigger operation on the text to be dubbed, the chat interface of the target virtual room is displayed; the chat interface displays information about multiple characters associated with the text to be dubbed.

[0123] The system acquires the first voice-over audio of the first user in the chat interface for the first character information, and the second voice-over audio of the second user in the chat interface for the second character information.

[0124] In one optional embodiment, the apparatus further includes:

[0125] The conversion module is configured to, in response to obtaining the target dialect type selected by the first user, convert the first dubbing audio into dubbing audio in the target dialect type; or,

[0126] Based on the authorized geographical location information of the first user, or the geographical location information contained in the text to be dubbed, a recommended dialect type is displayed; in response to a confirmation operation for the recommended dialect type, the first dubbing audio is converted into dubbing audio under the recommended dialect type.

[0127] In one optional embodiment, the apparatus further includes:

[0128] The first determining module is used to determine, for at least one audio distribution scenario, the aggregated dubbing audio associated with the target theme information based on the plot or character information associated with the target theme information in the audio distribution scenario;

[0129] The third display module is used to display the first audio identifier corresponding to the aggregated dubbing audio under the displayed target theme information; the first audio identifier is used to respond to the first trigger operation to display the text information and audio playback identifier corresponding to each dubbing audio of the aggregated dubbing audio, and play the corresponding dubbing audio after any audio playback identifier is triggered.

[0130] In an optional embodiment, the apparatus further includes a second determining module for determining the text to be dubbed, the second determining module being specifically used for:

[0131] Display multiple fragment dimensions associated with the target book; the fragment dimensions are used to indicate text fragments in the target book that match preset attribute features;

[0132] In response to the target segment dimension selected by the first user, the text in the target book that matches the target segment dimension is used as the text to be dubbed.

[0133] In one optional embodiment, the apparatus further includes:

[0134] The fifth acquisition module is used to acquire and display the voice-over aggregation information associated with the target book in response to a request to view the voice-over for the target book.

[0135] The voice-over aggregation information includes multiple segment dimensions, and a first audio identifier of the aggregated voice-over audio associated with each segment dimension; the segment dimensions are used to indicate text segments in the target book that match preset attribute features; the multiple segment dimensions include several of the following: popularity dimension, target character dimension, and target plot dimension;

[0136] The first audio identifier is used to respond to the second trigger operation by sequentially playing each dubbing audio of the aggregated dubbing audio under the segment dimension, or, after any audio playback identifier in the first audio identifier is triggered, playing the dubbing audio of the text segment corresponding to the audio playback identifier.

[0137] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0138] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 9 The diagram shows the structure of a computer device 900 provided in this embodiment, including a processor 901, a memory 902, and a bus 903. The memory 902 stores execution instructions and includes main memory 9021 and external memory 9022. The main memory 9021, also called internal memory, temporarily stores computational data in the processor 901 and data exchanged with external memory 9022 such as a hard disk. The processor 901 exchanges data with the external memory 9022 through the main memory 9021. When the computer device 900 is running, the processor 901 and the memory 902 communicate through the bus 903, causing the processor 901 to execute the following instructions:

[0139] Displays information about multiple characters associated with the text to be dubbed;

[0140] In response to a selection operation on the first character information displayed, the system obtains the first voice-over audio of the first user and associates the first voice-over audio with the first character information; the first voice-over audio is a voice-over of a text segment in the text to be voiced that is associated with the first character information.

[0141] Based on the first dubbing audio associated with the first character information, and the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, an aggregated dubbing audio corresponding to the text to be dubbed is obtained, and a first audio identifier corresponding to the aggregated dubbing audio is displayed; the first audio identifier indicates the character information corresponding to each dubbing audio in the aggregated dubbing audio.

[0142] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the voice-over interaction method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0143] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the voice-over interaction method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0144] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0145] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed device and method can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0147] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0148] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0149] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A voice-over interaction method, characterized in that, include: Display information about multiple characters associated with the text to be dubbed, wherein the multiple characters have corresponding text segments in the text to be dubbed, and the multiple characters include a first character and at least one second character; In response to a selection operation on the first character information displayed, the system obtains the first voice-over audio of the first user and associates the first voice-over audio with the first character information; the first voice-over audio is a voice-over of a text segment in the text to be voiced that is associated with the first character information. Based on the first dubbing audio associated with the first character information, and the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, an aggregated dubbing audio corresponding to the text to be dubbed is obtained, and a first audio identifier corresponding to the aggregated dubbing audio is displayed; the first audio identifier indicates the character information corresponding to each dubbing audio in the aggregated dubbing audio, and the second dubbing audio is a dubbing of a text segment in the text to be dubbed that is associated with the second character information. The text to be dubbed is obtained according to the following steps: Display multiple fragment dimensions associated with the target book; the fragment dimensions are used to indicate text fragments in the target book that match preset attribute features, and the multiple fragment dimensions include various ones such as popularity dimension, target character dimension, and target plot dimension; In response to the target segment dimension selected by the first user, the text in the target book that matches the target segment dimension is used as the text to be dubbed.

2. The method according to claim 1, wherein the first audio identifier includes audio playback identifiers for dubbing audio corresponding to multiple text segments; the order of the audio playback identifiers in the first audio identifier is related to the contextual order of the text segments in the text to be dubbed.

3. The method according to claim 1, characterized in that, The following steps are used to obtain the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed: The voice-over dynamic information of the first user is published; the voice-over dynamic information includes the first voice-over audio associated with the text to be voiced; Obtain the second voice-over audio corresponding to the second character information fed back by the second user based on the voice-over dynamic information.

4. The method according to claim 1, characterized in that, The following steps are used to obtain the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed: In response to an intelligent dubbing request, a second dubbing audio, generated in an artificial intelligence manner and corresponding to at least one second character information associated with the text to be dubbed, is obtained.

5. The method according to claim 1, characterized in that, The voice-over audio corresponding to each of the multiple character information is obtained using the following method: In response to a dubbing trigger operation on the text to be dubbed, the chat interface of the target virtual room is displayed; the chat interface displays information about multiple characters associated with the text to be dubbed. The system acquires the first voice-over audio of the first user in the chat interface for the first character information, and the second voice-over audio of the second user in the chat interface for the second character information.

6. The method according to claim 1, characterized in that, After obtaining the first dubbing audio, the process also includes: In response to obtaining the target dialect type selected by the first user, the first dubbing audio is converted into dubbing audio in the target dialect type; or... Based on the authorized geographical location information of the first user, or the geographical location information contained in the text to be dubbed, a recommended dialect type is displayed; in response to a confirmation operation for the recommended dialect type, the first dubbing audio is converted into dubbing audio under the recommended dialect type.

7. The method according to claim 1, characterized in that, The method further includes: For at least one audio distribution scenario, based on the plot or character information associated with the target theme information in the audio distribution scenario, determine the aggregated dubbing audio associated with the target theme information; Under the displayed target topic information, the first audio identifier corresponding to the aggregated dubbing audio is displayed in association; the first audio identifier is used to respond to the first trigger operation to display the text information and audio playback identifier corresponding to each dubbing audio of the aggregated dubbing audio, and the corresponding dubbing audio is played after any audio playback identifier is triggered.

8. The method according to claim 1, characterized in that, The method further includes: In response to a request to view the audio recording of a target book, obtain and display the aggregated audio recording information associated with the target book; The voice-over aggregation information includes multiple segment dimensions, and a first audio identifier of the aggregated voice-over audio associated with each segment dimension; The first audio identifier is used to respond to the second trigger operation by sequentially playing each dubbing audio of the aggregated dubbing audio under the segment dimension, or, after any audio playback identifier in the first audio identifier is triggered, playing the dubbing audio of the text segment corresponding to the audio playback identifier.

9. A voice-over interactive device, characterized in that, include: The first display module is used to display information about multiple characters associated with the text to be dubbed, wherein the multiple characters have corresponding text fragments in the text to be dubbed, and the multiple characters include a first character and at least one second character; The first acquisition module is configured to, in response to a selection operation on the displayed first role information, acquire the first voice-over audio of the first user and associate the first voice-over audio with the first role information; the first voice-over audio is a voice-over of a text segment in the text to be voiced that is associated with the first role information. The second display module is used to obtain aggregated dubbing audio corresponding to the text to be dubbed based on the first dubbing audio associated with the first character information and the second dubbing audio corresponding to at least one second character information associated with the text to be dubbed, and to display a first audio identifier corresponding to the aggregated dubbing audio; the first audio identifier indicates the character information corresponding to each dubbing audio in the aggregated dubbing audio, and the second dubbing audio is a dubbing of a text segment in the text to be dubbed that is associated with the second character information. The text to be dubbed is obtained according to the following steps: Display multiple fragment dimensions associated with the target book; the fragment dimensions are used to indicate text fragments in the target book that match preset attribute features, and the multiple fragment dimensions include various ones such as popularity dimension, target character dimension, and target plot dimension; In response to the target segment dimension selected by the first user, the text in the target book that matches the target segment dimension is used as the text to be dubbed.

10. A computer device, characterized in that, include: The computer device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and the machine-readable instructions, when executed by the processor, perform the steps of the voice-over interaction method as described in any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the voice-over interaction method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Media information processing method and device

    CN107659850A

  • Audio data merging method, audio data merging device, storage medium and processor

    CN107809666A

  • Social interaction method, device and system, equipment and storage medium

    CN112261435A

  • Audio book generation method and device, equipment, storage medium and program product

    CN114783403A

  • Video dubbing method, related equipment and computer readable storage medium

    CN115037975A