Dubbing method and device, equipment, storage medium and program product

By using intelligent sentence splitting and dubbing technology on terminal devices, multiple dubbing files are generated, solving the problem of low efficiency for multi-person dubbing and realizing efficient and flexible multimedia work generation.

CN120956969APending Publication Date: 2025-11-14BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510954658.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

When creating multimedia works with multiple voice actors, existing technologies require multiple text selection and voice-over operations, resulting in low generation efficiency and a lack of intelligent sentence segmentation and character recognition.

Method used

The terminal device uses intelligent sentence segmentation and dubbing technology to generate multiple dubbing files based on the text to be dubbed. It supports multiple dubbing objects, provides a dubbing file interface and generation controls, and allows users to modify and adjust the dubbing files.

Benefits of technology

It improves the efficiency and accuracy of multimedia production, enhances user flexibility and satisfaction with dubbing files, and simplifies the process of multi-person dubbing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956969A_ABST
    Figure CN120956969A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a dubbing method and device, equipment, a storage medium and a program product, and is used for enabling terminal equipment to intelligently perform sentence splitting and dubbing according to a to-be-dubbed text, generating a plurality of dubbing files and improving the efficiency of generating dubbed multimedia works. The method comprises the steps that in response to dubbing operation on the multimedia work, a dubbing text interface is displayed, and the dubbing text interface comprises a to-be-dubbed text used for dubbing the multimedia work and a first generation control; in response to a trigger operation on the first generation control, a dubbing file interface is displayed, the dubbing file interface comprises a plurality of dubbing files and a second generation control, the plurality of dubbing files are dubbing files generated by performing sentence splitting and dubbing on the to-be-dubbed text, and dubbing objects corresponding to the plurality of dubbing files comprise at least two dubbing objects; and in response to a trigger operation on the second generation control, generating a dubbed multimedia work according to the plurality of dubbing files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimedia, and more particularly to a dubbing method, apparatus, device, storage medium, and program product. Background Technology

[0002] When adding voiceovers to existing multimedia works, multiple text selection and voiceover operations are required to obtain the multimedia work with multiple voiceovers. This process is complex and inefficient. Summary of the Invention

[0003] This application provides a dubbing method, apparatus, device, storage medium, and program product, which can be used by terminal devices to intelligently split and dub the text to be dubbed, generate multiple dubbing files, and improve the efficiency of generating dubbed multimedia works.

[0004] The first aspect of this application provides a dubbing method, which may include: in response to an operation of dubbing a multimedia work, displaying a dubbing text interface, the dubbing text interface including text to be dubbed for dubbing the multimedia work and a first generation control; in response to a trigger operation of the first generation control, displaying a dubbing file interface, the dubbing file interface including multiple dubbing files and a second generation control, the multiple dubbing files being dubbing files generated by splitting and dubbing the text to be dubbed, and the dubbing objects corresponding to the multiple dubbing files including at least two dubbing objects; in response to a trigger operation of the second generation control, generating a dubbed multimedia work based on the multiple dubbing files.

[0005] In this technical solution, the terminal device can intelligently split sentences and dub the text to be dubbed, generating multiple dubbing files corresponding to multiple dubbing objects, thereby improving the efficiency of generating dubbed multimedia works.

[0006] In some possible implementations of this application, the step of displaying the dubbing file interface in response to a trigger operation on the first generation control may include: in response to a trigger operation on the first generation control, determining the plurality of dubbing files based on dubbing information, and displaying the dubbing file interface, wherein the dubbing information includes the text to be dubbed, or the dubbing information includes the text to be dubbed and the image objects in the multimedia work.

[0007] This technical solution provides a specific implementation method for a terminal device to display a dubbing file interface in response to a trigger operation of the first generation control. It provides a method to generate multiple dubbing files based on the text to be dubbed, or based on the text to be dubbed and the screen objects in the multimedia work, and then display them on the dubbing file interface, thereby improving the feasibility of the solution.

[0008] In some possible implementations of this application, the dubbing information includes the text to be dubbed, which includes the text to be dubbed corresponding to at least two character tags; the step of determining the plurality of dubbing files based on the dubbing information in response to the triggering operation of the first generation control may include: in response to the triggering operation of the first generation control, obtaining at least two sub-texts to be dubbed based on the text to be dubbed corresponding to the at least two character tags, each character tag corresponding to at least one sub-text to be dubbed, the at least one sub-text to be dubbed corresponding to each character tag being obtained based on the text to be dubbed corresponding to each character tag; determining at least two dubbing objects corresponding to the at least two character tags, each character tag corresponding to one dubbing object; dubbing the at least two sub-texts to be dubbed using the at least two dubbing objects to obtain the plurality of dubbing files, each dubbing file being obtained by dubbing one sub-text to be dubbed corresponding to each character tag using the dubbing object corresponding to each character tag.

[0009] In this technical solution, when the dubbing information includes the text to be dubbed, and the text to be dubbed includes the text to be dubbed corresponding to at least two character tags, the terminal device responds to the trigger operation of the first generation control and determines the multiple dubbing files according to the dubbing information. Specifically, based on the text to be dubbed corresponding to at least two character tags, at least two sub-texts to be dubbed and at least two dubbing objects can be obtained. Then, dubbing is performed on the at least two sub-texts to be dubbed using the at least two dubbing objects, resulting in multiple dubbing files displayed on the dubbing file interface. This improves the accuracy of generating multiple dubbing files, thereby enhancing the feasibility of the solution.

[0010] In some possible implementations of this application, the dubbing information includes the text to be dubbed and the screen objects in the multimedia work, wherein the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags; the step of determining the plurality of dubbing files based on the dubbing information in response to the triggering operation of the first generation control may include: in response to the triggering operation of the first generation control, obtaining at least two sub-texts to be dubbed based on the screen objects in the multimedia work and the text to be dubbed; determining the at least two dubbing objects based on the screen objects in the multimedia work and the text to be dubbed, wherein the number of the at least two dubbing objects is less than or equal to the number of screen objects in the multimedia work; dubbing the at least two sub-texts to be dubbed based on the at least two dubbing objects to obtain the plurality of dubbing files.

[0011] In this technical solution, when the dubbing information includes the text to be dubbed and the visual objects in the multimedia work, and the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags, the terminal device can obtain at least two sub-texts to be dubbed and at least two dubbing objects based on the visual objects in the multimedia work and the text to be dubbed. Then, the at least two sub-texts to be dubbed are dubbed using the at least two dubbing objects, resulting in multiple dubbing files displayed in the dubbing file interface. This can improve the accuracy of generating multiple dubbing files, thereby improving the feasibility of the solution.

[0012] In some possible implementations of this application, the step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control may include: processing the dubbing information through a target model in response to a trigger operation on the first generation control to obtain the plurality of dubbing files, and displaying the dubbing file interface; wherein the dubbing information includes the text to be dubbed, and the target model is obtained by model training based on historical text to be dubbed and historical dubbing files; or, the dubbing information includes the text to be dubbed and the image objects in the multimedia work, and the target model is obtained by model training based on historical text to be dubbed, image objects in historical multimedia works, and historical dubbing files.

[0013] In this technical solution, the terminal device responds to the trigger operation of the first generation control and determines the multiple dubbing files according to the dubbing information. It is explained that the text to be dubbed can be input into the target model and multiple dubbing files can be output; or the text to be dubbed and the screen objects in the multimedia work can be input into the target model and multiple dubbing files can be output. That is, multiple dubbing files can be obtained through the target model, which can improve the efficiency of generating multiple dubbing files.

[0014] In some possible implementations of this application, the method may further include: in response to a modification operation on the target dubbing file among the plurality of dubbing files, displaying the modified target dubbing file in the dubbing file interface; the step of generating a dubbed multimedia work based on the plurality of dubbing files in response to a trigger operation on the second generation control may include: in response to a trigger operation on the second generation control, generating a dubbed multimedia work based on the dubbing files other than the target dubbing file among the plurality of dubbing files, and the modified target dubbing file.

[0015] In this embodiment, the terminal device can intelligently split and dub the text to be dubbed, generating multiple dubbing files. If the user is not satisfied with the target dubbing file among the multiple dubbing files, the user can modify the target dubbing file. The modified target dubbing file is then displayed in the dubbing file interface. Subsequently, a dubbed multimedia work can be generated based on the modified target dubbing file and the multiple dubbing files excluding the target dubbing file. This improves the user's flexibility in modifying the target dubbing file among the multiple dubbing files, thereby increasing the efficiency of modifying the target dubbing file and the efficiency of generating dubbed multimedia works.

[0016] In some possible implementations of this application, displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file of the plurality of dubbing files may include: displaying the target dubbing file modified with respect to the target dubbing object corresponding to the target dubbing file of the plurality of dubbing files in the dubbing file interface; and / or, displaying the target dubbing file modified with respect to the target dubbing subtext corresponding to the target dubbing file of the plurality of dubbing files in the dubbing file interface.

[0017] In this technical solution, the user can modify the target dubbing object corresponding to the target dubbing file, and / or the target dubbing subtext. Then, in response to the modification operation of the target dubbing file in the multiple dubbing files, the terminal device displays the modified target dubbing file in the dubbing file interface. Subsequently, a dubbed multimedia work can be generated based on the dubbing files other than the target dubbing file in the multiple dubbing files, as well as the modified target dubbing file. This can improve the user's flexibility in modifying the target dubbing file in multiple dubbing files, thereby improving the user satisfaction rate of the subsequently generated dubbed multimedia work.

[0018] In some possible implementations of this application, the dubbing file interface further includes a batch control; before displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files, the method may further include: in response to a trigger operation on the batch selection control, the target dubbing file is selected in the dubbing file interface; the displaying the modified target dubbing file in response to a modification operation on the target dubbing file among the plurality of dubbing files may include: in response to a modification operation on the selected target dubbing file among the plurality of dubbing files, displaying the modified target dubbing file in the dubbing file interface.

[0019] In this technical solution, a batch selection control can be used to select target dubbing files, ensuring that each target dubbing file is in a selected state. The selected target dubbing files can be modified by the user, while other dubbing files cannot be modified. This improves the efficiency of user modification of target dubbing files, leading to the generation of a more satisfactory dubbed multimedia work.

[0020] In some possible implementations of this application, the dubbing file interface further includes a control for modifying the dubbing object; the step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation of the target dubbing object corresponding to the target dubbing file in the plurality of dubbing files may include: displaying an add dubbing object interface in response to a trigger operation of the modify dubbing object control, the add dubbing object interface including a plurality of dubbing object identifiers; in response to a selection operation of a first target dubbing object identifier, the first target dubbing object identifier is selected among the plurality of dubbing object identifiers included in the add dubbing object interface, the add dubbing object interface including a first completion control; in response to a trigger operation of the first completion control, modifying the target dubbing object corresponding to the target dubbing file to the first target dubbing object corresponding to the first target dubbing object identifier, and displaying the target dubbing file in the dubbing file interface with the target dubbing object identifier updated to the first target dubbing object identifier.

[0021] In this technical solution, in response to a modification operation on the target dubbing object corresponding to the target dubbing file in the multiple dubbing files, a detailed scheme of the modified target dubbing file is displayed in the dubbing file interface. This improves the feasibility of the solution and also increases the flexibility for users to modify the target dubbing file in multiple dubbing files, thereby improving the user satisfaction rate of the subsequently generated dubbed multimedia works.

[0022] In some possible implementations of this application, the step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing sub-text corresponding to the target dubbing file among the plurality of dubbing files may include: in response to a trigger operation on the first dubbing sub-text corresponding to the first dubbing file, the first dubbing sub-text is in an edit pre-selection state in the dubbing file interface, and the target dubbing file includes the first dubbing file; in response to a trigger operation on the first dubbing sub-text in the edit pre-selection state, a modification text interface is displayed, the modification text interface including the first dubbing sub-text in an editable state; in response to an editing operation on the first dubbing sub-text in the editable state, the first dubbing sub-text after the editing operation is displayed in the modification text interface, the modification text interface including a second completion control; in response to a trigger operation on the second completion control, the first dubbing file with the first dubbing sub-text updated to the first dubbing sub-text after the editing operation is displayed in the dubbing file interface.

[0023] In this technical solution, in response to the modification operation of the target dubbing sub-text corresponding to the target dubbing file in the multiple dubbing files, a detailed scheme of the target dubbing file after the target dubbing sub-text is displayed in the dubbing file interface. This improves the feasibility of the solution and also increases the flexibility of users to modify the target dubbing file in multiple dubbing files, thereby improving the user satisfaction rate of the subsequently generated dubbed multimedia works.

[0024] In some possible implementations of this application, displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files may include: displaying the adjusted target dubbing file in the dubbing file interface in response to an adjustment operation on the playback parameters of the target dubbing file among the plurality of dubbing files, wherein the playback parameters include at least one of volume and speech rate.

[0025] In this technical solution, after intelligently splitting and dubbing the text to be dubbed to generate multiple dubbing files and displaying them on the dubbing file interface, if the user is not satisfied with the multiple dubbing files, assuming they are not satisfied with the playback parameters corresponding to the target dubbing file among the multiple dubbing files, they can adjust the playback parameters corresponding to the target dubbing file. The target dubbing file with the adjusted playback parameters is then displayed on the dubbing file interface. Subsequently, based on the modified target dubbing file and the multiple dubbing files excluding the target dubbing file, a dubbed multimedia work can be generated. This can improve the user's flexibility in modifying the target dubbing file among multiple dubbing files, thereby improving the efficiency of modifying the target dubbing file.

[0026] In some possible implementations of this application, before displaying the dubbing text interface, the method may further include: displaying a first multimedia interface, the first multimedia interface including the multimedia work without subtitles and a dubbing control; and displaying a first text interface to be added in response to a trigger operation on the dubbing control.

[0027] The voiceover text display interface may include: when the first text interface to be added is a text interface to be added corresponding to multi-object voiceovers, displaying the voiceover text interface in response to an operation of adding text in the first text interface to be added; or, when the first text interface to be added is a text interface to be added corresponding to a single-object voiceover, and the first text interface to be added includes a multi-object voiceover control, displaying a second text interface to be added in response to a trigger operation on the multi-object voiceover control; and displaying the voiceover text interface in response to an operation of adding text in the second text interface to be added.

[0028] In this technical solution, when the multimedia work is a multimedia work without subtitles, the user can select the text interface to be added according to their own needs. When the text interface to be added is a text interface corresponding to multiple voice-overs, the voice-over text interface is displayed in response to the operation of adding text in the first text interface to be added. Thus, multiple voice-over files of multiple voice-over objects can be intelligently generated according to the text to be voice-over in the voice-over text interface, thereby improving the efficiency of generating dubbed multimedia works.

[0029] In some possible implementations of this application, the display of the dubbing text interface may include: displaying a second multimedia interface, the second multimedia interface including the multimedia work with added subtitles and dubbing controls;

[0030] The step of displaying the dubbing file interface in response to a trigger operation on the first generating control may include: displaying an add dubbing object interface in response to a trigger operation on the dubbing control, the add dubbing object interface including a full-text dubbing control; and displaying the dubbing file interface according to the screen with added subtitles in response to a trigger operation on the full-text dubbing control.

[0031] In this technical solution, when the multimedia work has added subtitles, the user can directly trigger the operation of generating multiple dubbing files based on the multimedia work with added subtitles using the full-text dubbing control, thereby improving the efficiency of generating dubbed multimedia works.

[0032] A second aspect of this application provides a dubbing device, which may include:

[0033] The display module is configured to display a dubbing text interface in response to an operation of dubbing a multimedia work. The dubbing text interface includes a text to be dubbed for dubbing the multimedia work and a first generation control. In response to a trigger operation of the first generation control, the display module is configured to display a dubbing file interface. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects.

[0034] The generation module is used to generate a dubbed multimedia work based on the plurality of dubbing files in response to a trigger operation of the second generation control.

[0035] In a third aspect, this application provides a terminal device including a memory, one or more processors, and a display, wherein the memory and the processors are coupled, the memory is used to store computer program code, the computer program code including computer instructions, and when the processor executes the computer instructions, the terminal device performs the method described in any of the first aspects.

[0036] In a fourth aspect, this application provides a chip including a processor, a memory, and a display. The processor may be a logic circuit, an integrated circuit, or a general-purpose processor, etc. The memory stores instructions. The processor may implement the method in the first aspect by reading the software code stored in the memory. The memory may be integrated into the processor or located outside the processor and exist independently.

[0037] In a fifth aspect of this application, a chip system is provided, the chip system being applied to a terminal device, the chip system including one or more interface circuits and one or more processors, and a display, the interface circuits and the processors being interconnected via lines, the interface circuits being configured to receive signals from a memory of the terminal device and send the signals to the processors, the signals including computer instructions stored in the memory, and when the processor executes the computer instructions, the terminal device performing the method described in any of the first aspects.

[0038] In a sixth aspect, this application provides a computer program product comprising: a computer program (also referred to as code or instructions) that, when run, causes a computer to perform the method described in the first aspect above.

[0039] A seventh aspect of this application provides a computer-readable storage medium storing a computer program (also referred to as code or instructions) that, when run on a computer, causes the computer to perform the methods described in the first aspect above. Attached Figure Description

[0040] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments and the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and other drawings can be obtained based on these drawings.

[0041] Figure 1 A schematic diagram illustrating multi-person voice-over for multimedia works in existing technology;

[0042] Figure 2 This is a schematic diagram of one embodiment of the dubbing method in this application;

[0043] Figure 3A This is a schematic diagram showing the voiceover text interface in an embodiment of this application;

[0044] Figure 3B This is a schematic diagram illustrating the switching from the text-to-be-added interface to the voice-over text interface in an embodiment of this application;

[0045] Figure 3C This is a schematic diagram illustrating the switching from the voiceover text interface to the voiceover file interface in an embodiment of this application;

[0046] Figure 3D This is a schematic diagram illustrating the process of obtaining multiple sub-texts to be dubbed from the text to be dubbed in an embodiment of this application;

[0047] Figure 3E This is a schematic diagram illustrating the generation of a dubbed multimedia work based on multiple dubbing files in an embodiment of this application.

[0048] Figure 4 This is a schematic diagram of another embodiment of the dubbing method in this application;

[0049] Figure 5A This is a schematic diagram illustrating a batch selection control in the dubbing file interface as described in this application embodiment;

[0050] Figure 5B This is a schematic diagram of a target dubbing object corresponding to a target dubbing file modified in an embodiment of this application;

[0051] Figure 5C This is another schematic diagram illustrating the modification of the target dubbing object corresponding to the target dubbing file in this embodiment of the application;

[0052] Figure 5D This is a schematic diagram illustrating the modification of the first dubbing subtext corresponding to the first dubbing file in an embodiment of this application;

[0053] Figure 5EThis is a schematic diagram illustrating the modification of playback parameters corresponding to the second dubbing file in an embodiment of this application;

[0054] Figure 5F This is a schematic diagram of the adjustment interface display text options in an embodiment of this application;

[0055] Figure 5G This is a schematic diagram illustrating the full-text modification of the dubbing sub-texts corresponding to multiple dubbing files in the embodiments of this application.

[0056] Figure 6 This is a schematic diagram illustrating another embodiment of the dubbing method in this application;

[0057] Figure 7A This is a schematic diagram of a first multimedia interface displayed as a first text interface to be added, as described in an embodiment of this application.

[0058] Figure 7B This is a schematic diagram of the first text interface to be added in an embodiment of this application, corresponding to the text interface to be added for multi-object dubbing;

[0059] Figure 7C This is a schematic diagram illustrating the switching from the text interface to be added corresponding to a single-object dubbing to the text interface to be added corresponding to multiple-object dubbing in an embodiment of this application.

[0060] Figure 8 This is a schematic diagram illustrating another embodiment of the dubbing method in this application;

[0061] Figure 9A This is a schematic diagram of an interface for adding voice-over objects displayed from the second multimedia interface in an embodiment of this application;

[0062] Figure 9B This is a schematic diagram showing the display of the dubbing file interface from the interface for adding dubbing objects in an embodiment of this application;

[0063] Figure 10A This is a schematic diagram of one embodiment of the dubbing device in this application;

[0064] Figure 10B This is a schematic diagram of one embodiment of the terminal device in this application;

[0065] Figure 11 This is a schematic diagram of another embodiment of the terminal device in this application. Detailed Implementation

[0066] This application provides a dubbing method, apparatus, device, storage medium, and program product, which can be used by terminal devices to intelligently split and dub the text to be dubbed, generate multiple dubbing files, and improve the efficiency of generating dubbed multimedia works.

[0067] To enable those skilled in the art to better understand the present application, the technical solutions of the embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. All embodiments based on the present application should fall within the scope of protection of the present application.

[0068] In existing technologies, when adding multiple voice-overs to existing multimedia works, multiple text selection and voice-over operations are required to obtain the multimedia work with multiple voice-overs. This process is complex, inefficient in generating multimedia works with multiple voice-overs, and lacks intelligent sentence segmentation; voice-over character recognition relies on manual intervention and is not intelligent. Figure 1 The diagram shown is a schematic representation of multi-person voice-over in existing technologies for multimedia works.

[0069] In this technical solution, the terminal device can intelligently split sentences and dub the text to be dubbed, generating multiple dubbing files corresponding to multiple dubbing objects, thereby improving the efficiency of generating dubbed multimedia works.

[0070] In the embodiments of this application, the terminal device may be a mobile phone, a tablet computer, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical care, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, or a wireless terminal device in a smart home, etc.

[0071] By way of example and not limitation, in this embodiment, the terminal device can also be a wearable device with a display interface. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0072] Additionally, it should be noted that the interface in this application embodiment can also be called a user interface (UI), which is the medium interface for interaction and information exchange between the application or operating system and the user. It realizes the conversion between the internal form of information and the form that the user can accept. The application's user interface is the source code written in a specific computer language such as Java or Extensible Markup Language (XML). The interface source code is parsed and rendered on the electronic device, and finally presented as content that the user can recognize, such as images, text, buttons, and other controls. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, images, and text. The attributes and content of controls in the interface are defined through tags or nodes, such as XML through... <textview> 、 <imgview> 、 <videoview>Nodes define the controls contained in the interface. A node corresponds to a control or property in the interface, and after being parsed and rendered, the node is presented as the content visible to the user.

[0073] The most common form of user interface is the graphical user interface (GUI), which refers to an interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0074] The technical solution of this application will be further described below by way of embodiments, such as... Figure 2 The diagram shown is a schematic representation of an embodiment of the dubbing method in this application, which may include the following steps:

[0075] 201. In response to the operation of dubbing a multimedia work, a dubbing text interface is displayed, the dubbing text interface including text to be dubbed for dubbing the multimedia work and a first generation control.

[0076] For example, if a user needs to add voiceover to a multimedia work, they can perform the voiceover operation. In response to this operation, the terminal device displays a voiceover text interface. This voiceover text interface includes the text to be voiced and a first generation control. The text to be voiced is the text used for adding voiceover to the multimedia work. This voiceover text interface is a text interface for multi-person voiceovers.

[0077] like Figure 3A The image shown is a schematic diagram illustrating the display of the voice-over text interface in an embodiment of this application. (Reference) Figure 3A The voiceover text interface includes the text to be voiced and a save control. That is, the first generation control in this application can be used as... Figure 3A The save control shown is illustrated as an example.

[0078] For example, the text to be dubbed is shown below:

[0079] Woman: Do you know which dynasty the Four Great Classical Novels were written in?

[0080] Man: It was written in the Ming Dynasty, but it doesn't seem to be entirely true.

[0081] Woman: Dream of the Red Chamber was written in the Qing Dynasty.

[0082] In some possible implementations, the text to be dubbed can be user-inputted text, text generated by artificial intelligence (AI), text obtained through a link, text extracted from a video, or other acquisition methods. This application embodiment does not impose specific limitations.

[0083] like Figure 3B The image shown is a schematic diagram illustrating the switching from the text-to-be-added interface to the voice-over text interface in an embodiment of this application. (Reference) Figure 3B The text-to-add interface includes AI-generated controls, link retrieval controls, and video extraction controls. Users can input the text to be dubbed for a multimedia work using a virtual keyboard, or click any of the controls to display the corresponding text. For example, when a user clicks... Figure 3B The video extraction control shown in the image displays a voiceover text interface, which shows the voiceover text extracted using the video extraction control.

[0084] 202. In response to the triggering operation of the first generation control, a dubbing file interface is displayed. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects.

[0085] For example, a user can trigger an operation on the first generation control, such as a click, press, double-click, or voice control operation. In response to the user's trigger operation on the first generation control, the terminal device displays a dubbing file interface. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. Moreover, the dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects, that is, the number of multiple dubbing files is greater than or equal to the number of at least two dubbing objects.

[0086] For example, in response to the triggering operation of the first generation control, the text to be dubbed, "Woman: Do you know which dynasty the Four Great Classical Novels were written in? Man: They were written in the Ming Dynasty, but I don't think so. Woman: Dream of the Red Chamber was written in the Qing Dynasty." can be intelligently segmented and dubbed to generate 4 dubbing files, which are then displayed on the dubbing file interface.

[0087] like Figure 3C The image shown is a schematic diagram illustrating the switching from a voiceover text interface to a voiceover file interface in an embodiment of this application. Figure 3C As shown, the dubbing file interface includes four dubbing files, and their corresponding dubbing sub-texts are as follows:

[0088] Dubbing sub-text 1: Do you know which dynasty the Four Great Classical Novels were written in?

[0089] Dubbing sub-text 2: It was written in the Ming Dynasty, right?

[0090] Dubbing subtext 3: It doesn't seem to be entirely true.

[0091] Dubbing sub-text 4: Dream of the Red Chamber was written in the Qing Dynasty.

[0092] Using the girl next door as an example, and the guy from Northeast China as an example, the voice-over objects corresponding to voice-over files 1 and 4 are used for illustration, while the voice-over objects corresponding to voice-over files 2 and 3 are used for illustration.

[0093] In some possible implementations of this application, the step of displaying the dubbing file interface in response to a trigger operation on the first generation control may include: in response to a trigger operation on the first generation control, determining the plurality of dubbing files based on dubbing information, and displaying the dubbing file interface, wherein the dubbing information includes the text to be dubbed, or the dubbing information includes the text to be dubbed and the image objects in the multimedia work.

[0094] For example, after the user triggers the first generation control, the terminal device responds to the user's triggering operation of the first generation control by displaying the dubbing file interface. Specifically, the terminal device may respond to the user's triggering operation of the first generation control by generating multiple dubbing files based on the text to be dubbed, or it may generate multiple dubbing files based on the text to be dubbed and the screen objects in the multimedia work. The specific implementation of this application is not limited. After generating multiple dubbing files, the multiple dubbing files are displayed on the dubbing file interface.

[0095] This technical solution provides a specific implementation method for a terminal device to display a dubbing file interface in response to a trigger operation of the first generation control. It provides a method to generate multiple dubbing files based on the text to be dubbed, or based on the text to be dubbed and the screen objects in the multimedia work, and then display them on the dubbing file interface, thereby improving the feasibility of the solution.

[0096] In some possible implementations of this application, the step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control can include, but is not limited to, the following implementations:

[0097] Implementation method 1: The dubbing information includes the text to be dubbed, and the text to be dubbed includes the text to be dubbed corresponding to at least two character tags.

[0098] The step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control may include: in response to a trigger operation on the first generation control, obtaining at least two sub-texts to be dubbed based on the text to be dubbed corresponding to the at least two character tags, wherein each character tag corresponds to at least one sub-text to be dubbed, and the at least one sub-text to be dubbed corresponding to each character tag is obtained based on the text to be dubbed corresponding to each character tag; determining at least two dubbing objects corresponding to the at least two character tags, wherein each character tag corresponds to one dubbing object; dubbing the at least two sub-texts to be dubbed using the at least two dubbing objects to obtain the plurality of dubbing files, wherein each dubbing file is obtained by dubbing one sub-text to be dubbed corresponding to each character tag using the dubbing object corresponding to each character tag.

[0099] For example, when the dubbing information includes text to be dubbed, if the text to be dubbed includes text corresponding to at least two character tags, after the user triggers the first generation control, the terminal device, in response to the user's triggering operation on the first generation control, can obtain at least two sub-texts to be dubbed based on the text to be dubbed corresponding to at least two character tags. It should be noted that here, the text to be dubbed can be split into sentences based on at least two character tags to obtain at least two sub-texts to be dubbed. Each character tag may include one sub-text to be dubbed, or multiple sub-texts to be dubbed; at least one sub-text to be dubbed corresponding to each character tag is obtained based on the text to be dubbed corresponding to each character tag.

[0100] The terminal device can also identify at least two voice-over objects corresponding to each of the at least two character tags, with each character tag corresponding to one voice-over object; then, it uses these at least two voice-over objects to dub at least two sub-texts to be dubbed, resulting in multiple dubbing files. Since the order of the at least two sub-texts to be dubbed is obtained by splitting them according to their original order, the device can dub at least two sub-texts according to the order of the sub-texts and the voice-over objects corresponding to their respective character tags, thus obtaining multiple dubbing files, which are then displayed on the dubbing file interface.

[0101] For example, the text to be dubbed is shown below:

[0102] Woman: Do you know which dynasty the Four Great Classical Novels were written in?

[0103] Man: It was written in the Ming Dynasty, but it doesn't seem to be entirely true.

[0104] Woman: Dream of the Red Chamber was written in the Qing Dynasty.

[0105] In the text to be dubbed, "female" and "male" can be understood as character tags.

[0106] In this technical solution, when the dubbing information includes the text to be dubbed, and the text to be dubbed includes dubbing text corresponding to at least two character tags, the terminal device responds to the trigger operation of the first generation control and determines the multiple dubbing files according to the dubbing information. Specifically, based on the dubbing text corresponding to at least two character tags, at least two sub-texts to be dubbed and at least two dubbing objects can be obtained. Then, dubbing is performed on at least two sub-texts to be dubbed using at least two dubbing objects, resulting in multiple dubbing files displayed on the dubbing file interface. This improves the accuracy of generating multiple dubbing files, thereby enhancing the feasibility of the solution. In other words, it can identify character tags in the text to be dubbed, break down sentences according to the character tags, and automatically match dubbing objects suitable for the character's image. This can also be understood as timbre, and the speech rate and pitch can be automatically adjusted.

[0107] Implementation method 2: The dubbing information includes the text to be dubbed and the visual objects in the multimedia work, and the text to be dubbed includes the text to be dubbed corresponding to at least two character tags.

[0108] The step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control may include: in response to a trigger operation on the first generation control, obtaining at least two sub-texts to be dubbed based on the text to be dubbed corresponding to the at least two character tags and the screen objects in the multimedia work, wherein each character tag corresponds to at least one sub-text to be dubbed, and the at least one sub-text to be dubbed corresponding to each character tag is obtained based on the text to be dubbed corresponding to each character tag, and the number of at least two character tags is less than or equal to the number of screen objects in the multimedia work; determining at least two dubbing objects corresponding to the at least two character tags and the screen objects in the multimedia work, wherein each character tag corresponds to one dubbing object, and the number of at least two dubbing objects is less than or equal to the number of screen objects in the multimedia work; dubbing the at least two sub-texts to be dubbed using the at least two dubbing objects to obtain the plurality of dubbing files, wherein each dubbing file is obtained by dubbing one sub-text to be dubbed corresponding to each character tag using the dubbing object corresponding to each character tag.

[0109] For example, when the dubbing information includes the text to be dubbed and the scene objects in the multimedia work, if the text to be dubbed includes text corresponding to at least two character tags, after the user triggers the first generation control, the terminal device, in response to the user's triggering operation on the first generation control, can obtain at least two sub-texts to be dubbed based on the text to be dubbed corresponding to at least two character tags and the scene objects in the multimedia work. It should be noted that the text to be dubbed can be segmented based on the text to be dubb ...

[0110] The terminal device can also identify at least two voice-over objects corresponding to at least two character tags and screen objects in the multimedia work, with each character tag corresponding to one voice-over object; then, it uses these at least two voice-over objects to dub at least two sub-texts to be dubbed, resulting in multiple dubbed text files. Since the order of the at least two sub-texts to be dubbed is obtained by splitting them according to the order of the sub-texts and the order of the screen objects in the multimedia work, it is possible to dub at least two sub-texts according to the order of the sub-texts to be dubbed and the order of the screen objects in the multimedia work, based on the voice-over objects corresponding to the character tags of the at least two sub-texts to be dubbed, thus obtaining multiple dubbed files, which are then displayed on the dubbed file interface.

[0111] In this technical solution, when the dubbing information includes the text to be dubbed and the visual objects in the multimedia work, and the text to be dubbed includes the text to be dubbed corresponding to at least two character tags, the terminal device responds to the trigger operation of the first generation control and determines the multiple dubbing files according to the dubbing information. Specifically, based on the text to be dubbed corresponding to at least two character tags and the visual objects in the multimedia work, at least two sub-texts to be dubbed and at least two dubbing objects can be obtained. Then, dubbing is performed on at least two sub-texts to be dubbed using at least two dubbing objects, resulting in multiple dubbing files displayed on the dubbing file interface. Generating multiple dubbing files using the text to be dubbed corresponding to at least two character tags and the visual objects in the multimedia work has higher accuracy and improves the feasibility of the solution.

[0112] Implementation method 3: The dubbing information includes the text to be dubbed and the image objects in the multimedia work, and the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags.

[0113] The step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control may include: in response to a trigger operation on the first generation control, obtaining at least two sub-texts to be dubbed based on the screen objects in the multimedia work and the text to be dubbed; determining the at least two dubbing objects based on the screen objects in the multimedia work and the text to be dubbed, wherein the number of the at least two dubbing objects is less than or equal to the number of screen objects in the multimedia work; and dubbing the at least two sub-texts to be dubbed based on the at least two dubbing objects to obtain the plurality of dubbing files.

[0114] For example, when the dubbing information includes the text to be dubbed and the visual objects in the multimedia work, if the text to be dubbed does not include the text corresponding to at least two character tags, after the user triggers the first generation control, the terminal device, in response to the user's triggering operation on the first generation control, can obtain at least two sub-texts to be dubbed and at least two dubbing objects based on the visual objects in the multimedia work and the text to be dubbed. Then, dubbing is performed on the at least two sub-texts to be dubbed using the at least two dubbing objects to obtain multiple dubbing files. That is, specifically, at least two sub-texts to be dubbed and at least two dubbing objects can be obtained based on the content of each frame of visual objects and the text to be dubbed in the multimedia work. Because the order of the at least two sub-texts to be dubbed is obtained by splitting the text according to its order, multiple dubbing files can be obtained by dubbing the at least two sub-texts according to the order of the at least two sub-texts to be dubbed, based on the at least two dubbing objects. These multiple dubbing files are then displayed on the dubbing file interface. It should be noted that some visual objects in a multimedia work may not have corresponding voice-over subtext, and therefore, there may be no corresponding voice-over object. This results in the number of at least two voice-over objects being less than or equal to the number of visual objects in the multimedia work.

[0115] In this technical solution, when the dubbing information includes the text to be dubbed and the visual objects in the multimedia work, and the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags, the terminal device responds to the trigger operation of the first generation control and determines the multiple dubbing files according to the dubbing information. Specifically, based on the visual objects in the multimedia work and the text to be dubbed, at least two sub-texts to be dubbed and at least two dubbing objects can be obtained. Then, dubbing is performed on at least two sub-texts to be dubbed using at least two dubbing objects, resulting in multiple dubbing files displayed on the dubbing file interface. This improves the accuracy of generating multiple dubbing files, thereby enhancing the feasibility of the solution.

[0116] Implementation method 4:

[0117] The dubbing information includes the screen objects in the multimedia work for which subtitles have been added, and the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags. The subtitles added in the multimedia work are the text to be dubbed.

[0118] The step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control may include: obtaining at least two sub-texts to be dubbed based on screen objects with added subtitles in the multimedia work; determining the at least two dubbing objects based on screen objects with added subtitles in the multimedia work, wherein the number of the at least two dubbing objects is less than or equal to the number of screen objects in the multimedia work; and dubbing the at least two sub-texts to be dubbed based on the at least two dubbing objects to obtain the plurality of dubbing files.

[0119] For example, when the dubbing information includes screen objects with added subtitles in the multimedia work, if the text to be dubbed does not include the text corresponding to at least two character tags, after the user triggers the first generation control, the terminal device, in response to the user's trigger operation on the first generation control, can obtain at least two sub-texts to be dubbed and at least two dubbing objects based on the screen objects with added subtitles in the multimedia work. Then, it can dub the at least two sub-texts to be dubbed using the at least two dubbing objects to obtain multiple dubbing files. That is, specifically, at least two sub-texts to be dubbed and at least two dubbing objects can be obtained based on the position of the screen object with added subtitles in each frame of the multimedia work and the corresponding added subtitles. Because the order of the at least two sub-texts to be dubbed is obtained by splitting according to the order of the added subtitles, here, the at least two sub-texts to be dubbed can be dubbed according to the order of the at least two sub-texts to be dubbed, based on the at least two dubbing objects, to obtain multiple dubbing files, and then display the multiple dubbing files on the dubbing file interface. It should be noted that some visual objects in a multimedia work may not have corresponding voice-over subtext, and therefore, there may be no corresponding voice-over object. This results in the number of at least two voice-over objects being less than or equal to the number of visual objects in the multimedia work.

[0120] In this technical solution, when the dubbing information includes screen objects with added subtitles in a multimedia work, and the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags, the terminal device, in response to the trigger operation of the first generation control, determines the multiple dubbing files based on the dubbing information. Specifically, based on the screen objects with added subtitles in the multimedia work, at least two sub-texts to be dubbed and at least two dubbing objects can be obtained. Then, dubbing is performed on at least two sub-texts to be dubbed using at least two dubbing objects, resulting in multiple dubbing files displayed in the dubbing file interface. This improves the accuracy of generating multiple dubbing files, thereby enhancing the feasibility of the solution. In other words, this technical solution can support recommending sub-texts to be dubbed and dubbing objects based on screen objects corresponding to screen objects with added subtitles in a multimedia work, thereby generating multiple dubbing files.

[0121] Implementation method 5:

[0122] The step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control may include: in response to a trigger operation on the first generation control, splitting the text to be dubbed into characters and / or words, and displaying the split text to be dubbed on the split text interface; in response to a selection operation on the split text to be dubbed, obtaining a plurality of sub-texts to be dubbed; and in response to the dubbing operation performed on the plurality of sub-texts to be dubbed, generating a plurality of dubbing files.

[0123] Exemplarily, the multiple dubbing files are obtained by selecting operation on the text to be dubbed that has been split into characters or words, and dubbing operation on the multiple sub-texts to be dubbed. That is, the multiple dubbing files can be dubbing files generated by intelligent sentence splitting and dubbing, or can be obtained by intelligently splitting the text to be dubbed into characters or words, then performing a selection operation on the text to be dubbed after character or word splitting to obtain multiple sub-texts to be dubbed, and then performing a dubbing operation on the multiple sub-texts to be dubbed to generate multiple dubbing files.

[0124] Such as Figure 3D shown, is a schematic diagram of obtaining multiple sub-texts to be dubbed according to the text to be dubbed in an embodiment of the present application. In Figure 3D shown, after triggering an operation on the save control in the dubbing text interface, a text-split interface is displayed. The text-split interface includes the text to be dubbed with characters and / or words split and a save control. For example: Do you know which dynasties the four great classical novels were written in? They were written in the Ming Dynasty. Well, not entirely. Dream of the Red Chamber was written in the Qing Dynasty. Suppose the user selects the words corresponding to "four", "great", "classical", "novels", "is", "which", "dynasty", "written", "?", and "Ming", "Dynasty", ",", "well", "like", "not", "entirely", "is", ".", and "Red", "Chamber", "Dream", "Qing", "Dynasty", "written", ".", then these selected words on the text-split interface are updated to the selected state; according to the text to be dubbed with the selected characters and / or words split, the multiple sub-texts to be dubbed can be:

[0125] Dubbing sub-text 1: Which dynasties were the four great classical novels written in?

[0126] Dubbing sub-text 2: Ming Dynasty,

[0127] Dubbing sub-text 3: Well, not entirely.

[0128] Dubbing sub-text 4: Dream of the Red Chamber was written in the Qing Dynasty.

[0129] Then perform a dubbing operation on these four dubbing sub-texts. For example, trigger an operation on the save control in the text-split interface to generate four dubbing files. Taking the dubbing objects corresponding to dubbing file 1 and dubbing file 4 as the neighboring girl for example, and taking the dubbing objects corresponding to dubbing file 2 and dubbing file 3 as the Northeast guy for example.

[0130] In this technical solution, multiple dubbing files are generated by selecting multiple sub-texts to be dubbed from the text to be dubbed (which has been split into characters or words), and then dubbing these sub-texts. Therefore, obtaining multiple dubbing files better meets user needs, reduces the probability of modifying multiple dubbing files, and improves the accuracy of generating multiple dubbing files.

[0131] In some possible implementations of this application, the step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control may include: processing the dubbing information through a target model in response to a trigger operation on the first generation control to obtain the plurality of dubbing files, and displaying the dubbing file interface; wherein the dubbing information includes the text to be dubbed, and the target model is obtained by model training based on historical text to be dubbed and historical dubbing files; or, the dubbing information includes the text to be dubbed and the image objects in the multimedia work, and the target model is obtained by model training based on historical text to be dubbed, image objects in historical multimedia works, and historical dubbing files.

[0132] For example, after the user triggers the first generation control, the terminal device responds to the user's triggering operation on the first generation control. If the target model is a model trained based on historical text to be dubbed and historical dubbing files, then the text to be dubbed can be directly input into the target model, outputting multiple dubbing files, and then displaying multiple dubbing files on the dubbing file interface. If the target model is a model trained based on historical text to be dubbed, screen objects in historical multimedia works, and historical dubbing files, then the text to be dubbed and screen objects in multimedia works can be directly input into the target model, outputting multiple dubbing files, and then displaying multiple dubbing files on the dubbing file interface.

[0133] In this technical solution, the terminal device responds to the trigger operation of the first generation control and determines the multiple dubbing files according to the dubbing information. It is explained that the text to be dubbed can be input into the target model and multiple dubbing files can be output; or the text to be dubbed and the screen objects in the multimedia work can be input into the target model and multiple dubbing files can be output. That is, multiple dubbing files can be obtained through the target model, which can improve the efficiency of generating multiple dubbing files.

[0134] In some possible implementations of this application, the method further includes: updating the target model to obtain an updated target model; the step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation of the first generation control may include: processing the dubbing information through the updated target model in response to a trigger operation of the first generation control to obtain the plurality of dubbing files, and displaying the dubbing file interface; wherein the dubbing information includes the text to be dubbed, and the updated target model is obtained by model training based on the updated historical text to be dubbed and the updated historical dubbing files; or, the dubbing information includes the text to be dubbed and the screen objects in the multimedia work, and the updated target model is obtained by model training based on the updated historical text to be dubbed, the screen objects in the updated historical multimedia work, and the updated historical dubbing files.

[0135] For example, the target model is not static, but can be updated based on updated historical data. For instance, the model can be trained based on the updated historical text to be dubbed and the updated historical dubbing files to obtain the updated target model; or the model can be trained based on the updated historical text to be dubbed, the visual objects in the updated historical multimedia works, and the updated historical dubbing files to obtain the updated target model.

[0136] After the user triggers the first generation control, the terminal device responds to the user's trigger operation. If the updated target model is a model trained based on the updated historical text to be dubbed and the updated historical dubbing files, then the text to be dubbed can be directly input into the updated target model, outputting multiple dubbing files, and then displaying multiple dubbing files on the dubbing file interface. If the updated target model is a model trained based on the updated historical text to be dubbed, the image objects in the updated historical multimedia works, and the updated historical dubbing files, then the text to be dubbed and the image objects in the multimedia works can be directly input into the updated target model, outputting multiple dubbing files, and then displaying multiple dubbing files on the dubbing file interface.

[0137] In this technical solution, the target model can be updated to ensure the accuracy of the updated target model. Furthermore, the terminal device's response to the triggering operation of the first generation control, and the determination of the multiple dubbing files based on the dubbing information, are specifically explained. The text to be dubbed can be input into the updated target model, outputting multiple dubbing files; alternatively, the text to be dubbed and visual objects from the multimedia work can be input into the updated target model, outputting multiple dubbing files. That is, multiple dubbing files are obtained through the updated target model, which ensures the accuracy of generating multiple dubbing files and also improves the efficiency of generating multiple dubbing files.

[0138] Different voice actors have different timbres. In voice acting, "timbre" refers to the unique quality or color of a voice, which makes each person's voice sound unique.

[0139] 203. In response to the triggering operation of the second generation control, generate a dubbed multimedia work based on the plurality of dubbing files.

[0140] For example, after displaying multiple dubbing files on the dubbing file interface, if the user is satisfied with these multiple dubbing files, then the user can directly trigger an operation on the second generation control on the dubbing file interface, such as a click operation, a press operation, a double-click operation, a voice control operation, etc. The terminal device responds to the user's trigger operation on the second generation control and generates a dubbed multimedia work based on the multiple dubbing files.

[0141] like Figure 3E The image shown is a schematic diagram illustrating the generation of a dubbed multimedia work based on multiple dubbing files in an embodiment of this application. Figure 3E In response to user actions triggering the control in the dubbing file interface, a dubbed multimedia work is generated based on the four dubbing files in the interface. This example uses a 35-second multimedia work. Optionally, the corresponding dubbing subtext (i.e., the subtitles) can be displayed in the subtitle display area of ​​the dubbed multimedia work, or it can be left undisplayed, depending on the user's needs. Optionally, the lower area of ​​the interface displaying the dubbed multimedia work can also display the video track (showing the corresponding playback video), the audio track (playing the sound of the corresponding dubbed object), and the subtitle track (displaying the corresponding dubbing subtext), for user viewing or editing.

[0142] In this embodiment, a voice-over text interface is displayed, comprising a text to be voiced and a first generation control. The text to be voiced is the text used to voice-over a multimedia work. In response to a trigger operation on the first generation control, a voice-over file interface is displayed, comprising multiple voice-over files and a second generation control. The multiple voice-over files are generated by splitting and voice-overing the text to be voiced, and the voice-over objects corresponding to the multiple voice-over files include at least two voice-over objects. In response to a trigger operation on the second generation control, a voice-over multimedia work is generated based on the multiple voice-over files. The terminal device can intelligently split and voice-over the text to be voiced, generating multiple voice-over files corresponding to multiple voice-over objects, thereby improving the efficiency of generating voice-over multimedia works. This solves the problem of high operational costs and low flexibility in multi-person voice-over scenarios in a single multimedia work, where users need to voice multiple speaking characters or add multiple voice-overs to the narration, thus improving the efficiency and quality of voice-over for multimedia works.

[0143] like Figure 4 The diagram shown is a schematic representation of another embodiment of the dubbing method in this application, which may include the following steps:

[0144] 401. In response to the operation of dubbing a multimedia work, a dubbing text interface is displayed, the dubbing text interface including text to be dubbed for dubbing the multimedia work and a first generation control.

[0145] 402. In response to the triggering operation of the first generation control, a dubbing file interface is displayed. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects.

[0146] It should be noted that steps 401-402 in the embodiments of this application are different from those in the present application. Figure 2 Steps 201-202 in the illustrated embodiment are similar and will not be repeated here.

[0147] 403. In response to the modification operation of the target dubbing file in the plurality of dubbing files, the modified target dubbing file is displayed in the dubbing file interface.

[0148] For example, after the intelligent system splits and dubs the text to be dubbed into sentences and generates multiple dubbing files, which are then displayed on the dubbing file interface, if the user is satisfied with the multiple dubbing files, they can directly generate a dubbed multimedia work based on these files. If the user is not satisfied with the multiple dubbing files, assuming they are not satisfied with the target dubbing file among the multiple dubbing files, they can modify the target dubbing file, and the modified target dubbing file will be displayed on the dubbing file interface. Subsequently, a dubbed multimedia work can be generated based on the modified target dubbing file and the multiple dubbing files excluding the target dubbing file. This can improve the user's flexibility in modifying the target dubbing file among the multiple dubbing files, thereby improving the efficiency of modifying the target dubbing file.

[0149] In some possible implementations of this application, displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files may include:

[0150] In response to a modification operation on a target dubbing object corresponding to a target dubbing file in the plurality of dubbing files, the modified target dubbing file is displayed in the dubbing file interface; and / or,

[0151] In response to a modification operation on the target dubbing subtext corresponding to the target dubbing file in the plurality of dubbing files, the target dubbing file after modification of the target dubbing subtext is displayed in the dubbing file interface.

[0152] For example, each of the multiple dubbing files includes a dubbing object and dubbing sub-text. A user can select a target dubbing file. The user can modify the dubbing object, the dubbing sub-text, or both of the dubbing object and dubbing sub-text in the target dubbing file; this embodiment does not impose specific limitations. In response to a modification operation on the target dubbing object corresponding to the target dubbing file, the terminal device displays the modified target dubbing file in the dubbing file interface; and / or, in response to a modification operation on the target dubbing sub-text corresponding to the target dubbing file, the terminal device displays the modified target dubbing file in the dubbing file interface.

[0153] In this technical solution, the user can modify the target dubbing object corresponding to the target dubbing file, and / or the target dubbing subtext. Then, in response to the modification operation of the target dubbing file in the multiple dubbing files, the terminal device displays the modified target dubbing file in the dubbing file interface. Subsequently, a dubbed multimedia work can be generated based on the dubbing files other than the target dubbing file in the multiple dubbing files, as well as the modified target dubbing file. This can improve the user's flexibility in modifying the target dubbing file in multiple dubbing files, thereby improving the user satisfaction rate of the subsequently generated dubbed multimedia work.

[0154] In some possible implementations of this application, the dubbing file interface further includes a batch control; before displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files, the method may further include: in response to a trigger operation on the batch selection control, the target dubbing file is selected in the dubbing file interface;

[0155] The step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files may include: displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file that is selected among the plurality of dubbing files.

[0156] For example, the dubbing file interface may also include a batch selection control. This batch selection control may include all selection controls, selection controls for multiple dubbing files corresponding to different dubbing object identifiers, and custom selection controls. Each dubbing file includes a dubbing object and dubbing sub-text. Users can use the batch selection control to select target dubbing files, thus selecting the target dubbing files. If the user selects the "All Selection" control, the terminal device responds to the user's trigger operation, and multiple dubbing files are selected in the dubbing file interface. If the user selects the "Girl Next Door" selection control, the terminal device responds to the user's trigger operation, and the corresponding dubbing files are selected in the dubbing file interface. If the user selects the "Northeast Guy" selection control, the terminal device responds to the user's trigger operation, and the corresponding dubbing files are selected in the dubbing file interface. If the user selects a custom selection control, which allows the user to select a target dubbing file according to their actual needs, the terminal device responds to the user's trigger operation, and multiple dubbing files are selectable in the dubbing file interface. The user can then select a target dubbing file from the selectable dubbing files, and the target dubbing file is selected in the dubbing file interface.

[0157] like Figure 5A The image shown is a schematic diagram illustrating a batch selection control in the dubbing file interface according to an embodiment of this application. (Used in...) Figure 5A As shown in the example, the user selects the "Northeast Xiaoge" control. In this case, all the dubbing files corresponding to "Northeast Xiaoge" are selected in the dubbing file interface. This example uses the target dubbing file as the dubbing file corresponding to "Northeast Xiaoge".

[0158] In this technical solution, a batch selection control can be used to select target dubbing files, ensuring that each target dubbing file is in a selected state. The selected target dubbing files can be modified by the user, while other dubbing files cannot be modified. This improves the efficiency of user modification of target dubbing files, leading to the generation of satisfactory dubbed multimedia works. In other words, it supports batch selection of target dubbing files, modification of dubbing objects and sub-text, resulting in higher modification efficiency.

[0159] In some possible implementations of this application, the dubbing file interface further includes a merging control; the step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the selected target dubbing file among the plurality of dubbing files may include: when the target dubbing file includes at least two dubbing files, in response to a trigger operation on the merging control, merging the selected target dubbing files, and displaying the merged target dubbing file in the dubbing file interface.

[0160] For example, the dubbing file interface may include a merge control, which can be used to merge target dubbing files. When the target dubbing file includes at least two dubbing files, the target dubbing files can be merged. In this case, the target dubbing sub-texts corresponding to the target dubbing files need to be merged, and the target dubbing objects corresponding to the target dubbing files also need to be merged. It should be noted that when merging the target dubbing objects corresponding to the target dubbing files, if the target dubbing objects corresponding to the target dubbing files are different, the user can be prompted which dubbing object will be the merged dubbing object. If the user is not satisfied with any of the target dubbing objects corresponding to the target dubbing files, they can also select the dubbing object by modifying the dubbing object control. If the target dubbing objects corresponding to the target dubbing files are the same, the same target dubbing object can be used as the merged dubbing object, or the dubbing object can be selected by modifying the dubbing object control. This application embodiment does not impose specific limitations. Optionally, the target dubbing file can be consecutive dubbing files from multiple dubbing files, or non-consecutive dubbing files from multiple dubbing files. This application embodiment does not impose specific limitations.

[0161] In this technical solution, when the target dubbing file includes at least two dubbing files, the target dubbing files can be merged, and the merged target dubbing file is displayed in the dubbing file interface. This improves the user's flexibility in modifying the target dubbing file, and can be used to generate a more satisfactory dubbed multimedia work later.

[0162] In some possible implementations of this application, the dubbing file interface further includes a control for modifying the dubbing object; the step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation of the target dubbing object corresponding to the target dubbing file in the plurality of dubbing files may include: displaying an add dubbing object interface in response to a trigger operation of the modify dubbing object control, the add dubbing object interface including a plurality of dubbing object identifiers; in response to a selection operation of a first target dubbing object identifier, the first target dubbing object identifier is selected among the plurality of dubbing object identifiers included in the add dubbing object interface, the add dubbing object interface including a first completion control; in response to a trigger operation of the first completion control, modifying the target dubbing object corresponding to the target dubbing file to the first target dubbing object corresponding to the first target dubbing object identifier, and displaying the target dubbing file in the dubbing file interface with the target dubbing object identifier updated to the first target dubbing object identifier.

[0163] For example, the dubbing file interface also includes a control for modifying the dubbing object. Users can trigger operations on this control, such as clicking, pressing, double-clicking, or voice control. The terminal device responds to the user's trigger operation on the dubbing object control by displaying an interface for adding dubbing objects. This interface includes multiple dubbing object identifiers, which users can use to select a first target dubbing object identifier. After the user selects the first target dubbing object identifier, it is selected among the multiple identifiers in the interface. The interface also includes a first completion control, which is used for user trigger operations, such as clicking, pressing, double-clicking, or voice control. Responding to the user's trigger operation on the first completion control, the terminal device can modify the target dubbing object corresponding to the target dubbing file to the first target dubbing object identified by the first target dubbing object identifier. The target dubbing object identifier corresponding to the target dubbing file is then updated and displayed as the first target dubbing object identifier in the dubbing file interface. This indicates that the target dubbing object corresponding to the target dubbing file has been modified, while the target dubbing subtext corresponding to the target dubbing file remains unchanged in the dubbing file interface. Optionally, the number of target dubbing files can be one or more, and this application embodiment does not specifically limit the number.

[0164] like Figure 5B The diagram shown is a schematic representation of modifying the target dubbing object corresponding to the target dubbing file in an embodiment of this application. The dubbing file interface also includes a control for modifying the dubbing object. Figure 5B The example shown uses the voice-over file corresponding to the Northeastern guy as the target voice-over file, and the voice-over object modification control as the speaker modification control. After the user triggers the speaker modification control, the "Add Voice-over Object" interface is displayed. The "Add Voice-over Object" interface includes multiple voice-over object identifiers. If the user selects the first target voice-over object identifier (e.g., Voice-over 2), Voice-over 2 will be selected. The voice-over object interface also includes a first completion control... Figure 5B The example of the checkmark control is as follows: After the user triggers the checkmark control, in the file display interface, the voice-over object identifiers in voice-over files 2 and 3, which correspond to the "Northeast Guy" identifier, are updated to the full name of the voice-over object identifier, i.e., voice-over object 2. The voice-over sub-texts corresponding to voice-over files 2 and 3 remain unchanged.

[0165] In this technical solution, in response to a modification operation on the target dubbing object corresponding to the target dubbing file in the multiple dubbing files, a detailed scheme of the modified target dubbing file is displayed in the dubbing file interface. This improves the feasibility of the solution and also increases the flexibility for users to modify the target dubbing file in multiple dubbing files, thereby improving the user satisfaction rate of the subsequently generated dubbed multimedia works.

[0166] In some possible implementations of this application, the dubbing file interface further includes dubbing object identifiers corresponding to the plurality of dubbing files; the step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file includes: in response to a trigger operation on the second dubbing object identifier corresponding to the second dubbing file, displaying an add dubbing object interface, the add dubbing object interface including a plurality of dubbing object identifiers; in response to a selection operation on the second target dubbing object identifier, the second target dubbing object identifier is selected among the plurality of dubbing object identifiers included in the add dubbing object interface, the add dubbing object interface including a first completion control; in response to a trigger operation on the first completion control, modifying the second dubbing object corresponding to the second dubbing file to the second target dubbing object identifier corresponding to the second target dubbing object identifier, and displaying the second dubbing file with the second dubbing object identifier updated to the second target dubbing object identifier in the dubbing file interface.

[0167] For example, the dubbing file interface also includes multiple dubbing object identifiers corresponding to dubbing files. Users can perform trigger operations on the dubbing object identifiers corresponding to different dubbing files, such as clicking, pressing, double-clicking, voice control, etc., to modify the dubbing object identifier corresponding to a single dubbing file. Assuming the user triggers an operation on the second dubbing identifier corresponding to the second dubbing file, the terminal device responds to the user's trigger operation and displays an add dubbing object interface. This add dubbing object interface includes multiple dubbing object identifiers, which can be used by the user to select a second target dubbing object identifier. If the user selects a second target dubbing object identifier, the add dubbing object interface includes... Among the multiple voice-over object identifiers, the second target voice-over object identifier is selected. The interface for adding a voice-over object includes a first completion control, which is used by the user to perform a trigger operation, such as a click operation, a press operation, a double-click operation, a voice control operation, etc. When the terminal device responds to the user's trigger operation on the first completion control, it can modify the second voice-over object corresponding to the second voice-over file to the second target voice-over object corresponding to the second target voice-over object identifier. In the voice-over file interface, the second voice-over object identifier corresponding to the second voice-over file is updated and displayed as the second target voice-over object identifier. At this time, it indicates that the second voice-over object corresponding to the second voice-over file has been modified. However, in the voice-over file interface, the second voice-over subtext corresponding to the second voice-over file remains unchanged.

[0168] like Figure 5C The diagram shown is another schematic representation of the target dubbing object corresponding to the target dubbing file modified in this embodiment of the application. Figure 5C In the illustration, using dubbing file 2 as an example, after the user triggers an operation on the dubbing object identifier corresponding to dubbing file 2, namely the "Northeast Guy" identifier, an "Add Dubbing Object" interface is displayed. This interface includes multiple dubbing object identifiers. If the user selects the first target dubbing object identifier from these identifiers (taking dubbing 2 as an example), dubbing 2 is selected. The dubbing object interface also includes a first completion control... Figure 5C The example of the checkmark control is as follows: After the user triggers the checkmark control, the voice-over object identifier of the voice-over file 2 corresponding to the Northeast Xiaoge identifier is updated to the full name of the voice-over 2 identifier, that is, voice-over object 2, in the file display interface. The voice-over subtext corresponding to voice-over file 2 remains unchanged.

[0169] This technical solution provides a method to display the modified second dubbing file in the dubbing file interface in response to a modification operation of the second dubbing object corresponding to the second dubbing file in the plurality of dubbing files. This improves the feasibility of the solution and also increases the flexibility of users to modify single dubbing files in multiple dubbing files, thereby improving the user satisfaction rate of the subsequently generated dubbed multimedia works.

[0170] In some possible implementations of this application, the step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing sub-text corresponding to the target dubbing file among the plurality of dubbing files may include: in response to a trigger operation on the first dubbing sub-text corresponding to the first dubbing file, the first dubbing sub-text is in an edit pre-selection state in the dubbing file interface, and the target dubbing file includes the first dubbing file; in response to a trigger operation on the first dubbing sub-text in the edit pre-selection state, a modification text interface is displayed, the modification text interface including the first dubbing sub-text in an editable state; in response to an editing operation on the first dubbing sub-text in the editable state, the first dubbing sub-text after the editing operation is displayed in the modification text interface, the modification text interface including a second completion control; in response to a trigger operation on the second completion control, the first dubbing file with the first dubbing sub-text updated to the first dubbing sub-text after the editing operation is displayed in the dubbing file interface.

[0171] For example, if a user wants to modify the first sub-text of the first dubbing file corresponding to the first dubbing file in the target dubbing file, they can first perform a trigger operation on the first sub-text of the first dubbing file, such as a click operation, a press operation, a double-click operation, a voice control operation, etc., so that the first sub-text of the first dubbing file is in an edit pre-selected state. That is, in response to the user's trigger operation on the first sub-text of the first dubbing file, the first sub-text of the first dubbing file is in an edit pre-selected state in the dubbing file interface. Then, the user can perform a trigger operation on the first sub-text of the first dubbing file in the edit pre-selected state. In response to the user's trigger operation on the first sub-text of the first dubbing file in the edit pre-selected state, the terminal device displays a text modification interface. The interface includes a first voice-over sub-text in an editable state, which the user can directly edit. Editing can include modifying the content of the first voice-over sub-text and adjusting at least one of its font, size, and color. In response to the editing operation on the editable first voice-over sub-text, the terminal device displays the edited first voice-over sub-text in the edited text interface. The edited text interface includes a second completion control for the user to save the edit, thus updating the first voice-over sub-text corresponding to the first voice-over file in the voice-over file interface to the edited first voice-over text. However, in the voice-over file interface, the first voice-over object identifier corresponding to the first voice-over file remains unchanged.

[0172] like Figure 5D The image shown is a schematic diagram illustrating the modification of the first dubbing sub-text corresponding to the first dubbing file in an embodiment of this application. Figure 5D In the illustration, taking dubbing file 3 as an example, after the user triggers an operation on the dubbing sub-text corresponding to dubbing file 3, the dubbing sub-text is in a pre-selected editing state, for example, adding a pen icon. If the user triggers an operation on the dubbing sub-text in the pre-selected editing state, a text modification interface will be displayed. The text modification interface includes a virtual keyboard for the user to edit the dubbing sub-text. After the user finishes editing, the text modification interface also includes a second completion control. Figure 5D The checkmark control shown is illustrated as an example. If the user triggers the checkmark control, the subtext of the dubbing file 3 displayed in the dubbing file interface will be the modified subtext of the dubbing file.

[0173] In this technical solution, in response to the modification operation of the target dubbing sub-text corresponding to the target dubbing file in the multiple dubbing files, a detailed scheme of the target dubbing file after the target dubbing sub-text is displayed in the dubbing file interface. This improves the feasibility of the solution and also increases the flexibility of users to modify the target dubbing file in multiple dubbing files, thereby improving the user satisfaction rate of the subsequently generated dubbed multimedia works.

[0174] In some possible implementations of this application, displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files may include: displaying the adjusted target dubbing file in the dubbing file interface in response to an adjustment operation on the playback parameters of the target dubbing file among the plurality of dubbing files, wherein the playback parameters include at least one of volume and speech rate.

[0175] For example, users can modify the target voice-over object, edit the corresponding voice-over sub-text, merge the target voice-over objects, and adjust the playback parameters of the target voice-over file. The details are as follows:

[0176] If a user adjusts the playback parameters of a target audio file among multiple audio files, the terminal device will respond to the user's adjustment by displaying the adjusted target audio file in the audio file interface. These playback parameters may include at least one of volume and speech rate. The number of target audio files can be one or more; typically, adjusting the playback parameters of a target audio file involves only one target audio file.

[0177] In this technical solution, after intelligently splitting and dubbing the text to be dubbed to generate multiple dubbing files and displaying them on the dubbing file interface, if the user is satisfied with the multiple dubbing files, they can directly generate a dubbed multimedia work based on the multiple dubbing files. If the user is not satisfied with the multiple dubbing files, assuming they are not satisfied with the playback parameters corresponding to the target dubbing file among the multiple dubbing files, they can adjust the playback parameters corresponding to the target dubbing file. The target dubbing file with adjusted playback parameters is then displayed on the dubbing file interface. Subsequently, a dubbed multimedia work can be generated based on the modified target dubbing file and the multiple dubbing files excluding the target dubbing file. This improves the user's flexibility in modifying the target dubbing file among the multiple dubbing files, thereby improving the efficiency of modifying the target dubbing file.

[0178] In some possible implementations of this application, the step of displaying the adjusted target dubbing file in the dubbing file interface in response to an adjustment operation on the playback parameters of the target dubbing file among the plurality of dubbing files may include: displaying an add dubbing object interface in response to a trigger operation on a second dubbing object identifier corresponding to a second dubbing file, wherein the target dubbing file includes the second dubbing file and the add dubbing object interface includes adjustment controls; displaying an adjustment interface in response to a trigger operation on the adjustment controls, wherein the adjustment interface includes at least one of a speech rate option and a volume option; and adjusting the playback parameters of the second dubbing object corresponding to the second dubbing object identifier in response to an adjustment operation on a target option, and displaying the adjusted target dubbing file in the dubbing file interface, wherein the target option is the speech rate option and / or the volume option.

[0179] For example, each of the multiple dubbing files includes a corresponding dubbing object identifier. If a user is dissatisfied with the playback parameters of a certain dubbing file, such as the second dubbing file, they can trigger an operation on the identifier of the second dubbing file. The terminal device, in response to the trigger operation on the second dubbing object identifier corresponding to the second dubbing file, displays an interface for adding dubbing objects. This interface may include adjustment controls, and optionally, multiple dubbing object identifiers. After triggering the adjustment controls, the user can adjust the playback parameters. After selecting multiple dubbing object identifiers, the user can modify the second dubbing object corresponding to the second dubbing file. In this embodiment, the user triggers the adjustment controls, and the terminal device, in response, displays an adjustment interface. This adjustment interface includes speech rate and volume options. If the user adjusts the speech rate option, the terminal device, in response to the user's speech rate adjustment, adjusts the speech rate of the second dubbing object corresponding to the second dubbing object identifier. If the user adjusts the volume option, the terminal device, in response to the user's volume adjustment, adjusts the volume of the second dubbing object corresponding to the second dubbing object identifier. The second dubbing file after the adjustment operation is displayed in the dubbing file interface.

[0180] like Figure 5E The diagram shown illustrates a modification of the playback parameters corresponding to the second dubbing file in an embodiment of this application. Figure 5E In the illustration, taking the second dubbing file, dubbing file 3, as an example, after the user triggers the operation on the dubbing object corresponding to dubbing file 3, such as the Northeastern guy icon corresponding to dubbing file 3, the Add Dubbing Object interface is displayed. The Add Dubbing Object interface includes adjustment controls and multiple dubbing object icons. Here, we take clicking the adjustment control as an example. If the user triggers the adjustment control, the adjustment interface is displayed. This adjustment interface displays the sliders corresponding to the speech rate and volume options, respectively. The speech rate and volume of dubbing file 3 can be adjusted by sliding the sliders.

[0181] In this technical solution, in response to the adjustment operation of the playback parameters of the target dubbing file in the multiple dubbing files, a detailed scheme of the target dubbing file after the adjustment operation is displayed in the dubbing file interface. This improves the feasibility of the solution and also increases the flexibility of users to modify the target dubbing file in multiple dubbing files, thereby improving the user satisfaction rate of the subsequently generated dubbed multimedia works.

[0182] In some possible implementations of this application, the adjustment interface further includes a text display option, which is used to determine whether to display the dubbing sub-text on the screen; the method may further include: in response to the operation of the text display option to display the dubbing sub-text on the screen, displaying the corresponding dubbing sub-text on the screen of the dubbed multimedia work; in response to the operation of the text display option to not display the dubbing sub-text on the screen, not displaying the corresponding dubbing sub-text on the screen of the dubbed multimedia work.

[0183] For example, the user can choose whether to display the corresponding voice-over subtext in subsequent dubbed multimedia works. Whether to display the corresponding voice-over subtext can be determined by displaying a text display option on the adjustment interface. The user can then manipulate this option to trigger the operation of displaying the voice-over subtext on the screen. If the user wants to display the voice-over subtext in the dubbed multimedia work, they can trigger the operation of displaying the voice-over subtext on the screen by adjusting the text display option. The terminal device will then display the corresponding voice-over subtext in the dubbed multimedia work in response to this operation. Conversely, if the user wants to not display the voice-over subtext in the dubbed multimedia work, they can trigger the operation of not displaying the voice-over subtext on the screen by adjusting the text display option. The terminal device will then not display the corresponding voice-over subtext in the dubbed multimedia work in response to this operation.

[0184] like Figure 5F The image shown is a schematic diagram of the text display options in the adjustment interface of this application embodiment. Figure 5F As shown, in addition to speech rate and volume options, the adjustment interface also includes text display options, which allow users to adjust whether to display the corresponding voice-over subtext (i.e., subtitles) in multimedia works according to their actual needs.

[0185] In this technical solution, users can also trigger the operation of whether to display the dubbing sub-text on the screen through the text display option, thereby displaying or not displaying the corresponding dubbing sub-text on the screen of the dubbed multimedia work, satisfying the user's need for whether to display the corresponding dubbing sub-text in the dubbed multimedia work, and improving the user experience.

[0186] In some possible implementations of this application, the dubbing file interface further includes a full-text editing control; the step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files may include: displaying the dubbing text interface in response to a trigger operation on the full-text editing control, wherein the text to be dubbed in the dubbing text interface is in an editable state; displaying the edited text to be dubbed in response to an editing operation on the editable text to be dubbed in the dubbing text interface; and displaying the edited plurality of dubbing files in response to a trigger operation on the first generation control; wherein the edited plurality of dubbing files are dubbing files generated by splitting and dubbing the edited text to be dubbed, and the dubbing objects corresponding to the edited plurality of dubbing files include at least two dubbing objects.

[0187] For example, after generating multiple voice-over files and displaying them on the voice-over file interface, if the user is not satisfied with the voice-over sub-texts corresponding to the multiple voice-over files, they can re-edit the voice-over sub-texts corresponding to the multiple voice-over files using the full-text editing control. That is, the user triggers the full-text editing control, such as by clicking, pressing, double-clicking, or using voice control. The terminal device responds to the user's triggering of the full-text editing control by displaying the voice-over text interface, in which the text to be voiced is in an editable state. The user can edit the text to be voiced in the editable state, and the terminal device responds to the editing of the text to be voiced by displaying the edited text to be voiced on the voice-over text interface. If the user is satisfied with the edited text to be voiced, they can trigger the first generation control to generate multiple edited voice-over files. The terminal device responds to the user's triggering of the first generation control by displaying the multiple edited voice-over files on the voice-over file interface. The multiple edited voice-over files are voice-over files generated by splitting and voice-overing the edited text to be voiced, and the voice-over objects corresponding to the multiple edited voice-over files include at least two voice-over objects.

[0188] like Figure 5G The image shown is a schematic diagram illustrating the full-text modification of dubbing sub-texts corresponding to multiple dubbing files in an embodiment of this application. Figure 5G As shown, if a user is not satisfied with the sub-texts of the multiple dubbing files generated based on the text to be dubbed, they can use the full-text editing control in the dubbing file interface to modify the sub-texts of the multiple dubbing files uniformly. This means the user will be redirected back to the dubbing text interface for operations such as re-editing the text to be dubbed or regenerating the entire text.

[0189] In this technical solution, multiple sub-texts corresponding to the generated multiple dubbing files, i.e., the text to be dubbed, can be edited using a full-text editing control. This allows for the generation of multiple edited dubbing files based on the edited text, and subsequently, a dubbed multimedia work can be generated based on these edited dubbing files. This provides a solution for full-text modification of the dubbing sub-texts corresponding to multiple dubbing files, thus improving the feasibility of the solution.

[0190] In some possible implementations of this application, the step of generating a dubbed multimedia work based on the plurality of dubbing files in response to a trigger operation on the second generation control may include step 405:

[0191] 405. In response to a trigger operation on the second generation control, generate a dubbed multimedia work based on the dubbing files other than the target dubbing file in the plurality of dubbing files and the modified target dubbing file.

[0192] For example, after displaying multiple dubbing files on the dubbing file interface, if the user is not satisfied with the target dubbing file among these multiple dubbing files, they can modify the target dubbing file. The terminal device responds to the user's modification operation on the target dubbing file among the multiple dubbing files by displaying the modified target dubbing file on the dubbing file interface. The dubbing file interface includes a second generation control. If the user is satisfied with the modified target dubbing file, they can directly trigger an operation on the second generation control on the dubbing file interface, such as a click operation, a press operation, a double-click operation, a voice control operation, etc. The terminal device responds to the user's trigger operation on the second generation control by generating a dubbed multimedia work based on the dubbing files other than the target dubbing file among the multiple dubbing files, and the modified target dubbing file.

[0193] In this embodiment, a dubbing text interface is displayed, which includes text to be dubbed and a first generation control. The text to be dubbed is text for dubbing a multimedia work. In response to a trigger operation of the first generation control, a dubbing file interface is displayed. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects. In response to a modification operation of a target dubbing file among the multiple dubbing files, the modified target dubbing file is displayed in the dubbing file interface. In response to a trigger operation of the second generation control, a dubbed multimedia work is generated based on the dubbing files other than the target dubbing file among the multiple dubbing files and the modified target dubbing file. The terminal device can intelligently break down and dub the text to be dubbed, generating multiple dubbing files. If the user is not satisfied with the target dubbing file among the multiple dubbing files, they can modify it. The modified target dubbing file is displayed in the dubbing file interface. Subsequently, based on the modified target dubbing file and the multiple dubbing files excluding the target dubbing file, a dubbed multimedia work can be generated. This increases the user's flexibility in modifying the target dubbing file among multiple dubbing files, thereby improving the efficiency of modifying the target dubbing file and the efficiency of generating dubbed multimedia works. It also improves the efficiency of configuring text-to-speech dubbing for the entire multimedia work; allows users to flexibly configure multiple dubbing objects for the same multimedia work at once, making the construction of multi-person dialogue scenes more efficient; and intelligently identifies the roles of the dubbing objects, presets the timbre and tone of the dubbing objects, and reduces selection costs.

[0194] like Figure 6 The diagram shown is a schematic representation of another embodiment of the dubbing method in this application, which may include:

[0195] 601. Display a first multimedia interface, the first multimedia interface including the multimedia work without subtitles and a dubbing control.

[0196] For example, when a user needs to publish a multimedia work and needs to add voice-over to the multimedia work, the terminal device displays a first multimedia interface. It is assumed that the first multimedia interface displays the multimedia work without subtitles and a voice-over control, so that the user can add voice-over to the multimedia work without subtitles.

[0197] 602. In response to the triggering operation of the dubbing control, display the first text interface to be added.

[0198] If a user wants to add voiceover to a multimedia work displayed on the first multimedia interface that has no subtitles, the first multimedia interface includes voiceover controls. The user can trigger these controls through actions such as clicking, pressing, double-clicking, or voice control. The terminal device responds to these actions by displaying the first text interface to be added. Figure 7A The image shown is a schematic diagram illustrating how a first multimedia interface is displayed as a first text-to-be-added interface in an embodiment of this application. Figure 7A As shown, the multimedia works displayed on the first multimedia interface do not have subtitles. If the user wants to add voiceover to the multimedia works, they can trigger the voiceover control included in the first multimedia interface to display the first text interface to be added. Here, we will take the text interface to be added as a single object voiceover, that is, the text interface to be added corresponding to single person voiceover (or single person broadcasting) as an example for explanation.

[0199] In some possible implementations of this application, displaying the voiceover text interface may include step 603 or 604, as shown below:

[0200] 603. If the first text interface to be added is a text interface to be added corresponding to multiple object dubbing, the dubbing text interface is displayed in response to the operation of adding text in the first text interface to be added.

[0201] For example, when the first text interface to be added is a multi-object voiceover, that is, a text interface to be added corresponding to multiple voiceovers, the user can directly add text in the first text interface to be added. In response to the operation of adding text in the first text interface to be added, the terminal device displays the voiceover text interface, which displays the text to be voiced.

[0202] like Figure 7B The image shown is a schematic diagram of the first text interface to be added in an embodiment of this application, corresponding to a multi-object dubbing interface. Figure 7B As shown, the multimedia works displayed on the first multimedia interface do not have subtitles. If the user wants to add voiceover to the multimedia works, they can trigger the voiceover control included in the first multimedia interface to display the first text interface to be added. Here, we take the first text interface to be added as an example of a text interface to be added for multiple voiceovers, that is, the text interface to be added for multiple voiceovers.

[0203] For example, the user can enter the text to be dubbed on the first text to be added interface, or the text to be dubbed can be generated by AI, or the text to be dubbed can be generated by a link, or the text to be dubbed can be generated by a video, or other generation methods. This application embodiment does not make specific limitations.

[0204] 604. If the first text interface to be added is a text interface to be added corresponding to a single object dubbing, and the first text interface to be added includes a multi-object dubbing control, in response to the triggering operation of the multi-object dubbing control, the second text interface to be added is displayed; in response to the operation of adding text in the second text interface to be added, the dubbing text interface is displayed.

[0205] For example, when the first text interface to be added is a single-object voiceover, that is, a text interface to be added corresponding to a single person's voiceover (or a single person's broadcast), the user can switch the first text interface to be added to the second text interface to be added, that is, a text interface to be added corresponding to multiple people's voiceover. The user can directly add text in the second text interface to be added. The terminal device responds to the operation of adding text in the second text interface to be added by displaying the voiceover text interface, which displays the text to be voiced.

[0206] like Figure 7C The image shown is a schematic diagram illustrating the switching between the text interface to be added corresponding to a single-object dubbing and the text interface to be added corresponding to multiple-object dubbing in an embodiment of this application. Figure 7C As shown, if the first text interface to be added is a single-object dubbing interface by default, that is, the text interface to be added corresponding to single-person dubbing (or single-person broadcasting), but the user wants to perform multi-object dubbing, the user can switch to multi-object dubbing, that is, the text interface to be added corresponding to multi-person dubbing, by triggering the multi-person dubbing control included in the first text interface to be added. Then, the user can perform the operation of obtaining the text to be dubbed and display the dubbing text interface.

[0207] 605. In response to the triggering operation of the first generation control, a dubbing file interface is displayed. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects.

[0208] 606. In response to a trigger operation on the second generation control, generate a dubbed multimedia work based on the plurality of dubbing files.

[0209] It should be noted that steps 605 and 606 in the embodiments of this application are different from those in the present application. Figure 2 Steps 202 and 203 in the illustrated embodiment are similar and will not be repeated here.

[0210] In this embodiment, a first multimedia interface is displayed, comprising the multimedia work without subtitles and a dubbing control. In response to a trigger operation on the dubbing control, a first text interface to be added is displayed. If the first text interface to be added is a text interface corresponding to multi-object dubbing, the dubbing text interface is displayed in response to an operation of adding text to the first text interface to be added. If the first text interface to be added is a text interface corresponding to single-object dubbing and includes multi-object dubbing controls, a second text interface to be added is displayed in response to a trigger operation on the multi-object dubbing controls. The dubbing text interface is displayed in response to an operation of adding text to the second text interface to be added. When the multimedia work is without subtitles, the user can select the text interface to be added according to their needs. When the text interface to be added is a text interface corresponding to multi-object dubbing, the dubbing text interface is displayed in response to an operation of adding text to the first text interface to be added. This allows for the intelligent generation of multiple dubbing files for multiple dubbing objects based on the text to be dubbed in the dubbing text interface, thereby improving the efficiency of generating dubbed multimedia works.

[0211] like Figure 8 The diagram shown is a schematic representation of another embodiment of the dubbing method in this application, which may include:

[0212] In some possible implementations of this application, the display of the voiceover text interface may include step 801:

[0213] 801. Display a second multimedia interface, the second multimedia interface including the multimedia work with added subtitles and a dubbing control.

[0214] For example, when a user needs to publish a multimedia work and needs to add voice-over to the multimedia work, the terminal device displays a second multimedia interface. It is assumed that the second multimedia interface displays the multimedia work with added subtitles and voice-over controls, so that the user can add voice-over to the multimedia work with added subtitles.

[0215] In some possible implementations of this application, the step of displaying the dubbing file interface in response to a trigger operation on the first generation control may include steps 802 and 803:

[0216] 802. In response to the triggering operation of the dubbing control, display the add dubbing object interface, the add dubbing object interface including the full text dubbing control.

[0217] For example, if a user wants to add voiceover to a multimedia work with subtitles displayed on the second multimedia interface, since the second multimedia interface includes a voiceover control, the user can trigger operations on the voiceover control, such as clicking, pressing, double-clicking, voice control, etc. The terminal device responds to the user's trigger operation on the voiceover control and displays an interface for adding voiceover objects. Figure 9A The image shown is a schematic diagram illustrating the display of an interface for adding voice-over objects from the second multimedia interface in an embodiment of this application. Figure 9A As shown, the second multimedia interface displays a multimedia work with subtitles added. Furthermore, the second multimedia interface also includes a dubbing control; here, the intelligent dubbing control is used as an example. The interface for adding dubbing objects is displayed, and this interface includes a full-text dubbing control used to dub the multimedia work with added subtitles. Optionally, the interface for adding dubbing objects may also include adjustment controls and at least one of multiple dubbing object identifiers.

[0218] In the case where the second multimedia interface includes the multimedia work with added subtitles, if the user performs intelligent dubbing on a single sentence with added subtitles, the dubbing object interface includes a full-text dubbing control. The user can choose to trigger the full-text dubbing control to perform full-text dubbing, which supports switching from single-sentence dubbing to full-text dubbing; or the user can choose a third target dubbing object among multiple dubbing objects in the dubbing object interface to dub the single sentence with added subtitles. This application embodiment does not make specific limitations.

[0219] 803. In response to the trigger operation of the full-text dubbing control, display the dubbing file interface according to the screen with added subtitles.

[0220] For example, because the second multimedia interface displays a multimedia work with added subtitles, when a user dubs multiple subtitle objects on the multimedia work, they do not need to add additional text to be dubbed. Instead, they use the already added subtitles as the text to be dubbed. They only need to trigger an operation on the full-text dubbing control, such as a click, press, double-click, or voice control operation. The terminal device responds to the trigger operation on the full-text dubbing control by displaying the dubbing file interface based on the screen with added subtitles. This dubbing file interface displays multiple dubbing files, and the dubbing objects corresponding to these multiple dubbing files include at least two dubbing objects. Figure 9B The image shown is a schematic diagram illustrating the display of a dubbing file interface from the interface for adding dubbing objects in an embodiment of this application. Figure 9B As shown, users can trigger the full-text dubbing control in the interface for adding dubbing objects. Then, the terminal device can intelligently generate multiple dubbing files based on the subtitles already added to the multimedia work and display them in the dubbing file interface.

[0221] 804. In response to a trigger operation on the second generation control, generate a dubbed multimedia work based on the plurality of dubbing files.

[0222] It should be noted that step 804 in the embodiments of this application is different from... Figure 2 Step 203 in the illustrated embodiment is similar and will not be repeated here.

[0223] In this embodiment, a second multimedia interface is displayed, including the multimedia work with added subtitles and a dubbing control. In response to a trigger operation on the dubbing control, an interface for adding dubbing objects is displayed, including a full-text dubbing control. In response to a trigger operation on the full-text dubbing control, an interface for displaying dubbing files is shown based on the screen with added subtitles. In response to a trigger operation on the second generation control, a dubbed multimedia work is generated based on the multiple dubbing files. When the multimedia work is one with added subtitles, the user can directly trigger the operation of generating multiple dubbing files from the multimedia work with added subtitles using the full-text dubbing control, thereby improving the efficiency of generating dubbed multimedia works.

[0224] like Figure 10A The diagram shown is a schematic representation of one embodiment of the dubbing device in this application, which may include:

[0225] Display module 1001 is configured to display a dubbing text interface in response to an operation of dubbing a multimedia work. The dubbing text interface includes text to be dubbed for dubbing the multimedia work and a first generation control. In response to a trigger operation of the first generation control, a dubbing file interface is displayed. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects.

[0226] The generation module 1002 is used to generate a dubbed multimedia work based on the plurality of dubbing files in response to a trigger operation of the second generation control.

[0227] In some possible implementations of this application, the dubbing device further includes a processing module 1003.

[0228] The processing module 1003 is used to respond to the trigger operation of the first generation control, determine the plurality of dubbing files according to the dubbing information, and the display module 1001 is used to display the dubbing file interface, wherein the dubbing information includes the text to be dubbed, or the dubbing information includes the text to be dubbed and the screen objects in the multimedia work.

[0229] In some possible implementations of this application, the dubbing information includes the text to be dubbed, and the text to be dubbed includes the text to be dubbed corresponding to at least two character tags;

[0230] The processing module 1003 is specifically configured to respond to a trigger operation on the first generation control, obtain at least two sub-texts to be dubbed based on the texts to be dubbed corresponding to the at least two character tags, wherein each character tag corresponds to at least one sub-text to be dubbed, and the at least one sub-text to be dubbed corresponding to each character tag is obtained based on the texts to be dubbed corresponding to each character tag; determine at least two dubbing objects corresponding to the at least two character tags, wherein each character tag corresponds to one dubbing object; and dub the at least two sub-texts to be dubbed using the at least two dubbing objects to obtain the plurality of dubbing files, wherein each dubbing file is obtained by dubbing one sub-text to be dubbed corresponding to each character tag using the dubbing object corresponding to each character tag.

[0231] In some possible implementations of this application, the dubbing information includes the text to be dubbed and the image objects in the multimedia work, wherein the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags;

[0232] The processing module 1003 is specifically configured to respond to a trigger operation on the first generation control, obtain at least two sub-texts to be dubbed based on the screen objects in the multimedia work and the text to be dubbed; determine the at least two dubbing objects based on the screen objects in the multimedia work and the text to be dubbed, wherein the number of the at least two dubbing objects is less than or equal to the number of screen objects in the multimedia work; and dub the at least two sub-texts to be dubbed based on the at least two dubbing objects to obtain the plurality of dubbing files.

[0233] In some possible implementations of this application, the processing module 1003 is specifically used to process the dubbing information through the target model in response to the trigger operation of the first generation control, to obtain the multiple dubbing files, and to display the dubbing file interface;

[0234] The dubbing information includes the text to be dubbed, and the target model is obtained by training the model based on historical texts to be dubbed and historical dubbing files; or...

[0235] The dubbing information includes the text to be dubbed and the visual objects in the multimedia work. The target model is obtained by training the model based on historical texts to be dubbed, visual objects in historical multimedia works, and historical dubbing files.

[0236] In some possible implementations of this application, the display module 1001 is further configured to display the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files;

[0237] The generation module 1002 is specifically used to generate a dubbed multimedia work in response to a trigger operation of the second generation control, based on the dubbing files other than the target dubbing file in the plurality of dubbing files and the modified target dubbing file.

[0238] In some possible implementations of this application, the display module 1001 is specifically used to display the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing object corresponding to the target dubbing file in the plurality of dubbing files; and / or,

[0239] The display module 1001 is specifically used to respond to the modification operation of the target dubbing subtext corresponding to the target dubbing file in the plurality of dubbing files, and to display the target dubbing file after the modification of the target dubbing subtext in the dubbing file interface.

[0240] In some possible implementations of this application, the dubbing file interface also includes a batch control;

[0241] The processing module 1003 is also configured to respond to a trigger operation on the batch selection control, wherein the target dubbing file is selected in the dubbing file interface;

[0242] Display module 1001 is specifically used to display the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file that is selected among the plurality of dubbing files.

[0243] In some possible implementations of this application, the dubbing file interface also includes a control for modifying the dubbing object;

[0244] The processing module 1003 is specifically configured to, in response to a trigger operation on the modified voice-over object control, display an add voice-over object interface, the add voice-over object interface including multiple voice-over object identifiers; in response to a selection operation on a first target voice-over object identifier, the first target voice-over object identifier is selected among the multiple voice-over object identifiers included in the add voice-over object interface, the add voice-over object interface including a first completion control; in response to a trigger operation on the first completion control, modify the target voice-over object corresponding to the target voice-over file to the first target voice-over object corresponding to the first target voice-over object identifier;

[0245] The display module 1001 is specifically used to display the target dubbing file in the dubbing file interface, wherein the target dubbing object identifier is updated to the first target dubbing object identifier.

[0246] In some possible implementations of this application, the processing module 1003 is specifically configured to respond to a trigger operation on a first dubbing sub-text corresponding to a first dubbing file, wherein the first dubbing sub-text is in an edit pre-selection state in the dubbing file interface, and the target dubbing file includes the first dubbing file; in response to a trigger operation on the first dubbing sub-text in the edit pre-selection state, a modified text interface is displayed, wherein the modified text interface includes the first dubbing sub-text in an editable state; in response to an editing operation on the first dubbing sub-text in the editable state, the modified text interface displays the first dubbing sub-text after the editing operation, wherein the modified text interface includes a second completion control;

[0247] Display module 1001 is specifically used to display the first dubbing file with the first dubbing subtext updated to the first dubbing subtext after the editing operation in response to the trigger operation of the second completion control in the dubbing file interface.

[0248] In some possible implementations of this application, the display module 1001 is specifically used to display the adjusted target dubbing file in the dubbing file interface in response to the adjustment operation of the playback parameters of the target dubbing file in the plurality of dubbing files, wherein the playback parameters include at least one of volume and speech rate.

[0249] In some possible implementations of this application, the display module 1001 is further configured to display a first multimedia interface, the first multimedia interface including the multimedia work without subtitles and a dubbing control; in response to a trigger operation on the dubbing control, a first text interface to be added is displayed;

[0250] The display module 1001 is specifically used to, in the case that the first text interface to be added is a text interface to be added corresponding to multiple object dubbing, respond to the operation of adding text in the first text interface to be added, and display the dubbing text interface; or...

[0251] The display module 1001 is specifically used to display a second text interface to be added in response to a trigger operation on the multi-object dubbing control when the first text interface to be added is a text interface to be added corresponding to a single object dubbing and the first text interface to be added includes a multi-object dubbing control; and to display the dubbing text interface in response to an operation of adding text in the second text interface to be added.

[0252] In some possible implementations of this application, the display module 1001 is specifically used to display a second multimedia interface, which includes the multimedia work with added subtitles and a dubbing control; in response to a trigger operation on the dubbing control, an interface for adding a dubbing object is displayed, which includes a full-text dubbing control; in response to a trigger operation on the full-text dubbing control, the interface for displaying the dubbing file is displayed according to the screen with added subtitles.

[0253] like Figure 10B The diagram shown is a schematic representation of an embodiment of the terminal device described in this application, which may include, for example: Figure 10A The dubbing device shown.

[0254] like Figure 11 The diagram shown is a schematic diagram of another embodiment of the terminal device in this application. The following is in conjunction with... Figure 11 A detailed introduction to the various components of a mobile phone in a terminal device:

[0255] RF circuit 1110 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with processor 1180; additionally, it transmits uplink data to the base station. Typically, RF circuit 1110 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 1110 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0256] The memory 1120 can be used to store software programs and modules. The processor 1180 executes various functions and data processing of the mobile phone by running the software programs and modules stored in the memory 1120. The memory 1120 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 1120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0257] The input unit 1130 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 1130 may include a touch panel 1131 and other input devices 1132. The touch panel 1131, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 1131), and drive the corresponding connection devices according to a pre-set program. Optionally, the touch panel 1131 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 1180, and can receive and execute commands sent by the processor 1180. In addition, the touch panel 1131 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1131, the input unit 1130 may also include other input devices 1132. Specifically, other input devices 1132 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0258] Display unit 1140 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. Display unit 1140 may include display panel 1141, optionally configured as a Liquid Crystal Display (LCD), Organic Light-Emitting Diode (OLED), or similar display panel 1141. Further, touch panel 1131 may cover display panel 1141. When touch panel 1131 detects a touch operation on or near it, it transmits the information to processor 1180 to determine the type of touch event. Subsequently, processor 1180 provides corresponding visual output on display panel 1141 based on the type of touch event. Although in Figure 11 In this embodiment, the touch panel 1131 and the display panel 1141 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 1131 and the display panel 1141 can be integrated to realize the input and output functions of the mobile phone.

[0259] The mobile phone may also include at least one sensor 1150, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 1141 according to the ambient light level, and the proximity sensor can turn off the display panel 1141 and / or the backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.

[0260] Audio circuit 1160, speaker 1161, and microphone 1162 provide an audio interface between the user and the mobile phone. Audio circuit 1160 converts received audio data into electrical signals and transmits them to speaker 1161, where speaker 1161 converts them into sound signals for output. On the other hand, microphone 1162 converts collected sound signals into electrical signals, which are received by audio circuit 1160, converted into audio data, and then processed by processor 1180 before being transmitted via RF circuit 1110 to, for example, another mobile phone, or the audio data can be output to memory 1120 for further processing.

[0261] Wi-Fi is a short-range wireless transmission technology. Through the Wi-Fi module 1170, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 11 Wi-Fi module 1170 is shown, but it is understood that it is not an essential component of a mobile phone and can be omitted as needed without changing the nature of the invention.

[0262] The processor 1180 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 1120, and calls data stored in the memory 1120 to perform various functions and process data, thereby providing overall monitoring of the phone. Optionally, the processor 1180 may include one or more processing units; preferably, the processor 1180 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1180.

[0263] The mobile phone also includes a power supply 1190 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 1180 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.

[0264] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.

[0265] In this embodiment, the display unit 1140 is configured to display a dubbing text interface in response to an operation of dubbing a multimedia work. The dubbing text interface includes a text to be dubbed for dubbing the multimedia work and a first generation control. In response to a trigger operation on the first generation control, a dubbing file interface is displayed. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects.

[0266] Processor 1180 is configured to generate a dubbed multimedia work based on the plurality of dubbing files in response to a trigger operation of the second generation control.

[0267] In some possible implementations of this application, the dubbing device also includes a processor 1180.

[0268] The processor 1180 is configured to, in response to a trigger operation on the first generation control, determine the plurality of dubbing files based on dubbing information, and the display unit 1140 is configured to display the dubbing file interface, wherein the dubbing information includes the text to be dubbed, or the dubbing information includes the text to be dubbed and a scene object in the multimedia work.

[0269] In some possible implementations of this application, the dubbing information includes the text to be dubbed, and the text to be dubbed includes the text to be dubbed corresponding to at least two character tags;

[0270] The processor 1180 is specifically configured to respond to a trigger operation on the first generation control, obtain at least two sub-texts to be dubbed based on the texts to be dubbed corresponding to the at least two character tags, wherein each character tag corresponds to at least one sub-text to be dubbed, and the at least one sub-text to be dubbed corresponding to each character tag is obtained based on the texts to be dubbed corresponding to each character tag; determine at least two dubbing objects corresponding to the at least two character tags, wherein each character tag corresponds to one dubbing object; and dub the at least two sub-texts to be dubbed using the at least two dubbing objects to obtain the plurality of dubbing files, wherein each dubbing file is obtained by dubbing one sub-text to be dubbed corresponding to each character tag using the dubbing object corresponding to each character tag.

[0271] In some possible implementations of this application, the dubbing information includes the text to be dubbed and the image objects in the multimedia work, wherein the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags;

[0272] The processor 1180 is specifically configured to respond to a trigger operation on the first generation control, obtain at least two sub-texts to be dubbed based on the screen objects in the multimedia work and the text to be dubbed; determine the at least two dubbing objects based on the screen objects in the multimedia work and the text to be dubbed, wherein the number of the at least two dubbing objects is less than or equal to the number of screen objects in the multimedia work; and dub the at least two sub-texts to be dubbed based on the at least two dubbing objects to obtain the plurality of dubbing files.

[0273] In some possible implementations of this application, the processor 1180 is specifically configured to, in response to a trigger operation on the first generation control, process the dubbing information through the target model to obtain the plurality of dubbing files, and display the dubbing file interface;

[0274] The dubbing information includes the text to be dubbed, and the target model is obtained by training the model based on historical texts to be dubbed and historical dubbing files; or...

[0275] The dubbing information includes the text to be dubbed and the visual objects in the multimedia work. The target model is obtained by training the model based on historical texts to be dubbed, visual objects in historical multimedia works, and historical dubbing files.

[0276] In some possible implementations of this application, the display unit 1140 is further configured to display the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files;

[0277] The processor 1180 is specifically configured to, in response to a trigger operation on the second generation control, generate a dubbed multimedia work based on the dubbing files other than the target dubbing file in the plurality of dubbing files, and the modified target dubbing file.

[0278] In some possible implementations of this application, the display unit 1140 is specifically used to display the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing object corresponding to the target dubbing file in the plurality of dubbing files; and / or,

[0279] The display unit 1140 is specifically used to display the target dubbing file after the target dubbing subtext has been modified in the dubbing file interface in response to the modification operation of the target dubbing subtext corresponding to the target dubbing file in the plurality of dubbing files.

[0280] In some possible implementations of this application, the dubbing file interface also includes a batch control;

[0281] The processor 1180 is also configured to respond to a triggering operation on the batch selection control, wherein the target dubbing file is selected in the dubbing file interface;

[0282] The display unit 1140 is specifically used to display the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file that is selected among the plurality of dubbing files.

[0283] In some possible implementations of this application, the dubbing file interface also includes a control for modifying the dubbing object;

[0284] The processor 1180 is specifically configured to, in response to a trigger operation on the modified voice-over object control, display an add voice-over object interface, the add voice-over object interface including multiple voice-over object identifiers; in response to a selection operation on a first target voice-over object identifier, the first target voice-over object identifier is selected among the multiple voice-over object identifiers included in the add voice-over object interface, the add voice-over object interface including a first completion control; in response to a trigger operation on the first completion control, modify the target voice-over object corresponding to the target voice-over file to the first target voice-over object corresponding to the first target voice-over object identifier;

[0285] The display unit 1140 is specifically used to display the target dubbing file in the dubbing file interface, wherein the target dubbing object identifier is updated to the first target dubbing object identifier.

[0286] In some possible implementations of this application, the processor 1180 is specifically configured to respond to a trigger operation on a first dubbing sub-text corresponding to a first dubbing file, wherein the first dubbing sub-text is in an edit pre-selected state in the dubbing file interface, and the target dubbing file includes the first dubbing file; in response to the trigger operation on the first dubbing sub-text in the edit pre-selected state, a modified text interface is displayed, wherein the modified text interface includes the first dubbing sub-text in an editable state; in response to an editing operation on the first dubbing sub-text in the editable state, the modified text interface displays the first dubbing sub-text after the editing operation, wherein the modified text interface includes a second completion control;

[0287] The display unit 1140 is specifically used to display the first dubbing file with the first dubbing subtext updated to the first dubbing subtext after the editing operation in response to the triggering operation of the second completion control in the dubbing file interface.

[0288] In some possible implementations of this application, the display unit 1140 is specifically used to display the adjusted target dubbing file in the dubbing file interface in response to an adjustment operation of the playback parameters of the target dubbing file in the plurality of dubbing files, wherein the playback parameters include at least one of volume and speech rate.

[0289] In some possible implementations of this application, the display unit 1140 is further configured to display a first multimedia interface, the first multimedia interface including the multimedia work without subtitles and a dubbing control; in response to a trigger operation on the dubbing control, a first text interface to be added is displayed;

[0290] The display unit 1140 is specifically configured to, in response to an operation of adding text to the first text interface to be added (where the first text interface to be added is a text interface to be added corresponding to multiple object dubbing), display the dubbing text interface; or...

[0291] The display unit 1140 is specifically used to display a second text interface to be added in response to a triggering operation of the multi-object dubbing control when the first text interface to be added is a text interface to be added corresponding to a single object dubbing and the first text interface to be added includes a multi-object dubbing control; and to display the dubbing text interface in response to an operation of adding text in the second text interface to be added.

[0292] In some possible implementations of this application, the display unit 1140 is specifically used to display a second multimedia interface, which includes the multimedia work with added subtitles and a dubbing control; in response to a trigger operation on the dubbing control, an interface for adding a dubbing object is displayed, which includes a full-text dubbing control; in response to a trigger operation on the full-text dubbing control, the interface for displaying the dubbing file is displayed according to the screen with added subtitles.

[0293] This application embodiment further provides a computer-readable storage medium storing computer instructions that, when executed on a terminal device, cause the terminal device to perform the above-described method embodiment.

[0294] The aforementioned computer-readable storage medium may take the form of any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0295] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0296] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, radio frequency (RF), etc., or any suitable combination thereof.

[0297] Computer program code for performing the operations described herein can be written in one or more programming languages ​​or a combination thereof, including resource-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0298] This application also provides a computer program product that, when run on a computer, causes the computer to perform some or all of the steps described in the method embodiments above.

[0299] This application provides a chip system including a processor and potentially a memory, for implementing the functions of the terminal device described in the aforementioned method. The chip system may be composed of chips or may include chips and other discrete components.

[0300] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0301] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0302] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0303] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0304] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0305] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.< / videoview> < / imgview> < / textview>

Claims

1. A dubbing method, characterized in that, include: In response to the operation of dubbing a multimedia work, a dubbing text interface is displayed, the dubbing text interface including text to be dubbed for dubbing the multimedia work and a first generation control; In response to the triggering operation of the first generation control, a dubbing file interface is displayed. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects. In response to a triggering operation on the second generation control, a dubbed multimedia work is generated based on the plurality of dubbing files.

2. The method according to claim 1, characterized in that, The step of displaying the dubbing file interface in response to a trigger operation on the first generation control includes: In response to a trigger operation on the first generation control, the plurality of dubbing files are determined based on the dubbing information, and the dubbing file interface is displayed. The dubbing information includes the text to be dubbed, or the dubbing information includes the text to be dubbed and the image objects in the multimedia work.

3. The method according to claim 2, characterized in that, The dubbing information includes the text to be dubbed, which includes the text to be dubbed corresponding to at least two character tags; The step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control includes: In response to the triggering operation of the first generation control, at least two sub-texts to be dubbed are obtained according to the dubbed texts corresponding to the at least two character tags, each character tag corresponds to at least one sub-text to be dubbed, and the at least one sub-text to be dubbed corresponding to each character tag is obtained based on the dubbed texts corresponding to each character tag; Identify at least two voice actors corresponding to the at least two character tags, with each character tag corresponding to one voice actor; The at least two voice-over objects are used to dub the at least two sub-texts to be dubbed, resulting in the plurality of dubbing files. Each dubbing file is obtained by dubbing one sub-text to be dubbed corresponding to each character tag using the voice-over object corresponding to each character tag.

4. The method according to claim 2, characterized in that, The dubbing information includes the text to be dubbed and the visual objects in the multimedia work, and the text to be dubbed does not include the text to be dubbed corresponding to at least two character tags; The step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control includes: In response to the triggering operation of the first generating control, at least two sub-texts to be dubbed are obtained based on the screen objects in the multimedia work and the text to be dubbed. Based on the visual objects in the multimedia work and the text to be dubbed, at least two dubbing objects are determined, wherein the number of the at least two dubbing objects is less than or equal to the number of visual objects in the multimedia work; The at least two sub-texts to be dubbed are dubbed based on the at least two dubbing objects to obtain the plurality of dubbing files.

5. The method according to any one of claims 2-4, characterized in that, The step of determining the plurality of dubbing files based on dubbing information in response to a trigger operation on the first generation control includes: In response to the triggering operation of the first generating control, the dubbing information is processed through the target model to obtain the multiple dubbing files, and the dubbing file interface is displayed; The dubbing information includes the text to be dubbed, and the target model is obtained by training the model based on historical texts to be dubbed and historical dubbing files; or... The dubbing information includes the text to be dubbed and the visual objects in the multimedia work. The target model is obtained by training the model based on historical texts to be dubbed, visual objects in historical multimedia works, and historical dubbing files.

6. The method according to any one of claims 1-4, characterized in that, The method further includes: In response to a modification operation on the target dubbing file among the plurality of dubbing files, the modified target dubbing file is displayed in the dubbing file interface; The step of generating a dubbed multimedia work based on the plurality of dubbing files in response to a trigger operation on the second generation control includes: In response to a triggering operation on the second generation control, a dubbed multimedia work is generated based on the dubbing files other than the target dubbing file in the plurality of dubbing files, and the modified target dubbing file.

7. The method according to claim 6, characterized in that, The step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files includes: In response to a modification operation on a target dubbing object corresponding to a target dubbing file in the plurality of dubbing files, the modified target dubbing file is displayed in the dubbing file interface; and / or, In response to a modification operation on the target dubbing subtext corresponding to the target dubbing file in the plurality of dubbing files, the target dubbing file after modification of the target dubbing subtext is displayed in the dubbing file interface.

8. The method according to claim 6, characterized in that, The dubbing file interface also includes a batch control; before displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on a target dubbing file among the multiple dubbing files, the method further includes: In response to the triggering operation of the batch selection control, the target dubbing file is selected in the dubbing file interface; The step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files includes: In response to a modification operation on the target dubbing file that is selected among the plurality of dubbing files, the modified target dubbing file is displayed in the dubbing file interface.

9. The method according to claim 7, characterized in that, The dubbing file interface also includes a control for modifying the dubbing object; the step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing object corresponding to the target dubbing file in the plurality of dubbing files includes: In response to a trigger operation on the modified voice-over object control, an add voice-over object interface is displayed, which includes multiple voice-over object identifiers; In response to the selection operation of the first target dubbing object identifier, the first target dubbing object identifier is selected among the multiple dubbing object identifiers included in the add dubbing object interface, and the add dubbing object interface includes a first completion control; In response to the triggering operation of the first completion control, the target dubbing object corresponding to the target dubbing file is modified to the first target dubbing object corresponding to the first target dubbing object identifier, and the target dubbing file with the target dubbing object identifier updated to the first target dubbing object identifier is displayed in the dubbing file interface.

10. The method according to claim 7, characterized in that, In response to a modification operation on the target dubbing subtext corresponding to the target dubbing file in the plurality of dubbing files, displaying the target dubbing file after modification of the target dubbing subtext in the dubbing file interface includes: In response to a trigger operation on the first dubbing sub-text corresponding to the first dubbing file, the first dubbing sub-text is in an edit pre-selection state in the dubbing file interface, and the target dubbing file includes the first dubbing file; In response to a trigger operation on the first dubbing sub-text that is in the edit pre-selected state, a text modification interface is displayed, the text modification interface including the first dubbing sub-text in an editable state; In response to an editing operation on the first voice-over sub-text in the editable state, the first voice-over sub-text after the editing operation is displayed in the modified text interface, the modified text interface including a second completion control; In response to a triggering operation on the second completion control, the first dubbing file with the first dubbing subtext updated to the first dubbing subtext after the editing operation is displayed in the dubbing file interface.

11. The method according to claim 6, characterized in that, The step of displaying the modified target dubbing file in the dubbing file interface in response to a modification operation on the target dubbing file among the plurality of dubbing files includes: In response to the adjustment operation of the playback parameters of the target dubbing file in the plurality of dubbing files, the target dubbing file after the adjustment operation is displayed in the dubbing file interface, wherein the playback parameters include at least one of volume and speech rate.

12. The method according to any one of claims 1-4, characterized in that, Before displaying the voiceover text interface, the method further includes: Display a first multimedia interface, the first multimedia interface including the multimedia work without subtitles and a dubbing control; In response to a trigger operation on the dubbing control, a first text interface to be added is displayed; The interface for displaying the voice-over text includes: If the first text interface to be added is a text interface corresponding to multiple object dubbing, the dubbing text interface is displayed in response to the operation of adding text to the first text interface to be added; or... When the first text interface to be added is a text interface to be added corresponding to a single object dubbing, and the first text interface to be added includes a multi-object dubbing control, in response to the triggering operation of the multi-object dubbing control, the second text interface to be added is displayed; in response to the operation of adding text in the second text interface to be added, the dubbing text interface is displayed.

13. The method according to any one of claims 1-4, characterized in that, The interface for displaying the voice-over text includes: Display a second multimedia interface, which includes the multimedia work with added subtitles and a dubbing control; The step of displaying the dubbing file interface in response to a trigger operation on the first generation control includes: In response to a trigger operation on the dubbing control, an interface for adding dubbing objects is displayed, the interface for adding dubbing objects including a full-text dubbing control; In response to the triggering operation of the full-text dubbing control, the dubbing file interface is displayed according to the screen with added subtitles.

14. A dubbing device, characterized in that, include: The display module is configured to display a dubbing text interface in response to an operation of dubbing a multimedia work. The dubbing text interface includes a text to be dubbed for dubbing the multimedia work and a first generation control. In response to a trigger operation of the first generation control, the display module is configured to display a dubbing file interface. The dubbing file interface includes multiple dubbing files and a second generation control. The multiple dubbing files are dubbing files generated by splitting and dubbing the text to be dubbed. The dubbing objects corresponding to the multiple dubbing files include at least two dubbing objects. The generation module is used to generate a dubbed multimedia work based on the plurality of dubbing files in response to a trigger operation of the second generation control.

15. A terminal device, characterized in that, include: The terminal device includes a memory, a processor, and a display, wherein the memory stores a computer program that can run on the processor, and the terminal device executes the program to implement the method of any one of claims 1-13.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-13.

17. A computer program product having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-13.

Citation Information

Patent Citations

  • Voice synthesis method and device, computer device and storage medium

    CN110459200A

  • Audio book generation method and device, equipment, storage medium and program product

    CN114783403A

  • Video dubbing method and related device, electronic equipment and storage medium

    CN117177024A

  • Method and device for generating video based on text

    CN119893165A