Method, apparatus, device, and medium for processing multimedia data

By dividing multimedia data into editing tracks for text, video, and audio segments, the method addresses inflexible editing issues, improving the quality and flexibility of multimedia content creation.

JP7715918B2Active Publication Date: 2025-07-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024503680
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-09-27
Publication Date
2025-07-30
Estimated Expiration
2043-09-27

AI Technical Summary

Technical Problem

Existing multimedia data creation methods lack flexibility in meeting user-specific editing needs and result in low-quality multimedia data due to inflexible processing of text, audio, and video segments.

Method used

A method and apparatus for processing multimedia data by dividing it into multiple editing tracks for text, video image, and audio segments, allowing for independent and synchronized editing of these segments to meet diverse user needs and improve quality.

Benefits of technology

Enriches the editing capabilities of multimedia data by aligning and synchronizing text, audio, and video segments, thereby enhancing the overall quality and flexibility of multimedia content creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007715918000001
    Figure 0007715918000001
  • Figure 0007715918000002
    Figure 0007715918000002
  • Figure 0007715918000003
    Figure 0007715918000003
Patent Text Reader

Abstract

A method, device, equipment and medium for processing multimedia data, the method including the steps of receiving text information input by a user, and in response to a processing instruction for the text information, generating multimedia data based on the text information and displaying a multimedia editing interface for editing and manipulating the multimedia data, the multimedia data including a plurality of multimedia segments, the multimedia editing interface including a first editing track, a second editing track and a third editing track, the first track segment, the second track segment and the third track segment with which the timeline is aligned in the editing tracks respectively identify corresponding text segments, video image segments and audio segments. In the embodiment of the present disclosure, the editing tracks corresponding to the multimedia data can be enriched to meet the diversified editing needs of the multimedia data, and the quality of the multimedia data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims the priority of Chinese Patent Application No. 202211295639.8 filed on October 21, 2022, and the entire content disclosed in the above - mentioned Chinese patent application is incorporated herein by reference as part of this application.

[0002] [Technical Field] The present disclosure relates to a method, apparatus, device, and medium for processing multimedia data.

Background Art

[0003] With the development of computer technology, the methods of sharing knowledge and information are becoming increasingly diverse. In addition to carriers of text and audio information, carriers of video information can now be found everywhere.

[0004] In related technologies, based on the text content that a user wants to share, related image videos containing the text content are generated. However, the user's ideas change at any time, and the current creation style is not flexible enough to meet the user's fine - grained needs for flexible processing, and the quality of multimedia data is not high.

Summary of the Invention

Means for Solving the Problems

[0005] To solve the above - mentioned technical problems or at least part of the above - mentioned technical problems, the present disclosure provides a method, apparatus, device, and medium for processing multimedia data.

[0006] Embodiments of the present disclosure provide a method for processing multimedia data. The method includes receiving text information input by a user, and in response to a processing instruction for the text information, generating multimedia data based on the text information and displaying a multimedia editing interface for editing the multimedia data. The multimedia data includes a plurality of multimedia segments, and the plurality of multimedia segments respectively correspond to a plurality of text segments divided from the text information. The plurality of multimedia segments include a plurality of audio segments generated by respective readings corresponding to the plurality of text segments, and a plurality of video image segments respectively matching the plurality of text segments. The multimedia editing interface includes a first editing track, a second editing track, and a third editing track. The first editing track includes a plurality of first track segments, and the plurality of first track segments are respectively used to identify the plurality of text segments. The second editing track includes a plurality of second track segments, and the plurality of second track segments are respectively used to identify the plurality of video image segments. The third editing track includes a plurality of third track segments, and the plurality of third track segments are respectively used to identify the plurality of audio segments. The first track segment, the second track segment, and the third track segment whose timelines are aligned in the editing track respectively identify the corresponding text segment, video image segment, and audio segment.

[0007] Embodiments of the present disclosure further provide an apparatus for processing multimedia data. The apparatus includes a receiving module for receiving text information input by a user, a generating module for generating multimedia data based on the text information in response to a processing instruction for the text information, and a display module for displaying a multimedia editing interface for editing the multimedia data. The multimedia data includes a plurality of multimedia segments, and the plurality of multimedia segments respectively correspond to a plurality of text segments divided from the text information. The plurality of multimedia segments include a plurality of audio segments generated by reading aloud respectively corresponding to the plurality of text segments, and a plurality of video image segments respectively matching the plurality of text segments. The multimedia editing interface includes a first editing track, a second editing track, and a third editing track. The first editing track includes a plurality of first track segments, and the plurality of first track segments are respectively used to identify the plurality of text segments. The second editing track includes a plurality of second track segments, and the plurality of second track segments are respectively used to identify the plurality of video image segments. The third editing track includes a plurality of third track segments, and the plurality of third track segments are respectively used to identify the plurality of audio segments. The first track segment, the second track segment, and the third track segment in which the timelines are aligned in the editing track respectively identify the corresponding text segment, video image segment, and audio segment.

[0008] Embodiments of the present disclosure further provide an electronic device. The electronic device includes a processor and a memory for storing instructions executable by the processor. The processor is used to read the executable instructions from the memory and implement the method for processing multimedia data provided by the embodiments of the present disclosure by executing the instructions.

[0009] Embodiments of the present disclosure further provide a computer-readable storage medium, on which a computer program is stored, and the computer program is used to execute the method for processing multimedia data provided by the embodiments of the present disclosure.

[0010] Embodiments of the present disclosure further provide a computer program product, and when instructions of the computer program product are executed by a processor, the method for processing multimedia data provided by the embodiments of the present disclosure is realized.

Advantages of the Invention

[0011] The technical solutions provided by the embodiments of the present disclosure have the following advantages.

[0012] The multimedia data processing means provided by the embodiments of the present disclosure receives the text information input by the user, generates multimedia data based on the text information in response to the processing instructions for the text information, displays a multimedia editing interface for editing the multimedia data, the multimedia editing interface includes a plurality of multimedia segments, the plurality of multimedia segments respectively correspond to a plurality of text segments divided from the text information, the plurality of multimedia segments include a plurality of audio segments generated by corresponding readings of the plurality of text segments, and a plurality of video image segments respectively matching the plurality of text segments, the multimedia editing interface includes a first editing track, a second editing track, and a third editing track, the first track segment corresponding to the first editing track where the timelines are aligned in the editing track, the second track segment corresponding to the second editing track, and the third track segment corresponding to the third editing track respectively identify the corresponding text segment, video image segment, and audio segment. In the embodiments of the present disclosure, the editing tracks corresponding to the multimedia data can be enriched, the diverse editing needs of the multimedia data can be satisfied, and the quality of the multimedia data can be improved.

[0013] With reference to the following specific embodiments in combination with the drawings, the above and other features, advantages, and aspects of each embodiment of the present disclosure will become clearer. Throughout the drawings, the same or similar reference numerals indicate the same or similar elements. It should be understood that the drawings are illustrative and the members and elements are not necessarily drawn to actual scale.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Embodiments for Carrying Out the Invention

[0015] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the drawings. Although several embodiments of the present disclosure are shown in the drawings, as should be understood, the present disclosure can be realized in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, the provision of these embodiments is for the purpose of understanding the present disclosure in more detail and completely. As should be understood, the drawings and embodiments of the present disclosure are merely exemplary and not intended to limit the protection scope of the present disclosure.

[0016] As should be understood, each step described in the embodiments of the method of the present disclosure can be executed in a different order and / or in parallel. Also, the embodiments of the method may include additional steps and / or may omit the execution of the steps shown. The scope of the present disclosure is not limited in this regard.

[0017] As used herein, the term "including" and its variations are open inclusion, meaning "including, but not limited to...". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment", the term "another embodiment" means "at least one another embodiment", and the term "some embodiments" means "at least some embodiments". Related definitions of other terms are given in the following description.

[0018] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only for distinguishing different devices, modules or units, and are not for limiting the order or interdependence of the functions executed by these devices, modules or units.

[0019] It should be noted that the modifications of "one" and "a plurality" mentioned in the present disclosure are merely exemplary and not restrictive. As those skilled in the art should understand, unless specifically pointed out in the context, it should be understood as "one or a plurality".

[0020] The names of the messages or information exchanged between multiple devices according to the embodiments of the present disclosure are only used for the purpose of description, and are not intended to limit the scope of these messages or information.

[0021] To solve the above problems, embodiments of the present disclosure provide a method for processing multimedia data. In this method, the multimedia data is divided into a plurality of editing tracks such as a text editing track, a video image editing track, and an audio editing track, and the corresponding information is edited by the editing operations of the editing tracks, so as to satisfy the diversified editing needs of the multimedia data and improve the quality of the multimedia data.

[0022] Hereinafter, a method for processing multimedia data will be described with reference to specific embodiments.

[0023] FIG. 1 is a flowchart of a method for processing multimedia data provided by an embodiment of the present disclosure. The method can be executed by a multimedia data processing device, which can be implemented by software and / or hardware and may generally be integrated in an electronic device such as a computer. As shown in FIG. 1, the method includes the following steps 101 to 102.

[0024] Step 101: Receive the text information input by the user.

[0025] In one embodiment of the present disclosure, as shown in FIG. 2, a text input interface for editing text can be provided. The input interface includes a text area and a link area. In this embodiment, according to the video creation needs, text information can be customized and edited. That is, the text information input by the user in the text area of the text editing interface is received, or a permitted link is pasted, and text information and the like are extracted from the link. That is, the link information input by the user in the link area is received, the link information is identified, and the text information of the corresponding text is obtained and displayed in the text area for the user to edit. That is, the text information displayed in the text area can be edited multiple times.

[0026] When creating a video, usually there is a time length limit. Correspondingly, in some possible embodiments, there is also a certain limit on the number of characters of the text information. For example, it does not exceed 2000 characters. Therefore, it is possible to check whether the number of characters of the text in the text area exceeds the limit. If it exceeds the limit, a pop-up window indicating that the number of characters exceeds the limit can be displayed to alert the user.

[0027] In one embodiment of the present disclosure, continuing to refer to FIG. 2, the text input interface may include, in addition to the text area and the link area, a video generation button and a voice color selection entry control. In response to a trigger operation by the user on the voice color selection entry control, a candidate voice color menu can be displayed. The candidate voice color menu includes one or more candidate voice colors (the plurality of candidate voice colors may include multiple voice color types such as uncle, boy, girl, lolita, etc.), and a audition control corresponding to the candidate voice color. When the user triggers the audition control, a part or all of the text information input by the user is played back in the corresponding candidate voice color.

[0028] In this embodiment, when the user inputs text information, the user can select a voice color from a diversified candidate voice color, thereby determining a first target voice color based on the selection operation of the user on the candidate voice color menu. Further, a plurality of voice segments generated by reading a plurality of text segments obtained by splitting the text information based on the first target voice color are obtained. At this time, the voice color of each voice segment is the first target voice color, and the selection efficiency for the first target voice color is improved while satisfying the personalized selection for the first target voice color.

[0029] When splitting the text information into a plurality of voice segments, delimiter processing can be performed on the text information according to the reading habit of the first target voice color, and the text segments included in each delimiter can be determined, or the text information can be split into a plurality of voice segments based on the semantic information of the text information, and the text information can be converted into a plurality of voice segments corresponding to the plurality of text segments with the text segment as the conversion granularity.

[0030] Step 102: In response to the processing instruction for the text information, generate multimedia data based on the text information, and display a multimedia editing interface for editing the multimedia data. The multimedia data includes a plurality of multimedia segments, the plurality of multimedia segments respectively correspond to a plurality of text segments split from the text information, and the plurality of multimedia segments include a plurality of voice segments generated by reading corresponding to the plurality of text segments, and a plurality of video image segments respectively matching the plurality of text segments. The multimedia editing interface includes a first editing track, a second editing track, and a third editing track. The first editing track includes a plurality of first track segments, and the plurality of first track segments are each used to identify a plurality of text segments. The second editing track includes a plurality of second track segments, and the plurality of second track segments are each used to identify a plurality of video image segments. The third editing track includes a plurality of third track segments, and the plurality of third track segments are each used to identify a plurality of audio segments. The first track segment, the second track segment, and the third track segment in which the timelines are aligned in the editing track respectively identify the corresponding text segment, video image segment, and audio segment.

[0031] In one embodiment of the present disclosure, an entry for editing text information is provided, and a processing instruction for the text information is obtained by the entry. For example, the entry for editing text information may be a "video generation" control displayed on the text editing interface. When it is detected that the user triggers the "video generation" control of the text editing interface, a processing instruction for the text information is obtained. Of course, in other possible embodiments, the entry for editing text information may be a gesture operation input entry, an audio information input entry, etc.

[0032] In one embodiment of the present disclosure, in response to the processing instruction for the text information, multimedia data is generated based on the text information. The multimedia data includes a plurality of multimedia segments, and the plurality of multimedia segments respectively correspond to a plurality of text segments divided from the text information. The plurality of multimedia segments include a plurality of audio segments generated by reading aloud corresponding to the plurality of text segments respectively, and a plurality of video image segments respectively matching the plurality of text segments.

[0033] That is, as shown in FIG. 3, in this embodiment, the generated multimedia data is composed of a plurality of multimedia segments with segments as the granularity. Each multimedia segment includes at least three information types, namely a text segment, an audio segment (the initial timbre of the audio segment may be the first target timbre selected in the text editing interface in the above embodiment, etc.), a video image segment, etc. (the video image segment may include a video stream composed of consecutive pictures, and the consecutive pictures can correspond to the pictures in the video stream matched in a preset video library, and can also correspond to one or more pictures matched in a preset picture material library).

[0034] For example, as shown in FIG. 4, when the text information input by the user is "In today's society, there are various types of women's clothing, such as common Hanfu, fashion wear, sportswear, etc.", after processing the text information, four multimedia segments A, B, C, and D are generated. Multimedia segment A includes text segment A1 "In today's society, there are various types of women's clothing", audio segment A3 corresponding to text segment A1, and video image segment A2 that matches text segment A1. Multimedia segment B includes text segment B1 "such as common Hanfu", audio segment B3 corresponding to text segment B1, and video image segment B2 that matches text segment B1. Multimedia segment C includes text segment C1 "fashion wear", audio segment C3 corresponding to text segment C1, and video image segment C2 that matches text segment C1. Multimedia segment D includes text segment D1 "sportswear, etc.", audio segment D3 corresponding to text segment D1, and video image segment D2 that matches text segment D1.

[0035] As described above, in this embodiment, the multimedia data includes at least three information types. In order to meet the diverse editing needs for the multimedia data, in one embodiment of the present disclosure, a multimedia editing interface for editing the multimedia data is displayed. The multimedia editing interface includes a first editing track, a second editing track, and a third editing track. The first editing track includes a plurality of first track segments, and the plurality of first track segments are respectively used to identify a plurality of text segments. The second editing track includes a plurality of second track segments, and the plurality of second track segments are respectively used to identify a plurality of video image segments. The third editing track includes a plurality of third track segments, and the plurality of third track segments are respectively used to identify a plurality of audio segments. In order to make it easier to visually represent the plurality of types of information segments corresponding to each multimedia data, the segments, the second track segments, and the third track segments are displayed so that the timelines are aligned in the editing track. The first track segments, the second track segments, and the third track segments of the first track respectively identify the corresponding text segments, video image segments, and audio segments.

[0036] In an embodiment of the present disclosure, after the multimedia data is divided into a plurality of multimedia segments, each multimedia segment is divided into editing tracks corresponding to a plurality of information types. Thereby, the user can not only edit a single multimedia segment, but also edit a certain information segment corresponding to a certain editing track in a single multimedia segment, meeting the diverse editing needs of the user and ensuring the quality of the generated multimedia data.

[0037] Note that in different application scenarios, the display mode of the multimedia editing interface is different. As a possible implementation form, as shown in FIG. 5, the multimedia editing interface may include a video playback area, an editing area, and an editing track display area. In the editing track display area, a first track segment corresponding to the first editing track, a second track segment corresponding to the second editing track, and a third track segment corresponding to the third editing track are displayed so that the time lines are aligned.

[0038] The editing area is used to display editing function controls corresponding to the currently selected information segment (the specific editing function controls can be set according to the needs of the experimental scenario). The video playback area is used to display the image, text information, etc. at the current playback time of the multimedia data (a reference line corresponding to the current playback time can be displayed in the editing track display area in a direction perpendicular to the time line. The reference line is used to indicate the current playback position of the multimedia data, and the reference line can also be dragged. The video playback area synchronously displays the image, text information, etc. of the multimedia segment at the real-time position corresponding to the reference line, thereby facilitating the user to realize frame-by-frame browsing of the multimedia data based on the forward and backward movement of the reference line, etc.). The video playback area can further play the video playback control. When the video playback control is triggered, the current multimedia data is displayed, so that the user can visually know the playback effect of the current multimedia data.

[0039] Continuing with the scene shown in FIG. 4 as an example, in FIG. 5, in the editing track display area, four multimedia segments A, B, C, D and A1A2A3B1B2B3C1C2C3D1D2D3 corresponding to the four multimedia segments A, B, C, D are displayed so that the timelines are aligned. When the text segment A1 is selected, an editing interface corresponding to the text segment A1 is displayed in the editing area, and the editable A1 and editing controls such as font and character size are included in the editing interface. At this time, multimedia data and the like at the position corresponding to the reference line of the current multimedia data are displayed in the video playback area.

[0040] In addition to the above three editing tracks, the multimedia editing interface of this embodiment may further include other editing tracks. The number of other editing tracks is not displayed, and each other editing track can be used to display other-dimensional information segments corresponding to the multimedia data. For example, when an editing track for editing the background sound of the multimedia data is included in the other editing track, that is, the multimedia editing interface may include a fourth editing track for identifying the background audio data. Accordingly, in response to a trigger operation on the fourth editing track, the current background sound used by the fourth editing track is displayed in a preset background sound editing area (for example, the editing area described in the above embodiment), and alternative candidate background sounds are displayed. The alternative candidate background sounds can be displayed in any style such as labels in the background sound editing area. Based on the target background sound generated by the user modifying the current background sound based on the candidate background sounds in the background sound editing area, the fourth editing track updates and identifies the target background sound.

[0041] As described above, the method for processing multimedia data according to the embodiments of the present disclosure divides multimedia data into a plurality of multimedia segments, and each multimedia segment has an editing track corresponding to each information type based on the information type included therein. In the editing track, an information segment in a certain information type included in the multimedia segment can be edited and corrected, thereby enriching the editing track corresponding to the multimedia data, meeting the diversified editing needs of the multimedia data, and improving the quality of the multimedia data.

[0042] Hereinafter, with reference to specific embodiments, it will be described how to edit information segments of different information types of multimedia information segments corresponding to multimedia data.

[0043] In one embodiment of the present disclosure, the text segment corresponding to the multimedia information segment can be edited and corrected independently.

[0044] In this embodiment, as shown in FIG. 6, the step of independently editing and correcting the text segment corresponding to the multimedia information segment includes the following steps 601 to 602.

[0045] Step 601, in response to the user selecting a first target track segment in a first editing track, display the text segment currently identified in the first target track segment in a text editing area.

[0046] In one embodiment of the present disclosure, in response to a user selecting a first target track segment in a first editing track, the first target track segment may be one or multiple. Further, the text segment currently identified in the first target editing track segment is displayed in a text editing area, which may be located in the editing area described in the above embodiment. In addition to the text segment currently identified in the editable first target track segment, the text editing area may include other functional editing controls for the text segment, such as font editing controls, character size editing controls, and the like.

[0047] Step 602: Based on the target text segment generated by the user modifying the text segment currently displayed in the text editing area, update and identify the target text segment in the first target track segment.

[0048] In this embodiment, based on the target text segment generated by the user modifying the text segment currently displayed in the text editing area, update and identify the target text segment in the first target track segment.

[0049] For example, as shown in FIG. 7, if the text segment currently identified in the first target track segment selected by the user is "In today's society, there are various types of women's clothing" and the user modifies the text segment currently displayed in the text editing area to "In today's society, there are various types of women's clothing and a large number of them", the text segment in the first target track can be updated to "In today's society, there are various types of women's clothing and a large number of them", thereby realizing the modification of a single text segment and meeting the need for modifying a single text segment of multimedia data.

[0050] In one embodiment of the present disclosure, in order to further improve the quality of multimedia data and ensure the matching of images and texts, the images in the video image segment can be synchronously updated based on the modification to the text segment. In this embodiment, in response to a text update operation on the target text segment in the first target track segment, a second target track segment corresponding to the first target track segment is determined in the second editing track, a target video image segment that matches the target text segment is obtained, and the target video image segment is updated and identified in the second target track segment.

[0051] Semantic matching can be performed on the target text and the pictures in the preset picture material library to determine the corresponding target video image, and further, a target video segment can be generated based on the target video image. Alternatively, a video segment that matches the target text segment can be directly determined as the target video image segment from the preset video segment material library, etc., without limitation here.

[0052] In one embodiment of the present disclosure, in order to ensure the synchronization of text and audio, the audio segment can be synchronously corrected based on the modification to the text segment.

[0053] That is, in this embodiment, in response to a text update operation on the target text segment in the first target track segment, a third target track segment corresponding to the first target track segment is determined in the third editing track. The third track segment includes an audio segment corresponding to the text segment in the first target track segment. A target audio segment corresponding to the target text segment is obtained. For example, the target audio segment is obtained by reading aloud the target text segment, and the target audio segment is updated and identified in the third target track segment, thereby realizing the synchronous correction of audio and text.

[0054] Continuing to refer to FIG. 7, if the text segment currently identified in the first target track segment selected by the user is "In today's society, there are various types of women's clothing", and the user modifies the text segment currently displayed in the text editing area to "In today's society, there are various types of women's clothing and a large variety", the text segment in the first target track can be updated to "In today's society, there are various types of women's clothing and a large variety", and the third target track segment corresponding to the first target track segment is determined in the third editing track, and the audio segment in the third target track segment is updated from "In today's society, there are various types of women's clothing" to "In today's society, there are various types of women's clothing and a large variety".

[0055] As described above, in the process of editing and modifying the text segment, the corresponding time length on the time axis of the modified text segment is different from the corresponding time length on the time axis of the text segment before modification. Therefore, in different application scenarios, different display processes can be performed on the editing track in the editing track display area based on this change in time length.

[0056] In an embodiment of the present disclosure, when it is necessary to define the video image segment corresponding to the multimedia segment as the main information segment with an unchangeable time length according to the scene, in order to ensure that the time length of the video image segment is unchangeable, when it is known by detection that the first update time length corresponding to the target text segment in the first editing track does not match the time length corresponding to the text segment before modification, the second editing track is maintained without being changed, that is, it is ensured that the time length of the corresponding video image segment is not changed, and the first update track segment corresponding to the first update time length is displayed in the preset first candidate area.

[0057] Identify the target text segment in the first update track segment. The first candidate region may be located in other regions such as the upper region of the text segment before modification. Thus, even when it is known that the first update time length corresponding to the target text segment in the first edit track does not match the time length corresponding to the text segment before modification, the target text segment can be displayed in a form such as "track up", and not only is the time length of the video image segment corresponding to the second edit track not modified correspondingly, but also the time length corresponding to the text information segment of other multimedia segments is not visually affected.

[0058] When it is known by detection that the third update time length corresponding to the target audio segment in the third edit track does not match the time length corresponding to the audio segment before modification, maintain the second edit track without change, display the third update track segment corresponding to the third update time length in a preset second candidate region, and identify the target audio segment with the third update track segment. The second candidate region may be located in other regions such as the lower region of the audio segment before modification. Thus, even when it is known that the third update time length corresponding to the target audio segment in the third edit track does not match the time length corresponding to the audio segment before modification, the target audio segment can be displayed in a form such as "track down", and not only is the time length of the video image segment corresponding to the second edit track not modified correspondingly, but also the time length corresponding to the audio segment of other multimedia information segments is not visually affected.

[0059] For example, as shown in FIG. 8, taking the scene shown in FIG. 7 as an example, the modified text segment "In today's society, there are various types of women's clothing and a large number of kinds" is clearly longer in the corresponding time length compared to the text segment before modification "In today's society, there are various types of women's clothing". Therefore, the modified text segment "In today's society, there are various types of women's clothing and a large number of kinds" can be displayed above the text segment before modification, and the time length of the corresponding video image segment can be maintained without being changed.

[0060] In this embodiment, the modified audio segment "In today's society, there are various types of women's clothing and a large number of kinds" is clearly longer in the corresponding time length compared to the audio segment before modification "In today's society, there are various types of women's clothing". Therefore, the modified audio segment "In today's society, there are various types of women's clothing and a large number of kinds" can be displayed below the audio segment before modification, and the time length of the corresponding video image segment can be maintained without being changed, meeting the need that the time length of the video image segment cannot be changed in the corresponding scene.

[0061] In an embodiment of the present disclosure, when it is necessary to synchronize the video image segment corresponding to the multimedia segment with other information segments on the timeline according to the scene, in order to ensure the synchronization of the video image segment on the timeline, when it is known by detection that the first updated time length corresponding to the target text segment in the first editing track does not match the time length corresponding to the text segment before modification, the length of the first target track segment is adjusted based on the first updated time length, that is, the length of the first target track segment is scaled at the original display position. Similarly, when it is known by detection that the third updated time length corresponding to the target audio segment in the third editing track does not match the time length corresponding to the audio segment before modification, the length of the third target track segment is adjusted based on the third updated time length.

[0062] Furthermore, the length of the second target track segment corresponding to the first target track segment and the third target track segment in the second edit track is adjusted correspondingly, and the time axes of the adjusted first target track segment, the adjusted second target track segment, and the adjusted third target track segment are aligned, thereby realizing the alignment of all information segments included in the multimedia segment in the timeline.

[0063] For example, as shown in FIG. 9, taking the scene shown in FIG. 7 as an example, the corrected text segment "In today's society, there are various types of women's clothing and a large variety" is clearly longer in the corresponding time length compared to the uncorrected text segment "In today's society, there are various types of women's clothing". Therefore, the length of the first target track segment of the uncorrected text segment can be increased for display, and the corrected text segment "In today's society, there are various types of women's clothing and a large variety" is displayed in the adjusted first target track segment.

[0064] In this embodiment, the corrected audio segment "In today's society, there are various types of women's clothing and a large variety" is clearly longer in the corresponding time length compared to the uncorrected audio segment "In today's society, there are various types of women's clothing". Therefore, the length of the second target track segment of the uncorrected audio segment can be increased for display, and the corrected audio segment "In today's society, there are various types of women's clothing and a large variety" is displayed in the adjusted second target track segment.

[0065] In order to realize the synchronization between the video image segment and other information segments, in this embodiment, the length of the second target track segment corresponding to the first target track segment and the third target track segment in the second edit track is adjusted correspondingly, and the time axes of the adjusted first target track segment, the adjusted second target track segment, and the adjusted third target track segment are aligned.

[0066] In one embodiment of the present disclosure, the audio segment corresponding to the multimedia information segment can be edited and modified independently.

[0067] In this embodiment, as shown in FIG. 10, the step of independently editing and modifying the audio segment corresponding to the multimedia information segment includes the following steps 1001 to 1003.

[0068] Step 1001, in response to the user selecting the third target track segment on the third editing track, the third target track segment correspondingly identifies the audio segment corresponding to the text segment displayed in the first target track segment.

[0069] In one embodiment of the present disclosure, in response to the user selecting the third target track segment on the third editing track, there may be one or more third target track segments. The third target track segment correspondingly identifies the audio segment corresponding to the text segment displayed in the first target track segment. That is, in this embodiment, the audio segment can be edited independently.

[0070] Step 1002, in the preset audio editing area, display the current tone color used for the audio segment in the third target track segment and display alternative candidate tone colors.

[0071] The preset audio editing area of this embodiment may be located in the editing area described in the above embodiment. In the preset audio editing area, display the current tone color used for the audio segment in the third target track segment and display alternative candidate tone colors. As shown in FIG. 11, the candidate alternative tone colors can be displayed in any style such as label form. For example, labels of candidate alternative tone colors such as "uncle", "girl", "old person", etc. can be displayed, and the user can realize the selection of candidate tone colors by triggering the corresponding label.

[0072] Step 1003: Based on the second target timbre generated by the user modifying the current timbre based on the candidate timbres in the audio editing area, update and identify the target voice segment in the third target track segment. The target voice segment is a voice segment generated by reading the text segment identified in the first target track segment using the second target timbre.

[0073] In this embodiment, the user can modify the current timbre by triggering the candidate timbres, modify the current timbre to the triggered candidate timbre, that is, the second target timbre. Thereby, the timbre of the voice segment in the third track segment is modified, meeting the user's need to modify the timbre of a certain voice segment. For example, the user can modify multiple voice segments corresponding to the third track segment to different timbres, thereby realizing an interesting voice playback effect.

[0074] As described above, the multimedia data processing method of the embodiment of the present disclosure can flexibly and independently edit and modify the text segment, voice segment, etc. corresponding to the multimedia segment, further meeting the diversified editing needs of multimedia data and improving the quality of multimedia data.

[0075] To implement the above embodiment, the present disclosure further provides a multimedia data processing apparatus.

[0076] FIG. 12 is a structural schematic diagram of a multimedia data processing apparatus provided by an embodiment of the present disclosure. The apparatus can be implemented by software and / or hardware, and generally integrated into an electronic device to perform multimedia data processing. As shown in FIG. 12, the apparatus includes a receiving module 1210, a generating module 1220, and a display module 1230. The receiving module 1210 is used to receive the text information input by the user. The generation module 1220 is used to generate multimedia data based on text information in response to a processing instruction for the text information. The display module 1230 is used to display a multimedia editing interface for editing multimedia data. The multimedia data includes a plurality of multimedia segments, and the plurality of multimedia segments respectively correspond to a plurality of text segments split from the text information. The plurality of multimedia segments include a plurality of audio segments generated by respective readings corresponding to the plurality of text segments, and a plurality of video image segments respectively matching the plurality of text segments. The multimedia editing interface includes a first editing track, a second editing track, and a third editing track. The first editing track includes a plurality of first track segments, and the plurality of first track segments are respectively used to identify a plurality of text segments. The second editing track includes a plurality of second track segments, and the plurality of second track segments are respectively used to identify a plurality of video image segments. The third editing track includes a plurality of third track segments, and the plurality of third track segments are respectively used to identify a plurality of audio segments. The first track segment, the second track segment, and the third track segment whose timelines are aligned in the editing track respectively identify the corresponding text segment, video image segment, and audio segment.

[0077] Optionally, the receiving module specifically receives text information input by the user in the text area and / or receives link information input by the user in the link area, identifies the link information, obtains the text information of the corresponding page, and is used to display it in the text area for the user to edit.

[0078] Optionally, it further includes a first display module, a second display module, a tone color determination module, and an audio segment acquisition module. The first display module is used to display a tone color selection entry control. The second display module is used to display a candidate tone color menu in response to a trigger operation by the user on the tone color selection entry control. The candidate tone color menu includes candidate tone colors and audition controls corresponding to the candidate tone colors. The tone color determination module is used to determine a first target tone color based on a selection operation by the user on the candidate tone color menu. The audio segment acquisition module is used to acquire a plurality of audio segments generated by reading a plurality of text segments obtained by splitting text information based on the first target tone color.

[0079] Optionally, it further includes a third display module and a text segment editing module. The third display module is used to display the text segment currently identified in the first target track segment in a text editing area in response to the user selecting the first target track segment in the first editing track. The text segment editing module is used to update and identify the target text segment in the first target track segment based on the target text segment generated by the user modifying the text segment currently displayed in the text editing area.

[0080] Optionally, it further includes a track segment determination module and an audio segment acquisition module. The track segment determination module is used to determine a third target track segment corresponding to the first target track segment in the third editing track in response to a text update operation on the target text segment in the first target track segment. The voice segment acquisition module is used to acquire a target voice segment corresponding to a target text segment, and update and identify the target voice segment with a third target track segment.

[0081] Optionally, when it is known by detection that the first updated time length corresponding to the target text segment in the first editing track does not match the time length corresponding to the text segment before correction, the second editing track is maintained without being changed, and a first updated track segment corresponding to the first updated time length is displayed in a preset first candidate area, and the target text segment is identified with the first updated track segment, when it is known by detection that the third updated time length corresponding to the target voice segment in the third editing track does not match the time length corresponding to the voice segment before correction, the second editing track is maintained without being changed, and a third updated track segment corresponding to the third updated time length is displayed in a preset second candidate area, and the target voice segment is identified with the third updated track segment, and further includes a first time length display processing module used for this.

[0082] Optionally, when it is known by detection that the first updated time length corresponding to the target text segment in the first editing track does not match the time length corresponding to the text segment before correction, the length of the first target track segment is adjusted based on the first updated time length, when it is known by detection that the third updated time length corresponding to the target voice segment in the third editing track does not match the time length corresponding to the voice segment before correction, the length of the third target track segment is adjusted based on the third updated time length, The second time length display processing module is further included, which is used to correspondingly adjust the length of the second target track segment corresponding to the first target track segment and the third target track segment in the second editing track, and align the time axes of the adjusted first target track segment, the adjusted second target track segment, and the adjusted third target track segment.

[0083] Optionally, in response to a text update operation on the target text segment in the first target track segment, determine the second target track segment corresponding to the first target track segment in the second editing track, The video image update module is further included, which is used to obtain a target video image segment that matches the target text segment, and update and identify the target video image segment with the second target track segment.

[0084] Optionally, responding to the user's selection of the third target track segment in the third editing track, where the third target track segment correspondingly identifies an audio segment corresponding to the text segment displayed in the first target track segment, in the preset audio editing area, display the current tone color used for the audio segment in the third target track segment and display alternative candidate tone colors, The tone color update module is further included, which is used to update and identify the target audio segment with the third target track segment based on the second target tone color generated by the user modifying the current tone color based on the candidate tone colors in the audio editing area, where the target audio segment is an audio segment generated by reading the text segment identified in the first target track segment using the second target tone color.

[0085] Optionally, the multimedia editing interface includes a fourth editing track used to identify background audio data, A background sound display module used to display the current background sound used by the fourth editing track and alternative candidate background sounds in a preset background sound editing area in response to a trigger operation on the fourth editing track; A background sound update processing module used to update and identify the target background sound on the fourth editing track based on the target background sound generated by the user modifying the current background sound based on the candidate background sound in the background sound editing area. It further includes.

[0086] The multimedia data processing device provided by the embodiments of the present disclosure can execute the multimedia data processing method provided by any embodiment of the present disclosure, has corresponding functional modules and beneficial effects for executing the method, and will not be repeatedly described here.

[0087] To implement the above embodiments, the present disclosure further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the multimedia data processing method of the above embodiments is realized.

[0088] FIG. 13 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure.

[0089] Hereinafter, specifically referring to FIG. 13, it shows a structural schematic diagram of an electronic device 1300 suitable for implementing the embodiments of the present disclosure. The electronic device 1300 of the embodiments of the present disclosure may include mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers, but is not limited thereto. The electronic device shown in FIG. 13 is merely an example and should not impose any limitations on the functions and usage ranges of the embodiments of the present disclosure.

[0090] As shown in FIG. 13, the electronic device 1300 may include a processor (e.g., a central processor, a graphics processor, etc.) 1301, which can execute various appropriate operations and processes based on a program stored in a read-only memory (ROM) 1302 or a program loaded from a memory 1308 into a random access memory (RAM) 1303. Various programs and data necessary for the operation of the electronic device 1300 are further stored in the RAM 1303. The processor 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0091] Normally, an input device 1306 including, for example, a touch panel, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc., an output device 1307 including, for example, a liquid crystal display (LCD), a speaker, an oscillator, etc., a memory 1308 including, for example, a tape, a hard disk, etc., and a communication device 1309 may be connected to the I / O interface 1305. The communication device 1309 can permit the electronic device 1300 to perform wireless or wired communication with other devices to exchange data. Although FIG. 13 shows an electronic device 1300 having various devices, it should be understood that it is not required to implement or include all the shown devices. Alternatively, more or fewer devices may be implemented or included.

[0092] In particular, based on the embodiments of the present disclosure, the above-described process described with reference to the flowchart may be implemented as a computer software program. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program mounted on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network by the communication device 1309, or installed from the memory 1308, or installed from the ROM 1302. When the computer program is executed by the processor 1301, the above-described functions limited to the method for processing multimedia data of the embodiments of the present disclosure are executed.

[0093] Note that the computer-readable medium of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having one or more conductors, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program, and the program may be used by an instruction execution system, apparatus, or device, or in combination therewith. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which has computer-readable program code. Such a propagated data signal can take various forms including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may further be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can transmit, propagate, or transmit a program used by or in combination with an instruction execution system, apparatus, or device. The program code included in the computer-readable medium may be transmitted by any suitable medium including, but not limited to, wires, cables, RF (radio frequency), or any suitable combination of the above.

[0094] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be connected to digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), networks of networks (e.g., the Internet) and end-to-end networks (e.g., ad hoc end-to-end networks), and any network currently known or future-developed.

[0095] The computer-readable medium may be included in the electronic device, may exist alone, or may not be assembled into the electronic device.

[0096] One or more programs are loaded on the computer-readable medium, and when the one or more programs are executed by the electronic device, the electronic device Receive the text information input by the user, generate multimedia data based on the text information in response to the processing instruction for the text information, display a multimedia editing interface for editing the multimedia data, the multimedia editing interface includes a plurality of multimedia segments, the plurality of multimedia segments respectively correspond to a plurality of text segments divided from the text information, the plurality of multimedia segments include a plurality of audio segments generated by respective readings corresponding to the plurality of text segments, and a plurality of video image segments respectively matching the plurality of text segments, the multimedia editing interface includes a first editing track, a second editing track, and a third editing track, a first track segment corresponding to the first editing track in which timelines are aligned in the editing track, a second track segment corresponding to the second editing track, and a third track segment corresponding to the third editing track respectively identify the corresponding text segment, video image segment, and audio segment. In the embodiments of the present disclosure, the editing track corresponding to the multimedia data can be enriched, the diversified editing needs of the multimedia data can be satisfied, and the quality of the multimedia data can be improved.

[0097] An electronic device can create computer program code for performing the operations of the present disclosure by one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and further include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on a user computer, partially on a user computer, executed as an independent software package, partially executed on a user computer and partially executed on a remote computer, or executed entirely on a remote computer or server. In a situation involving a remote computer, the remote computer can be connected to the user computer via any network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, connected via the Internet using an Internet service provider).

[0098] The flowcharts and block diagrams of the drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of code, and the module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. Note that in some alternative embodiments, the functions represented in the boxes may occur in a different order than the representation in the drawings. For example, two consecutive boxes can actually be executed basically in parallel, but in some cases, they may be executed in the reverse order, which is determined by the functions involved. Note that each box of the block diagram and / or flowchart, and combinations of boxes of the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that executes the specified function or operation, or may be implemented by a combination of dedicated hardware and computer instructions.

[0099] The units according to the embodiments of the present disclosure may be implemented in software or in hardware. The name of the unit does not constitute a limitation on the unit itself in a certain situation.

[0100] The functions described above in this specification can be executed at least partially by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0101] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be either a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0102] The above description is only an explanation of the preferred embodiments of the present disclosure and the technical principles used. As should be understood by those skilled in the art, the scope of the disclosure related to the present disclosure is not limited to the technical solutions formed by specific combinations of the above technical features. At the same time, without departing from the idea of the above disclosure, other technical means formed by arbitrarily combining the above technical features or their equivalent features should be included. For example, there may be mentioned technical solutions formed by replacing the above features with technical features having similar functions (not limited thereto) disclosed in the present disclosure.

[0103] Also, although the operations have been described in a particular order, it should not be understood that these operations are required to be performed in the particular order or sequential order shown. In some cases, multitasking and parallel processing may be advantageous. Similarly, although the above description includes some specific implementation details, these should not be construed as limitations on the scope of the present disclosure. Features described in the context of a single embodiment may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately or in any suitable sub-combination in multiple embodiments.

[0104] Although the subject matter has been described in a specific language with respect to structural features and / or method logic operations, it should be understood that the subject matter defined by the appended claims is not necessarily limited to the specific features or operations described above. On the contrary, the specific features and operations described above are merely exemplary forms for implementing the claims.

Claims

1. A method for processing multimedia data, comprising: receiving text information input by a user; responding to a processing instruction for the text information, generating multimedia data based on the text information, and displaying a multimedia editing interface for editing the multimedia data; wherein the multimedia data includes a plurality of multimedia segments, the plurality of multimedia segments respectively correspond to a plurality of text segments split from the text information, and the plurality of multimedia segments include a plurality of audio segments generated by respective readings corresponding to the plurality of text segments, and a plurality of video image segments respectively matching the plurality of text segments; the multimedia editing interface includes a first editing track, a second editing track, and a third editing track, the first editing track includes a plurality of first track segments, the plurality of first track segments are respectively used for identifying the plurality of text segments, the second editing track includes a plurality of second track segments, the plurality of second track segments are respectively used for identifying the plurality of video image segments, the third editing track includes a plurality of third track segments, the plurality of third track segments are respectively used for identifying the plurality of audio segments, and the first track segment, the second track segment, and the third track segment whose timelines are aligned in the first editing track, the second editing track, and the third editing track respectively identify the corresponding text segment, video image segment, and audio segment; the method for processing multimedia data further comprises: when it is known by detection that the first update time length corresponding to the target text segment in the first editing track does not match the time length corresponding to the text segment before modification, maintaining the second editing track unchanged and displaying a first update track segment corresponding to the first update time length in a preset first candidate area, the step of identifying the target text segment with the first update track segment; When it is known by detection that the third updated time length corresponding to the target voice segment in the third editing track does not match the time length corresponding to the voice segment before correction, maintaining the second editing track without change, and displaying a third updated track segment corresponding to the third updated time length in a preset second candidate area, the method for processing multimedia data further includes: identifying the target voice segment with the third updated track segment.

2. The step of receiving the text information input by the user includes: receiving the text information input by the user in the text area, and / or receiving the link information input by the user in the link area, identifying the link information, obtaining the text information of the corresponding page, and displaying the text information in the text area for the user to edit. The method according to claim 1.

3. displaying a timbre selection entry control; responding to a trigger operation on the timbre selection entry control by the user, and displaying a candidate timbre menu, where the candidate timbre menu includes candidate timbres and audition controls corresponding to the candidate timbres; determining a first target timbre based on a selection operation on the candidate timbre menu by the user; obtaining a plurality of voice segments generated by reading a plurality of text segments obtained by splitting the text information based on the first target timbre. The method according to claim 1 further includes this step.

4. responding to the user's selection of a first target track segment in the first editing track, and displaying the text segment currently identified by the first target track segment in a text editing area; updating and identifying the target text segment with the first target track segment based on the target text segment generated by the user's correction of the currently identified text segment displayed in the text editing area. The method according to claim 1 further includes these steps.

5. In response to a text update operation on the target text segment in the first target track segment, determining a third target track segment corresponding to the first target track segment in the third edit track; The method according to claim 4, further comprising: obtaining the target audio segment corresponding to the target text segment, and updating and identifying the target audio segment with the third target track segment.

6. When it is known by detection that the first update time length corresponding to the target text segment in the first edit track does not match the time length corresponding to the text segment before correction, adjusting the length of the first target track segment based on the first update time length; When it is known by detection that the third update time length corresponding to the target audio segment in the third edit track does not match the time length corresponding to the audio segment before correction, adjusting the length of the third target track segment based on the third update time length; The method according to claim 5, further comprising: correspondingly adjusting the lengths of the second target track segments corresponding to the first target track segment and the third target track segment in the second edit track, and aligning the time axes of the adjusted first target track segment, the adjusted second target track segment, and the adjusted third target track segment.

7. In response to a text update operation on the target text segment in the first target track segment, determining a second target track segment corresponding to the first target track segment in the second edit track; The method according to claim 4, further comprising: obtaining a target video image segment matching the target text segment, and updating and identifying the target video image segment with the second target track segment.

8. Responding to the user selecting a third target track segment in the third edit track, wherein the third target track segment correspondingly identifies an audio segment corresponding to the text segment displayed in the first target track segment; In a preset audio editing area, steps of displaying the current timbre used for the voice segment in the third target track segment and displaying alternative candidate timbres; Based on a second target timbre generated by the user modifying the current timbre based on the candidate timbres in the audio editing area, updating and identifying a target voice segment in the third target track segment, where the target voice segment is a voice segment generated by reading the text segment identified in the first target track segment using the second target timbre. The method according to claim 1 further includes this step.

9. The multimedia editing interface Further includes a fourth editing track used to identify background audio data, In response to a trigger operation on the fourth editing track, displaying the current background audio used by the fourth editing track and alternative candidate background audio in a preset background audio editing area, Based on a target background audio generated by the user modifying the current background audio based on the candidate background audio in the background audio editing area, updating and identifying the target background audio on the fourth editing track. The method according to claim 1.

10. A multimedia data processing device, A receiving module for receiving text information input by the user, A generating module for generating multimedia data based on the text information in response to a processing instruction for the text information, A display module for displaying a multimedia editing interface for editing operations on the multimedia data, and includes The multimedia data includes a plurality of multimedia segments, the plurality of multimedia segments respectively correspond to a plurality of text segments divided from the text information, and the plurality of multimedia segments include a plurality of voice segments generated by reading corresponding to the plurality of text segments respectively, and a plurality of video image segments respectively matching the plurality of text segments. The multimedia editing interface includes a first editing track, a second editing track, and a third editing track. The first editing track includes a plurality of first track segments, and the plurality of first track segments are respectively used to identify the plurality of text segments. The second editing track includes a plurality of second track segments, and the plurality of second track segments are respectively used to identify the plurality of video image segments. The third editing track includes a plurality of third track segments, and the plurality of third track segments are respectively used to identify the plurality of audio segments. The first track segment, the second track segment, and the third track segment in which the timelines are aligned in the first editing track, the second editing track, and the third editing track respectively identify the corresponding text segment, video image segment, and audio segment. The multimedia data processing apparatus When it is known by detection that the first update time length corresponding to the target text segment in the first editing track does not match the time length corresponding to the text segment before correction, maintaining the second editing track unchanged and displaying a first update track segment corresponding to the first update time length in a preset first candidate area, and identifying the target text segment with the first update track segment. When it is known by detection that the third update time length corresponding to the target audio segment in the third editing track does not match the time length corresponding to the audio segment before correction, maintaining the second editing track unchanged and displaying a third update track segment corresponding to the third update time length in a preset second candidate area, and identifying the target audio segment with the third update track segment. A multimedia data processing apparatus further including a first time length display processing module used for the above.

11. An electronic device A processor and A memory arranged to store executable instructions, and The processor is an electronic device arranged to read the executable instructions from the memory and execute the executable instructions to implement the method for processing multimedia data according to any one of claims 1 to 9 above. **Claim 12** A computer-readable storage medium having a computer program stored therein, the computer program being used to execute the method for processing multimedia data according to any one of claims 1 to 9 above.

Citation Information

Patent Citations

  • Video generation method and device, computer equipment and storage medium

    CN114513706A

  • System, method, and program for content generation

    JP2005062420A

  • Voice-attached animation production / distribution service system

    JP2011082789A

  • Method and apparatus for locating video playing node, device and storage medium

    US20220044703A1