Multimedia data generation method, device, electronic device, medium, and program product

The multimedia data generation method addresses low-quality video generation by enabling user-spontaneous recording and editing, improving emotional depth and production efficiency.

JP7758435B2Active Publication Date: 2025-10-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023577718
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-10-28
Filing Date
2022-10-27
Publication Date
2025-10-22
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

Existing technologies generate low-quality video data from text, as they rely on machine conversion, lacking emotional depth and user engagement.

Method used

A multimedia data generation method that includes user-spontaneous recording of text-to-voice conversion, allowing for emotional audio generation and intuitive video-image matching, with options for editing and re-recording to enhance quality.

Benefits of technology

Improves the quality and emotional depth of generated multimedia data, enhancing user experience and efficiency in video production by allowing for user interaction and intuitive multimedia fragment editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007758435000001
    Figure 0007758435000001
  • Figure 0007758435000002
    Figure 0007758435000002
  • Figure 0007758435000003
    Figure 0007758435000003
Patent Text Reader

Abstract

The present application discloses a multimedia data generating method, device, electronic device, medium, and program product for use in the field of multimedia data processing technology. The method includes receiving text information input by a user. In response to a recording trigger operation for the text information, displaying the text information and collecting a first reading voice according to the text information. Based on the text information and the first reading voice, first multimedia data is generated and presented, where the first multimedia data includes the first reading voice and a video image matching the text information. The first multimedia data includes a plurality of first multimedia fragments, and the plurality of first multimedia fragments respectively correspond to a plurality of text segments included in the text information. The present application can improve the quality of multimedia data generation.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application belongs to the field of data processing technology, and specifically relates to a multimedia data generating method, apparatus, electronic device, medium, and program product.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to a Chinese patent application filed with the State Intellectual Property Office of China on October 28, 2021, bearing application number 202111266196.5 and titled "Multimedia data generation method, device, electronic device, medium and program product," the entire contents of which are incorporated herein by reference. [Background technology]

[0003] With the development of the Internet, more and more users are creating videos to share content with other users. In related technologies, video data can be generated based on text edited by users, for example, text can be directly converted into audio by a machine, and video data can be generated based on the audio. However, the quality of the video generated by this related technology is relatively low. Summary of the Invention [Problem to be solved by the invention]

[0004] To solve the above technical problems or at least partially solve the above technical problems, the present application provides a multimedia data generating method, apparatus, electronic equipment, medium, and program product. [Means for solving the problem]

[0005] According to a first aspect of the present application, there is provided a method for generating multimedia data, the method comprising: receiving text information entered by a user; When responding to a recording trigger operation for the text information, displaying the text information and collecting a first reading voice based on the text information; generating and displaying (presenting) first multimedia data based on the text information and the first reading voice; The first multimedia data includes the first reading audio and a video image matching the text information, the first multimedia data includes a plurality of first multimedia fragments, each corresponding to a plurality of text segments included in the text information, the first target multimedia fragment includes a first target video fragment and a first target audio fragment, the first target multimedia fragment is a first multimedia fragment among the plurality of first multimedia fragments that corresponds to a first target text segment among the plurality of text segments, the first target video fragment includes a video image matching the first target text segment, and the first target audio fragment includes a reading audio of the first target text segment.

[0006] Optionally, the method further comprises: converting the text information into audio data in response to a multimedia composition operation; generating and presenting second multimedia data based on the text information and the audio data; Here, the second multimedia data includes a video image matching the audio data and the text information, the second multimedia data includes a plurality of second multimedia fragments, each corresponding to a plurality of text segments included in the text information, the second target multimedia fragment includes a second target video fragment and a second target audio fragment, the second target multimedia fragment is a second multimedia fragment among the plurality of second multimedia fragments that corresponds to a second target text segment among the plurality of text segments, the second target video fragment includes a video image matching the second target text segment, and the second target audio fragment includes a reading audio of the second target text segment.

[0007] Optionally, the method further comprises: After generating the second multimedia data, when responding to a recording trigger operation, displaying the text information and collecting a second reading voice according to the text information; generating and displaying third multimedia data based on the text information and the second reading voice, and overwriting the second multimedia data; The third multimedia data includes the second reading audio and a video image matching the text information, the third multimedia data includes a plurality of third multimedia fragments, each corresponding to a plurality of text segments included in the text information, the third target multimedia fragment includes a third target video fragment and a third target audio fragment, the third target multimedia fragment is a third multimedia fragment among the plurality of third multimedia fragments that corresponds to a third target text segment among the plurality of text segments, the third target video fragment includes a video image matching the third target text segment, and the third target audio fragment includes a reading audio of the third target text segment.

[0008] Optionally, the method further comprises: deleting the first target voice fragment in response to a re-recording operation for the first target voice fragment; displaying a first target text segment corresponding to the first target speech fragment and collecting a reading fragment of the first target text segment; The method further includes displaying the speech fragment in an area corresponding to the first target speech fragment.

[0009] Optionally, the method further comprises: The method further includes, when collecting the first reading speech, if it is detected that the matching rate between the first target speech fragment and the first target text segment is lower than a matching rate threshold, marking the first reading speech and the first target text segment.

[0010] Optionally, the method further comprises: The method further includes, in response to a swipe operation of an audio fragment, and when a first cursor indicating the first reading audio is swiped to the first target audio fragment, moving a second cursor indicating the text information to the first target text segment.

[0011] Optionally, the method further comprises: When collecting the first reading speech, the method further includes highlighting the text segment that the user is currently reading.

[0012] Optionally, the method further comprises: after generating and displaying the first multimedia data, modifying the text information in response to an editing operation on the text information to obtain modified target text information; In response to a recording trigger operation for the target text information, displaying the target text information and collecting a target reading voice according to the target text information; updating the first reading voice based on the target reading voice to obtain a third reading voice; generating and presenting fourth multimedia data based on the target text information and the third reading audio; The fourth multimedia data includes a video image matching the third reading audio and the text information, the fourth multimedia data includes a plurality of fourth multimedia fragments, each corresponding to a plurality of text segments included in the text information, the fourth target multimedia fragment includes a fourth target video fragment and a fourth target audio fragment, the fourth target multimedia fragment is a fourth multimedia fragment among the plurality of fourth multimedia fragments that corresponds to a fourth target text segment among the plurality of text segments, the fourth target video fragment includes a video image matching the fourth target text segment, and the fourth target audio fragment includes a reading audio of the fourth target text segment.

[0013] Optionally, the method further comprises: After collecting the first reading voice, the method further includes performing a sound change process and / or a speed change process on the first reading voice to obtain a fourth reading voice, generating and presenting first multimedia data based on the text information and the first reading voice, generating and presenting first multimedia data based on the text information and the fourth spoken audio.

[0014] According to a second aspect of the present application, there is provided a multimedia data generating apparatus, the apparatus comprising: a text information receiving module for receiving text information input by a user; a first reading voice collecting module for displaying the text information and collecting a first reading voice according to the text information in response to a recording trigger operation for the text information; a first multimedia data generation module for generating and presenting first multimedia data based on the text information and the first reading voice; The first multimedia data includes a video image matching the first reading audio and the text information, and the first multimedia data includes a plurality of first multimedia fragments, each corresponding to a plurality of text segments included in the text information, wherein the first target multimedia fragment includes a first target video fragment and a first target audio fragment, and the first target multimedia fragment is a multimedia fragment corresponding to a first target text segment among the plurality of text segments among the plurality of multimedia fragments, the first target video fragment includes a video image matching the first target text segment, and the first target audio fragment includes a reading audio of the first target text segment.

[0015] Optionally, the device comprises: an audio data conversion module for converting the text information into audio data in response to a multimedia composition operation; a second multimedia data generation module for generating and presenting second multimedia data based on the text information and the audio data; The second multimedia data includes a video image matching the audio data and the text information, the second multimedia data includes a plurality of second multimedia fragments, each corresponding to a plurality of text segments included in the text information, the second target multimedia fragment includes a second target video fragment and a second target audio fragment, the second target multimedia fragment is a multimedia fragment among the plurality of multimedia fragments that corresponds to a second target text segment among the plurality of text segments, the second target video fragment includes a video image matching the second target text segment, and the second target audio fragment includes a reading audio of the second target text segment.

[0016] Optionally, the device comprises: a second reading voice collecting module for displaying the text information and collecting a second reading voice according to the text information in response to a recording trigger operation after generating the second multimedia data; and a third multimedia data generating module for generating and displaying third multimedia data based on the text information and the second reading voice, and overwriting the second multimedia data; The third multimedia data includes the second reading audio and a video image matching the text information, the third multimedia data includes a plurality of third multimedia fragments, each corresponding to a plurality of text segments included in the text information, the third target multimedia fragment includes a third target video fragment and a third target audio fragment, the third target multimedia fragment is a multimedia fragment among the plurality of multimedia fragments that corresponds to a third target text segment among the plurality of text segments, the third target video fragment includes a video image matching the third target text segment, and the third target audio fragment includes a reading audio of the third target text segment.

[0017] Optionally, the device comprises: and a voice fragment re-recording module for use in, when responding to a re-recording operation for the first target voice fragment, deleting the first target voice fragment, displaying a first target text segment corresponding to the first target voice fragment, collecting reading fragments of the first target text segment, and displaying the reading fragments in an area corresponding to the first target voice fragment.

[0018] Optionally, the device comprises: The method further includes an error marking module for marking the first reading speech and the first target text segment when detecting that the matching rate between the first target speech fragment and the first target text segment is lower than a matching rate threshold when collecting the first reading speech.

[0019] Optionally, the device comprises: The device further includes an audio fragment swipe module for responding to an audio fragment swipe operation and for moving a second cursor indicating the text information to the first target text segment when a first cursor indicating the first spoken audio is swiped to the first target audio fragment.

[0020] Optionally, the device comprises: The system further includes a text segment highlighting module for highlighting a text segment currently being read by the user when collecting the first reading speech.

[0021] Optionally, the device comprises: a text information modifying module for modifying the text information in response to an editing operation on the text information after generating and displaying the first multimedia data, to obtain modified target text information; a target reading voice collection module for displaying the target text information in response to a recording trigger operation for the target text information and collecting a target reading voice according to the target text information; a third reading voice generation module for updating the first reading voice based on the target reading voice to obtain a third reading voice; a fourth multimedia data generation module for generating and presenting fourth multimedia data based on the target text information and the third reading voice; The fourth multimedia data includes a video image matching the third reading audio and the text information, the fourth multimedia data includes a plurality of fourth multimedia fragments, each corresponding to a plurality of text segments included in the text information, the fourth target multimedia fragment includes a fourth target video fragment and a fourth target audio fragment, the fourth target multimedia fragment is a multimedia fragment among the plurality of multimedia fragments that corresponds to a fourth target text segment among the plurality of text segments, the fourth target video fragment includes a video image matching the fourth target text segment, and the fourth target audio fragment includes a reading audio of the fourth target text segment.

[0022] Optionally, the device comprises: a voice processing module for performing a sound change process and / or a speed change process on the first reading voice after collecting the first reading voice to obtain a fourth reading voice; Specifically, the first multimedia data generating module is for generating and presenting first multimedia data based on the text information and the fourth reading voice.

[0023] According to a third aspect of the present application, there is provided an electronic device, the electronic device including a processor adapted to execute a computer program stored in a memory, the computer program being adapted to implement the method according to the first aspect when executed by the processor.

[0024] According to a fourth aspect of the present application, there is provided a computer readable storage medium having stored thereon a computer program which, when executed by a processor, performs the method according to the first aspect.

[0025] According to a fifth aspect of the present application there is provided a computer program product which, when run on a computer, causes the computer to carry out the method according to the first aspect. [Effects of the Invention]

[0026] The technical solutions according to the embodiments of the present application have the following advantages over the prior art:

[0027] After the user inputs the text information, a recording entry can be provided to the user, and the user can perform a recording trigger operation through this entry. In response to the recording trigger operation, the text information is displayed for the user to read aloud, and a first reading audio can be collected during the user's reading of the text information. First multimedia data can be generated and presented based on the text information and the first reading audio. The first multimedia data includes the first reading audio and a video image matching the text information, and the first multimedia data includes first multimedia fragments corresponding to multiple text segments in the text information. The present application can artificially record the first reading audio, and compared with machine-generated conversion of text information into audio, the artificially recorded first reading audio is more emotional. This improves the quality of the generated first multimedia data and the user's viewing experience. The first multimedia data is then displayed in the form of multiple first multimedia fragments, allowing the user to intuitively understand the correspondence between the text segments in the first multimedia fragments and the video images, thereby modifying a single multimedia fragment and improving the efficiency of video production and the user's experience. [Brief explanation of the drawings]

[0028] The drawings are used to provide a further understanding of the present invention, constitute a part of the specification, and are used to interpret the present invention together with the examples of the present invention, and are not intended to be limiting of the present invention. [Figure 1] 1 shows a schematic diagram of the system architecture of an exemplary application environment for a multimedia data generation method usable in an embodiment of the present application; [Figure 2] 1 is a flowchart of a multimedia data generating method according to an embodiment of the present application; [Figure 3] FIG. 2 is a schematic diagram of a text input interface in an embodiment of the present application. [Figure 4] FIG. 1 is a schematic diagram of a recording interface in an embodiment of the present application. [Figure 5] FIG. 2 is a schematic diagram of an interface for displaying multimedia data in an embodiment of the present application. [Figure 6] 10 is another flowchart of a multimedia data generating method according to an embodiment of the present application. [Figure 7] 10 is another flowchart of a multimedia data generating method according to an embodiment of the present application. [Figure 8] FIG. 10 is yet another schematic diagram of a recording interface in an embodiment of the present application. [Figure 9] FIG. 10 is yet another schematic diagram of a recording interface in an embodiment of the present application. [Figure 10] FIG. 10 is yet another schematic diagram of a recording interface in an embodiment of the present application. [Figure 11] 10 is another flowchart of a multimedia data generating method according to an embodiment of the present application. [Figure 12] 1 is a structural schematic diagram of a multimedia data generating device in an embodiment of the present application; [Figure 13] 1 is a structural schematic diagram of an electronic device according to an embodiment of the present application; DETAILED DESCRIPTION OF THE INVENTION

[0029] In order to make the above-mentioned objectives, features and advantages of the present application more clearly understood, the present application will be further described below. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0030] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application; however, it will be apparent that the present application may be implemented in other ways different from those described herein, and that the embodiments in the specification are merely some embodiments of the present application, but not all of the embodiments.

[0031] FIG. 1 shows a schematic diagram of the system architecture of an exemplary application environment for a multimedia data generation method usable in embodiments of the present application.

[0032] As shown in FIG. 1, system architecture 100 may include one or more of terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as a medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables. Terminal devices 101, 102, and 103 may be various electronic devices, including, but not limited to, desktop computers, portable computers, smart phones, and tablet computers. It should be understood that the number of terminal devices, networks, and servers in FIG. 1 is merely a rough guide. Any number of terminal devices, networks, and servers may be included as needed for implementation. For example, server 105 may be a server cluster consisting of multiple servers.

[0033] The multimedia data generation method provided by the embodiments of the present application is generally executed by an application installed on the terminal device 101, the terminal device 102, or the terminal device 103. Correspondingly, a multimedia data generation device may be installed in the terminal device 101, the terminal device 102, or the terminal device 103. For example, a user can input text information into an application installed on the terminal device 101, the terminal device 102, or the terminal device 103 and perform a recording operation. The application can display the text information for the user to read aloud. When the user reads the text information, the application can collect a first reading audio. The application uploads the text information and the first audio data to the server 105. Based on the text information, the server 105 retrieves video images matching the text information from a locally stored image library, synthesizes the video images and the first reading audio, obtains first multimedia data, and returns the first multimedia data to the application for presentation by the application. The user's spontaneous recording can make the generated first multimedia data more emotional, thereby improving the quality of the first multimedia data and the user's viewing experience.

[0034] Referring to FIG. 2, FIG. 2 is a flowchart of a multimedia data generating method in an embodiment of the present application, which is used in an application installed on a terminal device, and may include the following steps:

[0035] In step S210, text information entered by the user is received.

[0036] In an embodiment of the present application, a text input interface for editing text may be provided to the user. The user may customize the editing of text information according to the needs of video production, or may paste permitted links and extract text information from the links. When producing a video, there is generally a time limit, and accordingly, the number of characters in the text information is also limited, for example, not to exceed 2,000 characters. Therefore, during the process of the user inputting text information, it may be possible to check whether the number of characters has overrun, and if so, a character overrun pop-up window may be displayed to alert the user.

[0037]

[0033] Referring to Figure 3, Figure 3 is a schematic diagram of a text input interface in an embodiment of the present application, which includes a text input interface, a one-touch video generation button, and a spontaneous recording button. A user can edit text information through the text input interface. The one-touch video generation button and the spontaneous recording button are different ways of generating multimedia data, which will be introduced in detail below.

[0038] In step S220, in response to a recording trigger operation for the text information, the text information is displayed and a first reading voice based on the text information is collected.

[0039] In the embodiment of the present application, a user is provided with an entry for spontaneous recording, through which the user can create a video dubbed by himself. For example, the "spontaneous recording" button shown in Figure 3 is an entry for spontaneous recording. The user can click this "spontaneous recording" button to perform a recording trigger operation.

[0040] Alternatively, after the user performs a recording trigger operation, the application may directly display the text information, the user may read aloud based on the displayed text information, and the application may collect the first read-aloud voice of the text information, or the user may first enter the recording interface after performing a recording trigger operation, and then collect the first read-aloud voice of the text information after the user performs another trigger operation in the recording interface.

[0041] Referring to FIG. 4, FIG. 4 is a schematic diagram of a recording interface in an embodiment of the present application. The upper half of the recording interface is a text display area, and the lower half of the recording interface includes a record button. The text display area displays text information, allowing the user to easily record based on the text information. When the user first enters the recording interface, a guidance bubble may be presented to guide the user through recording. For example, the guidance bubble may display "Click the button below to start recording, and the subtitles will automatically scroll along with your reading," and the guidance bubble may automatically disappear after 5 seconds (s).

[0042] After the user performs a trigger operation (such as a click) on the record button, the recording function is turned on and the record button enters a recording state. Clicking the record button again stops the recording, and clicking the record button again resumes the recording. The user can recite the text information in the recording state, and the user can pause the recording as needed during the recording. After the user has completed the recitation, they can click the done button to obtain the first recitation audio.

[0043] In step S230, first multimedia data is generated and presented based on the text information and the first reading voice.

[0044] After recording and obtaining the first spoken audio, the application can send the first spoken audio and the text information to a server corresponding to the application. The server can select a video image matching the text information from a locally stored image library. The server can obtain first multimedia data by synthesizing the video image and the first spoken audio, and present the first multimedia data to the application. If the video generation fails, a pop-up window may be displayed stating "Video generation failed, please try again," and the user can click a retry button to regenerate the multimedia data.

[0045] It should be noted that the text information may be divided into multiple different text segments. For each text segment, an image matching the text segment may be selected. Therefore, the number of video images may be multiple. The more video images there are, the richer and more effective the generated first multimedia data content will be. Alternatively, the user may select a video image locally, and the application uploads the video image and the first reading audio together to the server. The server directly synthesizes the first multimedia data based on the video image and the first reading audio. During the synthesis, the video image and the first reading audio can be matched to improve the quality of the first multimedia data.

[0046] Here, the first multimedia data includes a first spoken audio and a video image matching the text information. That is, the first multimedia data is data including audio and images, and may be video data. As mentioned above, the text information may include multiple text segments. Therefore, the first multimedia data includes multiple first multimedia fragments, each corresponding to a multiple text segment included in the text information.

[0047] Accordingly, the first target multimedia fragment includes a first target video fragment and a first target audio fragment. The first target multimedia fragment is a first multimedia fragment corresponding to a first target text segment among a plurality of text segments among the plurality of first multimedia fragments. The first target video fragment includes a video image matching the first target text segment, and the first target audio fragment includes a spoken audio of the first target text segment. Here, the first target text segment may be any one text segment of the text information. Both the first target video fragment and the first target audio fragment correspond to the first target text segment.

[0048] Referring to FIG. 5, FIG. 5 is a schematic diagram of an interface for displaying multimedia data in an embodiment of the present application. The first multimedia data includes a plurality of first multimedia fragments, each corresponding to a different text segment, video image, and spoken audio. As can be seen, two multimedia fragments are displayed in the center of the interface, and the first of the two multimedia fragments is displayed enlarged at the top of the interface. The two multimedia fragments correspond to different video images, as well as different text segments and spoken audio. In this way, a user can intuitively view the correspondence between the text segments, video images, and spoken audio and modify one or more of the multimedia fragments as needed, thereby improving the efficiency of multimedia data generation.

[0049] In an embodiment of the multimedia data generation method of the present application, a user can input text information and then provide the user with a recording entry. The user can then trigger a recording operation via the entry. In response to the recording trigger operation, the text information is displayed for the user to read aloud, and a first reading audio can be collected during the user's reading of the text information. First multimedia data can be generated and presented based on the text information and the first reading audio. The first multimedia data includes the first reading audio and a video image matching the text information, and the first multimedia data includes first multimedia fragments corresponding to multiple text segments in the text information. The present application can artificially record the first reading audio, which is more emotional than machine-generated text information converted into audio. This improves the quality of the generated first multimedia data and the user's viewing experience. The first multimedia data is then displayed in the form of multiple first multimedia fragments, allowing the user to intuitively understand the correspondence between the text segments in the first multimedia fragments and the video images, thereby improving the efficiency of video production and the user's experience.

[0050] Referring to FIG. 6, FIG. 6 is another flowchart of a multimedia data generating method in an embodiment of the present application, which may include the following steps.

[0051] In step S610, text information entered by the user is received.

[0052] This step is the same as step S210 in the embodiment of FIG. 2, and for details, please refer to the description in the embodiment of FIG. 2, and no further description will be given here.

[0053] In step S620, if a multimedia composition operation is requested, the text information is converted into audio data, and second multimedia data is generated and presented based on the text information and the audio data.

[0054] In addition to supporting user-spontaneous recording, the present application also supports automatic dubbing, i.e., provides an entry for automatic dubbing. For example, the one-touch video generation button shown in FIG. 3 is an entry for automatic dubbing. A user can generate multimedia data with one touch by clicking the one-touch video generation button and performing a multimedia synthesis operation. For example, an application can send text information to a server, and the server can convert the text information into audio data using text-to-speech conversion technology. Similarly, a video image matching the text information is obtained, and the video image and audio data are synthesized to obtain second multimedia data. The second multimedia data is then sent to the application and presented to the application. As can be seen, this method is simple to operate and has a relatively high efficiency in generating multimedia data.

[0055] The second multimedia data may be similar to the first multimedia data and include audio data and a video image matching the text information. The second multimedia data may include a plurality of second multimedia fragments, each corresponding to a plurality of text segments included in the text information. The second target multimedia fragment may include a second target video fragment and a second target audio fragment, the second target multimedia fragment being a second multimedia fragment corresponding to a second target text segment among the plurality of text segments among the plurality of second multimedia fragments, the second target video fragment including a video image matching the second target text segment, and the second target audio fragment including audio of the second target text segment being read aloud.

[0056] In step S630, if a recording trigger operation is responded to, the text information is displayed and a second reading voice based on the text information is collected.

[0057] It should be noted that if the user is relatively satisfied with the second multimedia data, the user may store the second multimedia data in a local terminal device or share it on a social platform, etc. If the user is not satisfied with the second multimedia data, the user may voluntarily record it to improve the quality of the multimedia data.

[0058] The interface displaying the second multimedia data may include a spontaneous recording button, which has the same function as the spontaneous recording button in the embodiment of FIG. 2, and the user can click the spontaneous recording button to trigger the recording.

[0059] As shown in FIG. 5, the user may click the play button to play the second multimedia data. Alternatively, the user may click the spontaneous recording button to spontaneously record. After clicking the spontaneous recording button to trigger the recording, a confirmation pop-up window may first pop up, saying, "Using spontaneous recording will overwrite the existing synthesized voice. Do you want to continue recording?" If the user clicks continue, a second reading voice is collected. Because the volume, tone, and timbre may vary each time the same user recites the same text information, the second reading voice collected in this step may differ from the first reading voice.

[0060] In step S640, third multimedia data is generated based on the text information and the second reading voice, and is displayed to overwrite the second multimedia data.

[0061] Here, the third multimedia data includes a video image matching the second reading audio and the text information. The third multimedia data includes a plurality of third multimedia fragments, each corresponding to a plurality of text segments included in the text information. The third target multimedia fragment includes a third target video fragment and a third target audio fragment. The third target multimedia fragment is a third multimedia fragment corresponding to a third target text segment among the plurality of text segments among the plurality of third multimedia fragments, the third target video fragment includes a video image matching the third target text segment, and the third target audio fragment includes a reading audio of the third target text segment.

[0062] It should be noted that the method for generating the second and third multimedia data is the same as the method for generating the first multimedia data in the embodiment of Figure 2, and for details, please refer to the description in the embodiment of Figure 2, and no further description will be given here. Since the third multimedia data is generated by spontaneous recording based on the second multimedia data, the second multimedia data generated by one touch is overwritten by the third multimedia data.

[0063] In addition to spontaneously recording and directly generating the first multimedia data, the present application also allows for one-touch generation of second multimedia data, followed by spontaneous recording, and then updating the second multimedia data with spontaneously recorded third multimedia data, making the third multimedia data more emotional and improving the quality of the multimedia data.

[0064] Referring to FIG. 7, FIG. 7 is another flowchart of a multimedia data generating method in an embodiment of the present application, which may include the following steps:

[0065] In step S702, text information entered by the user is received.

[0066] This step is the same as step S210 in the embodiment of FIG. 2, and for details, please refer to the description in the embodiment of FIG. 2, and no further description will be given here.

[0067] In step S704, in response to a spontaneous recording operation, a recording interface including a recording button, a text display area, and an audio track area is displayed.

[0068] In an embodiment of the present application, the text information is divided into multiple text segments, each of which may be displayed on a separate line, with a new line being inserted when the text exceeds one line. Thus, the text information is segmented and displayed in the text display area. During the segmentation, a pop-up window may be displayed to indicate the segmentation progress, such as "Text processing is at xx%." This allows the user to view the text information more conveniently and intuitively during recording and avoid errors. The recording interface may further include an audio track area, which is used to display the reading audio already recorded by the user.

[0069] In step S706, in response to a trigger operation on the record button, a first reading voice based on the text information is collected and the first reading voice is displayed in the voice track area.

[0070] As described above, the text information in the text display area may be segmented and displayed according to each text segment. The user may recite each text segment in order, or the corresponding recitation audio may be displayed in the audio area each time a text segment is recited. Optionally, when collecting the first recitation audio, the text segment currently being recited by the user may be highlighted. For example, the text segment currently being recited by the user may be highlighted, and when it is detected that the recitation of the current text segment is complete, the currently being recited text segment may be scrolled up but may no longer be highlighted. To facilitate user reading during the entire recording, the currently being recited text segment may be maintained at the top of the text display area. For example, if the text display area is divided into four areas from top to bottom, the currently being recited text segment may be maintained in the first or second area.

[0071] Referring to Figure 8, Figure 8 is another schematic diagram of the recording interface in an embodiment of the present application. As can be seen, text information is segmented and displayed according to multiple different text segments, making it more convenient and intuitive for users when reading aloud. The text segment currently being read aloud may be highlighted. After reading is completed, the next unread text segment may be highlighted by scrolling up. Between the text display area and the record button is an audio track area, which displays multiple different reading voices.

[0072] The audio track area may include a play button, which is displayed when spoken audio is present and not yet played, and may be displayed near the end of the spoken audio indicated by the cursor. Clicking the play button starts playback of the spoken audio indicated by the cursor. During playback, the text information in the text display area may scroll as playback progresses.

[0073] If a flashback occurs during recording, the already recorded audio can be saved, and when re-entering the recording interface, a pop-up window may be displayed asking "There is an incomplete recording. Do you want to continue recording? Yes / No." If the user chooses to continue recording, the audio spoken before the flashback is loaded and the recording flow continues. If the user chooses to cancel, the audio spoken before the flashback is discarded. If no audio is identified during recording, a toast may be displayed saying, "No audio was identified. Please check your microphone or increase the volume."

[0074] In step S708, when collecting the first reading speech, if it is detected that the matching rate between the first target speech fragment and the first target text segment is lower than the matching rate threshold, the first reading speech and the first target text segment are marked.

[0075] During the user's reading, the accuracy of the user's reading may be detected. If it is detected that the matching rate between the collected first target speech fragment and the corresponding first target text segment is lower than a matching rate threshold (e.g., 85%, 90%, etc.), the first target speech fragment and the first target text segment are marked. For example, the first target text segment may be underlined, and when the user clicks on the first target text segment, a bubble may be displayed to warn the user that "the characters in this section do not correspond to the recording." At the same time, the first target speech fragment may be displayed in a dark red color, for example. Here, the first target text segment is any one text segment of the text information, and the first target speech fragment is the reading speech of the first target text segment.

[0076] Referring to Figure 9, Figure 9 is another schematic diagram of the recording interface in an embodiment of the present application, from which it can be seen that for the text segments where the user has read the text incorrectly, the mark is underlined, the reading voice corresponding to this text segment is displayed in a white background, and the other correct reading voices are displayed in a black background, so that the user can intuitively know the text segments where the reading is incorrect and re-record them.

[0077] In step S710, if a re-recording operation for the first target voice fragment is responded to, the first target voice fragment is deleted, the first target text segment corresponding to the first target voice fragment is displayed, a reading fragment of the first target text segment is collected, and the reading fragment is displayed in the area corresponding to the first target voice fragment.

[0078] If a user makes a reading error or is not satisfied with any of the reading audio, the user may re-record it. When the cursor points to the center of the track area, a record button may be presented with the message "Re-record this fragment." Referring to FIG. 10, FIG. 10 is another schematic diagram of the recording interface in an embodiment of the present application. The cursor is positioned in the center of the track area and points to the end of a reading audio. When the user clicks the record button, the reading audio can be deleted, leaving a gap in the area where the reading audio was located, and recording begins. When it is detected that the reading of the text segment is complete, the gap is filled with the generated reading fragment, and recording automatically ends.

[0079] In step S712, in response to the audio fragment swipe operation, and when the first cursor indicating the first reading audio is swiped to the first target audio fragment, the second cursor indicating the text information is moved to the first target text segment.

[0080] In an embodiment of the present application, when the spoken audio in the audio track area is dragged horizontally to the left or right, the text information in the text display area is scrolled simultaneously. For example, when the cursor in the audio track area is positioned on a first target audio fragment, the cursor in the text display area is positioned on a first target text fragment corresponding to the first target audio fragment.

[0081] In step S714, the first reading voice is subjected to sound modification processing and / or speed modification processing to obtain a fourth reading voice.

[0082] In an embodiment of the present application, after generating a spoken voice, a voice modulation button and a speed change button may be displayed on the recording interface. When recording is stopped, the voice modulation button can be used to perform voice modulation on the spoken voice, including various voice modulation types such as old man, boy, girl, and loli. The speed change button can be used to adjust the audio speed, including various speed changes such as 0.5X, 1X, 1.5X, and 2X, which the user can select according to actual needs. The voice modulation and speed change processes may be applied to the entire spoken voice or to only a part of the spoken voice.

[0083] In step S716, first multimedia data is generated and presented based on the text information and the fourth reading voice.

[0084] In this step, the method for generating the first multimedia data is the same as the method for generating the first multimedia data in the embodiment of Fig. 2, and specifically, please refer to the description in the embodiment of Fig. 2, and no further description will be given here. It should be noted that the first multimedia data generated in this step is different from the first multimedia data in the embodiment of Fig. 2 because the fourth reading voice is different from the first reading voice.

[0085] In the multimedia data generation method of the embodiment of the present application, text information in the text display area can be segmented and presented, and the user can record the text segments in order and display the reading audio of each text segment in the audio track area. The user can play the reading audio collected in the audio track area and re-record it if they are not satisfied with the reading audio. The accuracy of the user's reading can also be detected during the collection of reading audio. If there are many reading errors, the text containing the user's reading errors can be segmented, and the corresponding reading audio can be marked and presented to the user. The user can perform voice modification and / or speed modification on the collected reading audio according to actual needs. As can be seen, the present application provides users with a convenient and intuitive operating interface, improves the efficiency of video generation, and enhances the user experience.

[0086] Referring to FIG. 11, FIG. 11 is another flowchart of a multimedia data generating method in an embodiment of the present application, which may further include the following steps after step S230 based on the embodiment of FIG.

[0087] In step S1110, in response to an editing operation on the text information, the text information is corrected to obtain corrected target text information.

[0088] After generating the first multimedia data, the user can generate new multimedia data by modifying the text information. The interface displaying the first multimedia data may provide an entry for modifying the text information. The interface displaying the first multimedia data may refer to FIG. 6, and the interface may include a text edit button. When the user clicks the text edit button, a text input interface is displayed. The user can modify the text information displayed in the text input interface to obtain target text information, i.e., the text information modified by the user.

[0089] In step S1120, in response to a recording trigger operation for the target text information, the target text information is displayed, and a target reading voice based on the target text information is collected.

[0090] After the user modifies the text information, a pop-up window may be displayed to the user, asking, "The text content has been modified. Do you need to re-dub it?" If the user clicks "Yes," the recording interface will be entered, and the text display area will automatically be positioned to the text segment modified by the user, and the audio track area will be positioned to the corresponding reading audio. The user can recite the text segment modified by the user based on the position positioned for the user in the text display area and collect the corresponding reading audio, i.e., the target reading audio. In other words, the user only needs to record the modified text segment, thereby avoiding repeated recording and improving the efficiency of multimedia data updating.

[0091] In step S1130, the first reading voice is updated based on the target reading voice to obtain a third reading voice.

[0092] After the user recites each text segment corrected by the user, the pre-correction reading voice is automatically deleted, and the recollected reading voice is replaced with the position where the pre-correction reading voice is located, and finally the first reading voice can be updated to the third reading voice.

[0093] In step S1140, fourth multimedia data is generated and presented based on the target text information and the third reading voice.

[0094] Here, the fourth multimedia data includes a video image matching the third reading audio and the text information. The fourth multimedia data includes a plurality of fourth multimedia fragments, each corresponding to a plurality of text segments included in the text information. The fourth target multimedia fragment includes a fourth target video fragment and a fourth target audio fragment. The fourth target multimedia fragment is a fourth multimedia fragment corresponding to a fourth target text segment among the plurality of text segments among the plurality of fourth multimedia fragments, the fourth target video fragment including a video image matching the fourth target text segment, and the fourth target audio fragment including a reading audio of the fourth target text segment.

[0095] In the multimedia data generating method of the embodiment of the present application, after generating the multimedia data, the user can not only re-record the text information, but also re-edit the text information. After modifying the text information, the user only needs to re-record the text segment modified by the user, and the first multimedia data is updated based on the re-recorded reading voice, and finally new multimedia data is generated. This method can improve the efficiency of updating multimedia data.

[0096] Corresponding to the above method embodiment, the embodiment of the present application further provides a multimedia data generating device. Referring to FIG. 12, the multimedia data generating device 1200 includes: a text information receiving module 1210 for receiving text information input by a user; a first reading voice collection module 1220 for displaying the text information and collecting a first reading voice according to the text information in response to a recording trigger operation for the text information; a first multimedia data generation module 1230 for generating and presenting first multimedia data based on the text information and the first spoken audio; Here, the first multimedia data includes a first reading audio and a video image matching the text information, the first multimedia data includes a plurality of first multimedia fragments, each corresponding to a plurality of text segments included in the text information, and wherein the first target multimedia fragment includes a first target video fragment and a first target audio fragment, and the first target multimedia fragment is a first multimedia fragment among the plurality of first multimedia fragments that corresponds to a first target text segment among the plurality of text segments, the first target video fragment includes a video image matching the first target text segment, and the first target audio fragment includes a reading audio of the first target text segment.

[0097] Optionally, the multimedia data generating device 1200 a voice data conversion module for converting the text information into voice data in response to a multimedia composition operation; a second multimedia data generation module for generating and presenting second multimedia data based on the text information and the audio data; Here, the second multimedia data includes audio data and a video image matching the text information, the second multimedia data includes a plurality of second multimedia fragments, each corresponding to a plurality of text segments included in the text information, the second target multimedia fragment includes a second target video fragment and a second target audio fragment, the second target multimedia fragment is a second multimedia fragment corresponding to a second target text segment among the plurality of text segments among the plurality of second multimedia fragments, the second target video fragment includes a video image matching the second target text segment, and the second target audio fragment includes a reading audio of the second target text segment.

[0098] Optionally, the multimedia data generating device 1200 a second reading voice collecting module for displaying text information and collecting second reading voices according to the text information when responding to a recording trigger operation after generating the second multimedia data; a third multimedia data generating module for generating and displaying third multimedia data based on the text information and the second reading voice, and overwriting the second multimedia data; Here, the third multimedia data includes a video image matching the second reading audio and the text information, the third multimedia data includes a plurality of third multimedia fragments, each corresponding to a plurality of text segments included in the text information, the third target multimedia fragment includes a third target video fragment and a third target audio fragment, the third target multimedia fragment is a third multimedia fragment corresponding to a third target text segment among the plurality of text segments among the plurality of third multimedia fragments, the third target video fragment includes a video image matching the third target text segment, and the third target audio fragment includes a reading audio of the third target text segment.

[0099] Optionally, the multimedia data generating device 1200 and a speech fragment re-recording module for, in response to a re-recording operation for the first target speech fragment, deleting the first target speech fragment, displaying a first target text segment corresponding to the first target speech fragment, collecting a reading fragment of the first target text segment, and displaying the reading fragment in an area corresponding to the first target speech fragment.

[0100] Optionally, the multimedia data generating device 1200 The method further includes an error marking module for marking the first reading speech and the first target text segment when detecting that the matching rate between the first target speech fragment and the first target text segment is lower than a matching rate threshold when collecting the first reading speech.

[0101] Optionally, the multimedia data generating device 1200 The device further includes an audio fragment swipe module responsive to a swipe operation of the audio fragment and configured to move a second cursor indicating the text information to the first target text segment when the first cursor indicating the first spoken audio is swiped to the first target audio fragment.

[0102] Optionally, the multimedia data generating device 1200 The system further includes a text segment highlighting module for highlighting the text segment currently being read by the user when collecting the first reading speech.

[0103] Optionally, the multimedia data generating device 1200 a text information modifying module for modifying the text information in response to an editing operation on the text information after generating and displaying the first multimedia data to obtain modified target text information; a target reading voice collection module for displaying the target text information in response to a recording trigger operation on the target text information and collecting a target reading voice according to the target text information; a third reading speech generation module for updating the first reading speech based on the target reading speech to obtain a third reading speech; a fourth multimedia data generation module for generating and presenting fourth multimedia data based on the target text information and the third reading audio; Here, the fourth multimedia data includes a video image matching the third reading audio and the text information, the fourth multimedia data includes a plurality of fourth multimedia fragments, each corresponding to a plurality of text segments included in the text information, the fourth target multimedia fragment includes a fourth target video fragment and a fourth target audio fragment, the fourth target multimedia fragment is a fourth multimedia fragment corresponding to a fourth target text segment among the plurality of text segments among the plurality of fourth multimedia fragments, the fourth target video fragment includes a video image matching the fourth target text segment, and the fourth target audio fragment includes a reading audio of the fourth target text segment.

[0104] Optionally, the multimedia data generating device 1200 a speech processing module for, after collecting the first reading speech, performing a sound change process and / or a speed change process on the first reading speech to obtain a fourth reading speech; Specifically, the system further includes a first multimedia data generating module for generating and presenting first multimedia data based on the text information and the fourth reading voice.

[0105] The specific details of each module or unit in the above device have already been described in detail in a corresponding manner, so they will not be further described here.

[0106] It should be noted that although the above detailed description refers to several modules or units of an apparatus for performing operations, such classification is not mandatory. In fact, according to embodiments of the present application, the features and functions of two or more of the modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied in multiple modules or units.

[0107] An exemplary embodiment of the present application further provides an electronic device, the electronic device including a processor and a memory for storing processor-executable instructions, wherein the processor is configured to perform the multimedia data generation method described above in this exemplary embodiment.

[0108] 13 is a structural schematic diagram of an electronic device in an embodiment of the present application. It should be noted that the electronic device 1300 shown in FIG. 13 is merely an example and should not impose any limitations on the function and scope of use of the embodiment of the present application.

[0109] 13, electronic device 1300 includes a central processing unit (CPU) 1301, which can perform various appropriate operations and processes based on programs stored in read-only memory (ROM) 1302 or programs loaded from storage portion 1308 into random access memory (RAM) 1303. RAM 1303 further stores various programs and data necessary for system operation. Central processing unit 1301, ROM 1302, and RAM 1303 are connected to each other via bus 1304. Input / output (I / O) interface 1305 is also connected to bus 1304.

[0110] The following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including, for example, a network interface card such as a local area network (LAN) card or a modem. The communication section 1309 performs communication processing via a network, for example, the Internet. A drive 1310 may be connected to the I / O interface 1305 as needed. A removable medium 1311, for example, a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., may be mounted on the drive 1310 as needed to facilitate installation of a computer program read from the drive 1310 in the storage section 1308 as needed.

[0111] In particular, according to an embodiment of the present application, the processes described in the above-referenced flowcharts may be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program stored on a computer-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network via the communication portion 1309 and / or from removable media 1311. When the computer program is executed by the central processing unit 1301, it performs various functions specific to the device of the present application.

[0112] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, the computer program implementing the above multimedia data generating method when executed by a processor.

[0113] It should be noted that a computer-readable storage medium referred to in this application may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more conductors, a portable computer magnetic disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact magnetic disk read-only memory (CD-ROM), an optical memory device, a magnetic memory device, or any suitable combination of the above. In this application, a computer-readable storage medium may be any tangible medium that contains or stores a program, which may be used by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted over any suitable medium, including, but not limited to, wireless, electrical wire, optical cable, radio frequency, or the like, or any suitable combination of the above.

[0114] In an embodiment of the present application, a computer program product is further provided, which, when running on a computer, causes the computer to perform the above multimedia data generating method.

[0115] It should be explained that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another and do not necessarily require or imply the existence of any such actual relationship or ordering between those entities or operations. Furthermore, the terms "comprise," "include," "includes," or any other variation thereof, are intended to override the non-exclusive "comprise," whereby a process, method, article, or apparatus that includes a set of elements not only includes those elements, but also other elements not expressly listed, or further includes elements inherent in such process, method, article, or apparatus. Absent further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0116] The foregoing are merely specific embodiments of the present application that will enable those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not intended to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating multimedia data by computer software, comprising: receiving text information entered by a user; When responding to a recording trigger operation for the text information, displaying the text information and collecting a first reading voice based on the text information; generating and presenting first multimedia data based on the text information and the first spoken audio; the first multimedia data includes a video image that matches the first spoken voice and the text information; the first multimedia data includes a plurality of first multimedia fragments; the plurality of first multimedia fragments respectively correspond to a plurality of text segments included in the text information; the first target multimedia fragment includes a first target video fragment and a first target audio fragment; the first target multimedia fragment is a first multimedia fragment among the plurality of first multimedia fragments that corresponds to a first target text segment among the plurality of text segments; the first target video fragment includes a video image that matches the first target text segment; the first target audio fragment comprises a reading of the first target text segment; and further comprising: marking the first reading speech and the first target text segment when detecting that the matching rate between the first target speech fragment and the first target text segment is lower than a matching rate threshold when collecting the first reading speech. A method characterized by:

2. The method further comprises: converting the text information into audio data in response to a multimedia composition operation; generating and presenting second multimedia data based on the text information and the audio data; the second multimedia data includes video images that match the audio data and the text information; the second multimedia data includes a plurality of second multimedia fragments; the second multimedia fragments respectively correspond to a plurality of text segments included in the text information; the second target multimedia fragment includes a second target video fragment and a second target audio fragment; the second target multimedia fragment is a second multimedia fragment among the plurality of second multimedia fragments that corresponds to a second target text segment among the plurality of text segments; the second target video fragment includes a video image that matches the second target text segment; the second target audio fragment comprises a reading of the second target text segment; 2. The method of claim 1.

3. The method further comprises: After generating the second multimedia data, when responding to a recording trigger operation, displaying the text information and collecting a second reading voice according to the text information; generating and displaying third multimedia data based on the text information and the second reading voice, and overwriting the second multimedia data; the third multimedia data includes a video image that matches the second spoken voice and the text information; the third multimedia data includes a plurality of third multimedia fragments; the third multimedia fragments respectively correspond to a plurality of text segments included in the text information; the third target multimedia fragment includes a third target video fragment and a third target audio fragment; the third target multimedia fragment is a third multimedia fragment among the plurality of third multimedia fragments that corresponds to a third target text segment among the plurality of text segments; the third target video fragment includes a video image that matches the third target text segment; the third target audio fragment comprises a reading of the third target text segment. The method of claim 2.

4. The method further comprises: receiving a re-recording operation for the first target voice fragment from the user; deleting the first target voice fragment in response to the re-recording operation for the first target voice fragment; displaying a first target text segment corresponding to the first target speech fragment and collecting a reading fragment of the first target text segment; and displaying the speech fragment in an area corresponding to the first target speech fragment.

2. The method of claim 1.

5. The method further comprises: receiving a swipe operation of the user on the audio fragment, and swiping a first cursor indicating the first spoken audio to the first target audio fragment; in response to a swipe operation of an audio fragment, and when a first cursor indicating the first spoken audio is swiped to the first target audio fragment, moving a second cursor indicating the text information to the first target text segment.

2. The method of claim 1.

6. The method further comprises: highlighting the text segment currently being read by the user when collecting the first reading audio.

2. The method of claim 1.

7. The method further comprises: after generating and displaying the first multimedia data, modifying the text information in response to an editing operation on the text information to obtain modified target text information; In response to a recording trigger operation for the target text information, displaying the target text information and collecting a target reading voice according to the target text information; updating the first reading voice based on the target reading voice to obtain a third reading voice; generating and presenting fourth multimedia data based on the target text information and the third spoken audio; the fourth multimedia data includes a video image matching the third spoken voice and the text information; the fourth multimedia data includes a plurality of fourth multimedia fragments; the plurality of fourth multimedia fragments respectively correspond to a plurality of text segments included in the text information; the fourth target multimedia fragment includes a fourth target video fragment and a fourth target audio fragment; the fourth target multimedia fragment is a fourth multimedia fragment among the plurality of fourth multimedia fragments that corresponds to a fourth target text segment among the plurality of text segments; the fourth target video fragment includes a video image that matches the fourth target text segment; the fourth target audio fragment comprises a reading of the fourth target text segment; 2. The method of claim 1.

8. The method further comprises: After collecting the first reading voice, performing a sound change process and / or a speed change process on the first reading voice to obtain a fourth reading voice, generating and presenting first multimedia data based on the text information and the first reading voice, generating and presenting first multimedia data based on the text information and the fourth spoken audio; 2. The method of claim 1.

9. A multimedia data generating device, a text information receiving module for receiving text information input by a user; a first reading voice collecting module for displaying the text information and collecting a first reading voice according to the text information in response to a recording trigger operation for the text information; a first multimedia data generation module for generating and presenting first multimedia data based on the text information and the first reading voice; the first multimedia data includes a video image that matches the first spoken voice and the text information; the first multimedia data includes a plurality of first multimedia fragments; the plurality of first multimedia fragments respectively correspond to a plurality of text segments included in the text information; the first target multimedia fragment includes a first target video fragment and a first target audio fragment; the first target multimedia fragment is a multimedia fragment among the plurality of first multimedia fragments that corresponds to a first target text segment among the plurality of text segments; the first target video fragment includes a video image that matches the first target text segment; the first target audio fragment comprises a reading of the first target text segment; and further comprising: marking the first reading speech and the first target text segment when detecting that the matching rate between the first target speech fragment and the first target text segment is lower than a matching rate threshold when collecting the first reading speech. An apparatus characterized in that

10. 1. An electronic device including a processor, the processor is adapted to execute a computer program stored in a memory; The computer program, when executed by the processor, performs the method of any one of claims 1 to 8. An electronic device characterized by:

11. A computer-readable storage medium having a computer program stored thereon, The computer program, when executed by a processor, causes the computer to carry out the method of any one of claims 1 to 8. A computer-readable storage medium comprising:

12. A computer program comprising computer instructions, When executed by a processor, the method causes the computer to perform the method of any one of claims 1 to 8. A computer program comprising:

Citation Information

Patent Citations

  • Program transmission system and device thereof

    JP2002300434A

  • Apparatus, system, method and program for information delivery

    JP2003016093A

  • System and method for recording

    JP2007328849A

  • Reproduction device, method of controlling reproduction device and control program

    JP2009230468A

  • Electronic device and method of controlling the same

    US20190146834A1