Media data processing method, apparatus, device, medium, and product
By introducing audio data editing items and media processing logic data into media data processing, the visualization and automated processing of audio data are realized, solving the problems of low efficiency, high difficulty and monotonous effect in media file production, and improving the production efficiency and presentation effect of media files.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-01-14
- Publication Date
- 2026-07-14
AI Technical Summary
Existing technologies for creating media files are characterized by low efficiency, high difficulty, narrow applicability, and monotonous presentation. This is mainly due to their reliance on image resources and limited use of audio data, resulting in high professional requirements and difficulty in meeting diverse user needs.
By displaying media data editing information, including audio data editing items, receiving audio data editing operations, determining audio data and its associated display text data, and generating target media files using pre-set media processing logic data, the system achieves visual editing and automated processing of audio data.
It simplifies the creation of media files, improves production efficiency, enriches the presentation of media files, supports flexible association and display of audio and text data, and meets diverse user needs.
Smart Images

Figure CN122387359A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer processing technology, and more particularly to a media data processing method, apparatus, device, medium, and product. Background Technology
[0002] With the rapid development of computer technology, users have placed higher demands on the presentation of media data. Adding media data to enhance its artistry and appeal has become an important technical approach. Media data can bring richer visual effects to media, making it more attractive and valuable for viewing.
[0003] In related technologies, media files mostly rely on image resources and image processing methods, using audio data less frequently, resulting in a relatively monotonous presentation and difficulty in meeting diverse user needs. Furthermore, methods of creating media files from audio data typically require professionals to manually write programs related to the audio data, leading to low production efficiency. This method also demands a high level of expertise from the media data creators, making it difficult to produce and limiting its applicability. Summary of the Invention
[0004] This disclosure provides a media data processing method, apparatus, device, medium, and product to reduce the difficulty of media data production, improve the efficiency of media data production, and enrich the presentation effects of media data.
[0005] In a first aspect, embodiments of this disclosure provide a media data processing method, the method comprising:
[0006] Display media data editing information, wherein the media data editing information includes audio data editing items;
[0007] The audio data editing item receives an audio data editing operation, determines the first audio data based on the audio data editing operation, and determines the first display data of the target display text associated with the first audio data.
[0008] In response to a file generation request, pre-set media processing logic data corresponding to the first audio data is obtained, and a target media file is generated based on the first audio data, the first display data, and the media processing logic data.
[0009] Secondly, embodiments of this disclosure also provide a media data processing apparatus, the apparatus comprising:
[0010] An editing information display module is used to display media data editing information, wherein the media data editing information includes audio data editing items;
[0011] The media data editing module is used to receive audio data editing operations through the audio data editing item, determine first audio data according to the audio data editing operations, and determine first display data of target display text associated with the first audio data;
[0012] The media file generation module is used to respond to a file generation request, obtain pre-set media processing logic data corresponding to the first audio data, and generate a target media file based on the first audio data, the first display data, and the media processing logic data.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the media data processing method as described in any of the embodiments of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the media data processing method as described in any of the embodiments of this disclosure.
[0018] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the media data processing method as described in any of the embodiments of this disclosure.
[0019] The technical solution of this disclosure embodiment realizes the visual presentation of media data editing information by displaying media data editing information. Since the media data editing information includes audio data editing items, it realizes the visual editing of audio data in the media data production scenario. Then, it receives audio data editing operations through the audio data editing items, determines the first audio data according to the audio data editing operations, and determines the first display data of the target display text associated with the first audio data. This enables flexible editing of audio data and can automatically determine the text data that matches the audio data. Then, in response to the file generation request, it obtains the pre-set media processing logic data corresponding to the first audio data, and generates the target media file according to the first audio data, the first display data, and the media processing logic data. This solves the technical problems of low efficiency, high difficulty, narrow applicability, and monotonous presentation effect of media files in related technologies. Media files can be automatically, quickly, and conveniently generated through simple interactive operations, reducing the difficulty of media file production, improving the efficiency of media file production, and supporting the association and display of audio data with its corresponding text data in the media file, enriching the presentation effect of the media file. Attached Figure Description
[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0021] Figure 1A A schematic flowchart illustrating a media data processing method provided in an embodiment of this disclosure;
[0022] Figure 1B A schematic diagram of an audio data editing interface for a media data processing method according to an embodiment of the present disclosure is provided.
[0023] Figure 1C A schematic diagram of an audio data editing interface for another media data processing method provided in this embodiment of the present disclosure;
[0024] Figure 2A A schematic flowchart of another media data processing method provided in this embodiment of the disclosure;
[0025] Figure 2B A schematic diagram of an editing interface for a media data processing method according to an embodiment of the present disclosure, comprising an editing item for displaying text data associated with audio data;
[0026] Figure 3AA schematic flowchart of another media data processing method provided in this embodiment of the disclosure;
[0027] Figure 3B A schematic diagram of an editing interface for a media data processing method provided in this embodiment of the present disclosure, comprising display information editing items for target interactive data corresponding to the target interactive data input when applying the target media file;
[0028] Figure 4 This is a schematic diagram of the structure of a media data processing apparatus provided in an embodiment of the present disclosure;
[0029] Figure 5 This is a schematic diagram of the structure of an electronic device for implementing an embodiment of the present disclosure. Detailed Implementation
[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0036] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0037] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0038] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0039] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0040] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0041] Figure 1A This is a flowchart illustrating a media data processing method provided in an embodiment of this disclosure. This embodiment is applicable to media file creation scenarios, particularly those involving the visual editing of audio data. The method can be executed by a media data processing device, which can be implemented in software and / or hardware, optionally through an electronic device such as a mobile terminal, PC, or server. Figure 1A As shown, the method in this embodiment may specifically include:
[0042] S110. Display media data editing information, wherein the media data editing information includes audio data editing items.
[0043] In this embodiment of the disclosure, media data editing information can be understood as information used to edit media files. Specifically, media data editing information can be all parameters and settings related to media data that the user can see and manipulate within application software with media data editing capabilities, such as media data creation applications (e.g., special effects creation applications). Simply put, media data editing information can be information within a media data editing application. For example, media data editing information can be interactive information displayed in the media data editing interface of a media data creation application. Media data editing information allows users to edit media files to achieve desired effects, such as video filters, transition effects, animation effects, and audio processing effects. Users can adjust the edited media data as needed based on the media data editing information to achieve the desired visual or auditory effects.
[0044] In media data editing information, the audio data editing item can be understood as an information item used to edit audio data. The audio data editing item can be presented as a component in the media data editing interface. The audio data editing item can be used to perform at least one of the following audio editing operations: audio data setting, audio data trimming, and audio data splicing. The audio data setting operation is used to set the audio data used to generate the media file. The audio data setting operation includes at least one of the following operations: audio data adding, audio data switching, audio data receiving, and audio data uploading. The audio data adding operation can be understood as an operation to add audio data. The audio data trimming operation is used to trim the audio data so that partial data of the audio data can be obtained as needed, realizing flexible selection of audio data. The audio data receiving operation can be understood as receiving audio data transmitted by a preset object. The audio data transmitted by the preset object can be audio data processed by the preset object, or audio data transmitted through the preset object. For example, the preset object can include at least one of the following objects: a preset component, a preset terminal, and a preset interface. The audio data uploading operation can be understood as an operation to upload audio data used to create the media file.
[0045] In media data editing scenarios, the media data editing information (including audio data editing items) can be presented to the user in the form of a graphical user interface. Users can edit media data parameters by dragging sliders, clicking buttons, entering values, or selecting presets. Furthermore, users can view and edit various media data parameters and settings related to the audio data through the media data editing information to create the desired media data effects.
[0046] As an optional technical solution in this disclosure embodiment, the display of media data editing information may specifically include: receiving a media data editing request and displaying media data editing information according to the media data editing request. Further, receiving a media data editing request and displaying media data editing information according to the media data editing request may include at least one of the following operations: receiving a control trigger operation on an application-enabled control of a media data editing application, and in response to the control trigger operation, displaying a media data editing interface of the media data editing application, the media data editing interface including media data editing information; receiving a media data editing request for a generated target media file and displaying media data editing information corresponding to the target media file; receiving a control trigger operation on a preset media data editing trigger control and displaying media data editing information; etc.
[0047] S120. Receive audio data editing operation through the audio data editing item, determine first audio data according to the audio data editing operation, and determine first display data of the target display text associated with the first audio data.
[0048] In this embodiment of the disclosure, audio data editing operations may include one or more of the following: audio data setting operations, audio data trimming operations, and audio data splicing operations. Therefore, the presentation format of the audio data editing item receiving the audio data editing operation can also be varied. For example, the audio data editing item may include at least one of the following: audio data selection item, audio data upload item, audio data trimming item, and audio data splicing item.
[0049] Furthermore, receiving audio data editing operations through the audio data editing item may specifically include, but is not limited to, at least one of the following operations: receiving an audio data selection operation for at least one segment of preset audio data through the audio data selection item; receiving an audio data upload operation through the audio data upload item; receiving an audio data trimming operation for at least one segment of preset audio data through the audio data trimming item; receiving an audio data splicing operation for at least two segments of preset audio data through the audio data splicing item; etc.
[0050] As an optional implementation of this disclosure, an audio data selection operation for at least one segment of preset audio data can be received through the audio data selection item, and the first audio data can be determined based on the audio data selection operation. Specifically, it may include: receiving a data viewing operation for preset audio data through the audio data selection control, displaying at least one segment of preset audio data; and, in response to the audio data selection operation for at least one segment of preset audio data, determining the selected audio data as the first audio data. Preset audio data can be understood as pre-set audio data that is available for user selection (optional audio data). This technical solution can conveniently and quickly determine the first audio data used to generate a media file by selecting preset audio data, simplifying the audio data editing operation and improving the efficiency of media file production.
[0051] As another optional implementation of this disclosure, an audio data upload operation can be received through the audio data upload item, and the first audio data can be determined based on the audio data upload operation. Exemplarily, receiving an audio data upload operation through the audio data upload item and determining the first audio data based on the audio data upload operation may include: dragging the audio data to be uploaded to the audio data upload area corresponding to the audio data upload item, and determining the audio data dragged into the audio data upload area as the first audio data; and / or, receiving an upload trigger operation for the audio data upload item, displaying at least one segment of uploadable audio data, receiving a selection operation for at least one segment of uploadable audio data, and determining the selected audio data as the first audio data; etc. It should be noted that the audio data uploaded through the audio data upload operation can be one or more segments. When the first audio data is one segment of audio data, the audio data most recently dragged into the audio data upload area can be determined as the first audio data, or the most recently selected audio data can be determined as the first audio data. Specifically, the audio data upload area can be configured to add a maximum of one audio data segment, and similarly, the selected audio data segment can be configured to be a maximum of one audio data segment. When the first audio data consists of multiple audio segments, the audio data upload area can be configured to add a maximum of a first number of audio data segments, and similarly, the selected audio data segment can be configured to be a second number. In this case, multiple audio segments (not exceeding the first number) dragged and dropped into the audio data upload area can be defined as the first audio data, or multiple selected audio segments (not exceeding the second number) can be defined as the first audio data. Using this technical solution, custom settings for the first audio data can be achieved, better meeting the differentiated editing needs of users, enriching the audio data used to generate media files, making media file editing more flexible, and improving the media file creation experience.
[0052] As another optional implementation of this disclosure, the audio data trimming item can receive an audio data trimming operation for at least one segment of preset audio data, thereby determining the trimmed audio data as the first audio data. Optionally, the audio data selection item can receive an audio data selection operation for at least one segment of preset audio data, display an audio data trimming item corresponding to the selected preset audio data; the audio data trimming item can receive an audio data trimming operation for the selected preset audio data, thereby determining the trimmed audio data as the first audio data. Alternatively, the audio data upload item can receive an audio data upload operation, obtain the uploaded audio data, and the audio data trimming item can receive an audio data trimming operation for the uploaded audio data, determining the trimmed audio data as the first audio data. This technical solution enables flexible trimming of audio data, allowing the audio data to achieve the desired effect. Moreover, since the trimmed audio data is usually smaller than the original audio data, storage space is saved. This makes this technical solution particularly suitable for scenarios where audio data needs to be stored and played on mobile devices or devices with limited storage space, thus expanding the applicability of the generated media files.
[0053] In another optional implementation of this disclosure, the audio data splicing item receives an audio data splicing operation for at least two preset audio data segments, and determines the spliced audio data as the first audio data. By adopting this technical solution, a media file can contain multiple audio data segments, enriching the presentation effect of the first audio data and thus enriching the presentation effect of the media file.
[0054] Understandably, in practical applications, one or more of the above audio data editing operations can be performed on the audio data to determine the first audio data. For example, first select the preset audio data and then trim it; or first trim multiple audio data segments and then splice them together, etc.
[0055] For ease of understanding, combined with Figure 1B and Figure 1C Here's an example illustrating audio data editing. In this example, the audio data editing options support viewing and selecting preset audio data, and further support cropping the selected preset audio data to obtain the first audio data. For example... Figure 1B As shown, after triggering the audio data editing control, multiple audio data segments from a pre-built target music library can be displayed, including audio data A, audio data B, audio data C, audio data D, audio data E, audio data F, audio data G, etc. When the mouse cursor or touch point moves to a specific audio data segment (e.g., ...), the display will show the audio data segments. Figure 1BWhen audio data A is located within the trigger area, audio data A is highlighted, and the corresponding selection control "Use" is displayed. Triggering "Use" displays the audio data trimming control (scissors graphic icon) and the deselection control "Cancel" corresponding to audio data A. Furthermore, in... Figure 1B When the audio data trimming control is triggered, a trimming editing area corresponding to audio data A can be displayed. Within this area, trimming editing information is shown, including graphical audio data, the trimming box, the audio playback time corresponding to the trimming box, the duration of the audio segment within the trimming box, and a trimming confirmation control "Done." Figure 1C As shown. Users can adjust the positions of the edges of the cropping box to crop audio segments of the desired length and content. After triggering the "Complete" confirmation control, audio data A can be cropped according to the position of the cropping box edges to obtain the first audio data. For ease of operation, during the adjustment of the cropping box, the audio playback time corresponding to the cropping box and the duration of the audio segment within the cropping box can be displayed inside the cropping box, such as 12 seconds already selected. In addition, when the cropping box is not adjusted, the duration of the audio segment within the cropping box can also be displayed outside the adjustment box (e.g., lower left), such as 12 seconds already selected.
[0056] In this embodiment of the disclosure, the target display text may be pre-set text data associated with the first audio data. The display time information of the target display text is associated with the playback time information of the audio data. For example, the audio data may be the accompaniment data of a song, and the target display text associated with the audio data may be the lyrics corresponding to the accompaniment data. The first audio data may be the accompaniment data of the entire song or the accompaniment data of a song segment. Assuming that the first audio data may be a segment of the accompaniment data of a song, then the target display text associated with the first audio data is the lyrics corresponding to the first audio data. In the application scenario of the target media file, the target display text is displayed in association with the first audio data. The first display data of the target display text can be understood as the visual presentation data of the target display text, including the display style data (first style data) and / or display position data of the target display text. The display style data may include, but is not limited to, at least one of the following: display shape (e.g., font), display size (e.g., font size), display color, display transparency, and display posture (e.g., tilt).
[0057] Optionally, the first display data of the target display text associated with the first audio data can be preset, random, or customized by the user through interactive operations. Therefore, determining the first display data of the target display text associated with the first audio data may include: obtaining preset first display data of the target display text associated with the first audio data; or, randomly obtaining a set of display data from preset display data as the first display data of the target display text associated with the first audio data; or, determining the first display data of the target display text associated with the first audio data according to the display data setting operation for the target display text, etc.
[0058] As an optional implementation of this disclosure, after determining the first display data of the target display text associated with the first audio data, the method further includes: displaying display effect preview data corresponding to the first display data. The advantage of this setting is that it allows for a direct understanding of whether the display effect of the target display text meets expectations, thereby ensuring the presentation effect of the target media file.
[0059] S130. In response to the file generation request, obtain the pre-set media processing logic data corresponding to the first audio data, and generate the target media file according to the first audio data, the first display data and the media processing logic data.
[0060] In this embodiment of the disclosure, a file generation request can be understood as a request to generate a media file. Exemplarily, a file generation request can be generated in response to a control triggering operation on a preset media data generation control. The media processing logic data can be understood as defining how, when the first audio data is played or triggered, the input media data to be processed is combined to generate file application effect data to display the media data effect of the target media file. The media processing logic data corresponding to different first audio data can be the same or different.
[0061] In this embodiment, the target media file is configured to be invoked by an application to present a desired effect in conjunction with the input media data to be processed. The target media file can be invoked in real-time while the media data to be processed is being captured, or it can be invoked during the post-processing of the data after it has been uploaded. After the target media file is invoked, it can generate a fusion display effect of first audio data, target display text, and the input media data to be processed. Taking, for example, the first audio data as the accompaniment data of a song, the target display text as the lyrics corresponding to the accompaniment data, and the input media data to be processed as second audio data, after the target media file is invoked, it can play the accompaniment data and display the lyrics corresponding to the accompaniment data, and record the accompaniment data, the lyrics, and the second audio data to generate a song performance video. If the media data to be processed also includes video data or image data, after the target media file is invoked, it can play the accompaniment data, display the lyrics corresponding to the accompaniment data, and record the accompaniment data, the lyrics, the video data or image data, and the second audio data to generate a song performance video.
[0062] Specifically, the media processing logic data can be configured to: control the playback and recording of the associated first audio data, the display of the target display text, and generate file application effect data based on the first audio data, the target display text, and the media data to be processed input when applying the target media file, when applying the target media file. For example, the media data to be processed input when applying the target media file may include at least one of the following: second audio data, input image data, animation sequence frames, and input video data. The second audio data includes at least audio data recorded by a preset recording device. The preset recording device may be a microphone or similar device provided by the terminal applying the target media file.
[0063] More specifically, the media processing logic data includes audio processing logic data for first audio data and second audio data from the media data to be processed input when applying the target media file, as well as text processing logic data for the target display text; the audio processing logic data is configured to play the first audio data and record the first audio data and the second audio data from the media data to be processed input when applying the target media file; the text processing logic data is configured to display the target display text. Using this technical solution, the first audio data, the second audio data, and the target display text can be automatically processed into file application effect data through the media processing logic data to display the media data effects of the target media file.
[0064] To facilitate understanding of the above scheme, taking the first audio data as accompaniment data and the target display text as lyrics as an example, the media processing logic data will be introduced. The media processing logic data includes the first audio data and the audio processing logic data of the second audio data in the media data to be processed input when applying the target media file. This can be configured to: play the first audio data in the target media file and record the first audio data into the media data video; acquire the second audio data recorded by the microphone, and generate target interactive data based on the second audio data and the first audio data, and record the target interactive data into the media data video. Furthermore, the audio processing logic data can also be configured to: play the third audio data in response to a switching playback operation on the first audio data.
[0065] In this embodiment of the disclosure, the text processing logic data can be further configured to: acquire preset text data corresponding to the first audio data; determine target text data based on the preset text data; and display target text based on the first audio data, the target text data, and the first display data. The target text data includes the target display text and display time data for each character in the target display text; the display time data includes the start display time, display duration, and end display time.
[0066] Optionally, determining the target text data based on the preset text data includes: determining the preset text data of the target format as the target text data; or, obtaining the preset text data of the first format and the preset text data of the second format, and determining the target text data based on the preset text data of the first format and the preset text data of the second format, wherein the recording method of the display time data of the characters recorded in the preset text data under the first format and the second format is different; etc.
[0067] As an optional implementation of this disclosure, determining the target text data based on the preset text data in the first format and the preset text data in the second format includes: converting the preset text data in the first format and the preset text data in the second format into reference text data in the target format, wherein the reference text data includes the target display text and the display time data of each character in the target display text; the display time data includes the start display time, the display duration, and the end display time; and cropping the reference text data according to the first audio playback data to obtain the target text data.
[0068] Continuing with the previous example, suppose the lyrics file corresponding to the accompaniment data of each song has two formats. In one format, the lyrics are recorded line by line, recording the content of each line of lyrics, the start display time (the offset time from the start time of the entire song's accompaniment when the line of lyrics begins to play), and the duration of display. In the other format, the lyrics are recorded character by character, recording the content of each character in the lyrics, the start display time of the corresponding line of lyrics, and the duration of display. The text processing logic data of the target display text included in the media processing logic data can be used to obtain the lyrics files of these two formats, parse out the lyrics file of the target format (i.e., the reference text data corresponding to the accompaniment data), and in the target format, the lyrics are recorded character by character, recording the content of each character in the lyrics, the start display time (the offset time from the start time of the entire song's accompaniment when the character begins to play), the duration of display, and the end display time (the offset time from the start time of the entire song's accompaniment when the character ends to play).
[0069] It should be noted that the continuous display time is used to indicate that the accompaniment data is playing the audio segment corresponding to the character, and does not limit the display of the character to this time period. For example, the text processing logic data can be used to control the character to be highlighted during its continuous display time when the accompaniment data is playing the audio segment corresponding to the character, based on the playback time data of the target format lyrics file and the accompaniment data. Specifically, when playing the first audio data, the start timestamp is recorded, and the difference between the current time and the start timestamp is calculated for each subsequent frame. The target format lyrics file is presented as a JSON array type, with the index starting from 0, and the lyrics information of the index is obtained. The start time of the current sentence of lyrics is compared with the difference of the current timestamp. If the difference of the current timestamp is greater than the start time of the current sentence of lyrics, it means that the current sentence of lyrics is being sung, and the next step is to calculate the actual progress of each sentence of lyrics. The information of each character in the lyrics is obtained in a loop using the time of a single character. The current timestamp is compared with each character to see if it is greater than the start display time of the character, and the timestamp minus the start display time of the character is less than the continuous display time of the character. If the above conditions are met, calculate the progress percentage of each character at the current timestamp, multiply it by the reciprocal of the sentence's character count, and finally add the index of the current character to the percentage of all characters in the sentence. This gives the overall progress of the current sentence.
[0070] When the first audio data (i.e., accompaniment data) is obtained by cropping, or when the first audio data is part of the accompaniment data of a song, the text processing logic data is configured to: perform time matching based on the time range of the first audio data cropped and the lyrics file of the target format; when the start time of each line of lyrics is within this range, it is extracted, and the extracted lyrics are used to form the target text data.
[0071] As an optional implementation of this disclosure, the media processing logic data may further include interactive processing logic data. Specifically, the interactive processing logic data may be configured to display target interactive data corresponding to the input media data to be processed when the target media file is applied. More specifically, the interactive processing logic data may be configured to: generate target interactive data corresponding to the input media data to be processed based on the first audio data and the input media data to be processed when the target media file is applied, and display the target interactive data. By adopting this technical solution, adding interactive logic to the target media file enables the presentation of target interactive data when the media data is applied, further enriching the media data presentation effect.
[0072] Continuing with the previous example, taking the target interactive data as the song's performance score and pitch representation data, the interactive processing logic data can be used as follows: after the song's accompaniment data begins playing, acquire the pitch of the second audio data input from the user's microphone and the current pitch of the song, execute a preset scoring algorithm, and obtain the performance score for each line and the total score for each frame of the algorithm result. Based on the above information, the song's performance score, as well as the pitch fluctuations of the second audio data and / or the first audio data, can be displayed during the song's playback.
[0073] The technical solution of this disclosure embodiment realizes the visual presentation of media data editing information by displaying media data editing information. Since the media data editing information includes audio data editing items, it realizes the visual editing of audio data in the media data production scenario. Then, it receives audio data editing operations through the audio data editing items, determines the first audio data according to the audio data editing operations, and determines the first display data of the target display text associated with the first audio data. This enables flexible editing of audio data and can automatically determine the text data that matches the audio data. Then, in response to the file generation request, it obtains the pre-set media processing logic data corresponding to the first audio data, and generates the target media file according to the first audio data, the first display data, and the media processing logic data. This solves the technical problems of low efficiency, high difficulty, narrow applicability, and monotonous presentation effect of media files in related technologies. Media files can be automatically, quickly, and conveniently generated through simple interactive operations, reducing the difficulty of media file production, improving the efficiency of media file production, and supporting the association and display of audio data with its corresponding text data in the media file, enriching the presentation effect of the media file.
[0074] Figure 2A This is a flowchart illustrating another media data processing method provided in this embodiment. Based on the above embodiments, this embodiment further refines the method for determining the first display data of the target display text associated with the first audio data. Optionally, the media data editing information further includes a first display editing item; further, determining the first display data of the target display text associated with the first audio data includes: receiving a first display editing operation through the first display editing item, and determining the first display data of the target display text associated with the first audio data based on the first display editing operation. For detailed implementation, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 2A As shown, the method in this embodiment may specifically include:
[0075] S210. Display media data editing information, wherein the media data editing information includes audio data editing items and a first display editing item.
[0076] In this embodiment of the disclosure, the first display editing item is used to edit the first display data of the target text data displayed in association with the first audio data. Optionally, the first display editing item includes a first style editing item. The first style editing item is used to set the display style data of the target display text displayed in association with the first audio data, that is, the first style data. The first style editing item includes template identifiers and / or style data setting items for multiple preset first style templates. The first style template can be understood as a pre-set template that gives the target display text pre-fetched visual and layout features, corresponding to multiple preset style data. The template identifier of the first style template is used to distinguish different first style templates and can be composed of graphic identifiers and / or text identifiers. The style data setting item can be understood as an interactive information item for editing multiple preset style data. One or more display style data of the target display text displayed in association with the first audio data can be set through interactive operations with the style data setting item.
[0077] S220. Receive audio data editing operation through the audio data editing item, and determine the first audio data according to the audio data editing operation.
[0078] S230: Receive a first display editing operation through the first display editing item, and determine the first display data of the target display text associated with the first audio data based on the first display editing operation.
[0079] In this embodiment of the disclosure, the first display editing operation can be understood as an interactive operation with the first display editing item, used to edit the first display data of the target display text displayed in association with the first audio data. As mentioned above, the first display data of the target display text includes display position data and / or first style data.
[0080] As an optional implementation of this disclosure, when the first style editing item includes a style data setting item, the step of receiving a first display editing operation through the first display editing item and determining the first display data of the target display text associated with the first audio data based on the first display editing operation includes: receiving a first display editing operation through the style data setting item and determining the first display data of the target display text associated with the first audio data based on the set display style data.
[0081] As another optional technical solution of this disclosure embodiment, when the first style editing item includes template identifiers of multiple preset first style templates, the step of receiving a first display editing operation through the first display editing item and determining the first display data of the target display text associated with the first audio data according to the first display editing operation includes: receiving a first trigger operation for the template identifier of the preset first style template, determining a first target template among the multiple first style templates according to the first trigger operation, and determining the first display data of the target display text associated with the first audio data according to the first target template.
[0082] Optionally, determining the first style data of the target text data to be displayed in association with the first audio data according to the first target template includes: determining the preset style data of the first target template as the first style parameter; and / or, displaying the preset style data of the first target template, receiving a data adjustment operation for at least one of the preset style data, and determining the first style parameter according to the preset style data and the data adjustment operation.
[0083] Optionally, after determining the first target template among the multiple first style templates based on the first triggering operation, the method further includes: displaying style effect preview data corresponding to the first target template. Using this technical solution, it is possible to intuitively and conveniently observe whether the style effect corresponding to the selected first target template meets expectations.
[0084] As an optional implementation of this disclosure, when the first style editing item includes a style data setting item and template identifiers of multiple preset first style templates, if no first trigger operation is received for the template identifier of the preset first style template and a first display editing operation is received through the style data setting item, the first display data of the target display text associated with the first audio data is determined according to the display style data set through the style data setting item.
[0085] As another optional implementation of this disclosure, when the first style editing item includes a style data setting item and template identifiers of multiple preset first style templates, after receiving a first trigger operation for the template identifier of the preset first style template and determining a first target template among the multiple first style templates based on the first trigger operation, the preset style data of the first target template can be displayed through the style data setting item. Further, determining the first style data of the target text data associated with the first audio data based on the first target template includes: receiving a data adjustment operation for the preset style data of the first target template through the style data setting item, and determining the first style data of the target text data associated with the first audio data based on the data adjustment operation and the preset style data. By adopting this technical solution, the display style data of the target text can be quickly set through the first style template, and the display style data of the first style template can be adjusted, making the setting method of the display style data of the target text more flexible and improving the editing experience of media files.
[0086] It is understood that the data adjustment operation for the preset style data of the first target template can be an adjustment of part of the preset style data of the first target template, or an adjustment of all the preset style data of the first target template. When part of the preset style data of the first target template is adjusted, the first style data of the target text data associated with the first audio data can be determined based on the unadjusted preset style data and the adjusted preset style data. Similarly, when part of the preset style data of the first target template is adjusted, the first style data of the target text data associated with the first audio data can be determined based on the adjusted preset style data.
[0087] To facilitate operation, after receiving the data adjustment operation of the preset style data for the first target template through the style data setting item, the method further includes: updating the style effect preview data according to the data adjustment operation. Using this technical solution, it is possible to intuitively and conveniently observe whether the style effect of the adjusted display style data meets expectations.
[0088] As another optional technical solution in this disclosure, when the first display editing item includes a display position editing item, the step of receiving a first display editing operation through the first display editing item and determining the first display data of the target display text associated with the first audio data based on the first display editing operation may specifically include: receiving a display position editing operation through the display position editing item and determining the first display data of the target display text associated with the first audio data based on the display position editing operation. By adopting this technical solution, the display position of the target display text can be flexibly adjusted, supporting diverse presentations of the target display text's display position to enrich the display effect of the target display text, thereby enriching the presentation effect of the target media file.
[0089] For example, the display position editing item includes a display position selection control and / or a position change control, etc. When the display position editing item includes a display position selection control, the step of receiving a display position editing operation through the display position editing item and determining the first display data of the target display text associated with the first audio data based on the display position editing operation may specifically include: receiving an editing trigger operation for the display position selection control, displaying multiple candidate position adjustment items, wherein the position adjustment items correspond to a preset display position; receiving an adjustment trigger operation for the position adjustment items, and adjusting the display position data of the target display text according to the position adjustment item corresponding to the preset display position.
[0090] For example, when the display position editing item includes a position transformation control, the step of receiving a display position editing operation through the display position editing item and determining the first display data of the target display text associated with the first audio data based on the display position editing operation may specifically include: receiving a display position editing operation through the display position transformation control, determining the display position data of the target display text associated with the first audio data before transformation, and determining the display position data of the target display text after transformation based on the display position data of the target display text before transformation and the preset transformation rule corresponding to the display position transformation control. For example, the preset transformation rule may include switching between multiple preset display positions according to a preset switching order; or moving the display position of the target display text according to a preset trajectory, etc.
[0091] like Figure 2BAs shown, the first display editing item corresponding to the target display text is displayed in the first parameter panel. The first display editing item includes an object information component, a display position component, and a text component. The object information component is used to display the object identification information corresponding to the edited object in the media data display, and is used to characterize the effective object of the parameters in the first parameter panel. The display position component is used to edit the display position data corresponding to the target display text. The text component includes a template selection component for selecting multiple text style templates, a font component, a color component, and an slant component. The template identifier includes graphic identifiers and text identifiers. The graphic identifier can display a preset image or a style effect image, etc. By triggering the graphic identifier, text identifier, or a preset area including graphic identifiers and text identifiers in the template identifier, the template identifier corresponding to the triggered template identifier (e.g., style template A) can be highlighted, and a style preview effect (not shown in the figure) corresponding to the triggered template identifier (e.g., style template A) can be displayed. The preset style parameters (including font, color, and slant) corresponding to the triggered template identifier (e.g., style template A) can be displayed in the font component, color component, and slant component. Furthermore, the display data in the font, color, and tilt components can be adjusted, such as switching fonts, changing colors, and adjusting the tilt degree. The audio data component displays a data identifier for the first audio data being edited, in order to determine the association between the edited first display data and the audio data, thus determining which audio data's target display text to render, which audio data to generate the target media file with, or in other words, which target media file to encapsulate it in.
[0092] S240. In response to the file generation request, obtain the pre-set media processing logic data corresponding to the first audio data, and generate the target media file based on the first audio data, the first display data and the media processing logic data.
[0093] The technical solution of this disclosure embodiment supports visual editing of the first display data of the target text data associated with the first audio data by adding a first display editing item to the media data editing information. Furthermore, by receiving a first display editing operation through the first display editing item and determining the first display data of the target display text associated with the first audio data according to the first display editing operation, flexible editing of the first display data can be achieved through a simple and convenient interactive method, so as to better meet the user's differentiated display needs for the target text data, enrich the effect of the target text data associated with the first audio data when the target media file is applied, and thus improve the presentation effect of the target media file.
[0094] Figure 3AThis is a flowchart illustrating another media data processing method provided in this embodiment. Based on the above embodiments, this embodiment adds a technical feature of customizing the display data of the target interactive data. Optionally, the media data editing information includes an interactive data editing item corresponding to the target interactive data; the interactive data editing item includes a second display editing item; after the display media data editing information, it further includes: receiving a second display editing operation through the second display editing item, and determining, based on the second display editing operation, second display data of the target interactive data corresponding to the media data to be processed input when applying the target media file. Based on this, generating the target media file according to the first audio data, the first display data, and the media processing logic data includes: generating the target media file according to the first audio data, the first display data, the second display data, and the media processing logic data. For detailed implementation, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 3A As shown, the method in this embodiment may specifically include:
[0095] S310. Display media data editing information, wherein the media data editing information includes an audio data editing item and a second display editing item.
[0096] In this embodiment of the disclosure, the first display editing item is used to edit second display data of target interactive data associated with the first audio data. Optionally, the second display editing item includes a second style editing item. Further, the second style editing item includes a plurality of preset second style templates.
[0097] S320. Receive audio data editing operation through the audio data editing item, determine first audio data according to the audio data editing operation, and determine first display data of the target display text associated with the first audio data.
[0098] S330. Receive a second display editing operation through the second display editing item, and determine second display data corresponding to the media data entered when applying the target media file based on the second display editing operation.
[0099] In this embodiment of the disclosure, the second display editing operation can be understood as an interactive operation with the second display editing item, used to edit the second display data of the target interactive data corresponding to the media data input when applying the target media file. The media data input when applying the target media file may include at least one of the following: second audio data, media data image data, animation sequence frames, and file application effect data, to display the media data effect data of the target media file. The first display data of the target interactive data corresponding to the media data input when applying the target media file includes display position data and / or second style data.
[0100] As an optional implementation of this disclosure, when the second display editing item includes a second style editing item and the second style editing item includes multiple preset second style templates, the step of receiving a second display editing operation through the second display editing item and determining the second display data of the target interactive data corresponding to the media data input when applying the target media file according to the second display editing operation includes: receiving a second selection operation for a template identifier of a preset second preset style template; determining a second target template among multiple second style templates according to the second selection operation and the template identifier; and determining the second style data of the target interactive data corresponding to the media data input when applying the target media file according to the second target template. By adopting this technical solution, the display style data setting of the target interactive data can be conveniently and quickly realized through simple interactive operations, supporting differentiated display of target interactive data in different media files, further enriching the presentation effect of media files, and improving the editing experience of media files.
[0101] As another optional technical solution in this disclosure, when the second display editing item includes a display position editing item, the step of receiving a second display editing operation through the second display editing item and determining the second display data of the target interactive data corresponding to the media data input when applying the target media file based on the second display editing operation may specifically include: receiving a display position editing operation through the display position editing item and determining the second display data of the target interactive data corresponding to the media data input when applying the target media file based on the display position editing operation. By adopting this technical solution, the display position of the target interactive data can be flexibly adjusted, supporting diversified presentation of the display position of the target interactive data to enrich the display effect of the target interactive data, thereby enriching the presentation effect of the target media file.
[0102] For example, the display position editing item includes a display position selection control and / or a position change control, etc. When the display position editing item includes a display position selection control, the step of receiving a display position editing operation through the display position editing item and determining the second display data corresponding to the target interactive data input when applying the target media file based on the display position editing operation may specifically include: receiving an editing trigger operation for the display position selection control, displaying multiple candidate position adjustment items, wherein the position adjustment items correspond to a preset display position; receiving an adjustment trigger operation for the position adjustment items, and adjusting the display position data of the target interactive data according to the position adjustment item corresponding to the preset display position.
[0103] For example, when the display position editing item includes a position transformation control, the step of receiving a display position editing operation through the display position editing item and determining the second display data of the target interactive data corresponding to the media data input when applying the target media file, based on the display position editing operation, may specifically include: receiving a display position editing operation through the display position transformation control, determining the display position data of the target interactive data corresponding to the input media data before transformation, and determining the display position data of the target interactive data after transformation based on the display position data of the target interactive data before transformation and the preset transformation rule corresponding to the display position transformation control. For example, the preset transformation rule may include switching between multiple preset display positions according to a preset switching order; or moving the display position of the target interactive data according to a preset trajectory, etc.
[0104] like Figure 3BAs shown, the first display editing item corresponding to the target display text is displayed in the first parameter panel. The first display editing item includes an object information component, a display position component, and an interactive data component. The object information component is used to display the object identification information corresponding to the edited object in the media data display, and is used to characterize the effective object of the parameters in the first parameter panel. The display position component is used to edit the display position data corresponding to the target interactive data. For example, it can be located above or below the target display text, and / or displayed in association with the target display text in a preset manner, etc. The interactive data component includes template identifiers for multiple interactive style templates. Among them, the template identifiers include graphic identifiers and text identifiers. The graphic identifiers can display preset images or style effect images, etc. By triggering the graphic identifiers, text identifiers, or preset areas including graphic identifiers and text identifiers in the template identifiers, the template identifier corresponding to the triggered template identifier (e.g., style template L) can be highlighted, and a style preview effect (not shown in the figure) corresponding to the triggered template identifier can be displayed. The audio data component displays the data identifier of the first audio data being edited in order to determine the association between the second display data being edited and the audio data, in order to determine which audio data is used to render the target interactive data, with which audio data generates the target media file, or in other words, in which target media file it is encapsulated.
[0105] S340. In response to the file generation request, obtain the pre-set media processing logic data corresponding to the first audio data, and generate the target media file according to the first audio data, the first display data, the second display data and the media processing logic data.
[0106] In this embodiment of the disclosure, upon receiving a file generation request, a target media file can be automatically generated by executing a preset media data generation logic file based on the first audio data, the first display data, the second display data, and the media processing logic data. Because the target media file includes second display data for the target interactive data, and interactive processing logic data is added to the media processing logic data, the media data-related content in the target media file is enriched, enabling the target media file to present the target interactive data during application, further enriching the media data presentation effect.
[0107] The technical solution of this disclosure embodiment, by receiving a second display editing operation through the second display editing item after the display media data editing information, and determining the second display data of the target interactive data corresponding to the media data to be processed according to the second display editing operation, can realize flexible setting of the display data of the target interactive data, support the diversified presentation of the target interactive data in different target media files, and better meet the differentiated setting needs of different users, enrich the presentation effect of media files, and improve the editing experience of media files.
[0108] Figure 4 This is a schematic diagram of the structure of a media data processing apparatus provided in an embodiment of the present disclosure, as shown below. Figure 4 As shown, the device includes: an editing information display module 410, a media data editing module 420, and a media file generation module 430. The editing information display module displays media data editing information, including audio data editing items. The media data editing module 420 receives audio data editing operations through the audio data editing items, determines first audio data based on the audio data editing operations, and determines first display data of target display text associated with the first audio data. The media file generation module 430, in response to a file generation request, acquires pre-set media processing logic data corresponding to the first audio data, and generates a target media file based on the first audio data, the first display data, and the media processing logic data.
[0109] The technical solution of this embodiment displays media data editing information through the editing information display module 410, realizing the visual presentation of media data editing information. Since the media data editing information includes audio data editing items, it realizes the visual editing of audio data in the media data production scenario. Then, the media data editing module 420 receives audio data editing operations through the audio data editing items, determines the first audio data according to the audio data editing operations, and determines the first display data of the target display text associated with the first audio data. This enables flexible editing of audio data and can automatically determine the text data that matches the audio data. Then, the media file generation module 430 responds to the file generation request, obtains the pre-set media processing logic data corresponding to the first audio data, and generates the target media file according to the first audio data, the first display data, and the media processing logic data. This solves the technical problems of low efficiency, high difficulty, narrow applicability, and monotonous presentation effects in related technologies. Media files can be automatically, quickly, and conveniently generated through simple interactive operations, reducing the difficulty of media file production, improving the efficiency of media file production, and supporting the association and display of audio data with its corresponding text data in the media file, enriching the presentation effect of the media file.
[0110] Based on any optional technical solution in the embodiments of this disclosure, the audio data editing item may optionally include at least one of an audio data selection item, an audio data upload item, an audio data trimming item, and an audio data splicing item. Further, the media data editing module 420 may include an audio data editing unit. The audio data editing unit can be used to perform at least one of the following operations: receiving an audio data selection operation for at least one segment of preset audio data through the audio data selection item; receiving an audio data upload operation through the audio data upload item; receiving an audio data trimming operation for at least one segment of preset audio data through the audio data trimming item; and receiving an audio data splicing operation for at least two segments of preset audio data through the audio data splicing item.
[0111] Optionally, based on any of the optional technical solutions in the embodiments of this disclosure, the media data editing information may further include a first display editing item. Further, the media data editing module 420 may include a text display editing unit. The text display editing unit can be used to receive a first display editing operation through the first display editing item, and determine first display data of the target display text associated with the first audio data based on the first display editing operation.
[0112] Based on any optional technical solution in the embodiments of this disclosure, optionally, the first display editing item includes a first style editing item; the first style editing item may include template identifiers of multiple preset first style templates. Further, the text display editing unit may include a text style editing subunit, wherein the text style editing subunit is specifically configured to receive a first trigger operation for a template identifier of a preset first style template, determine a first target template among multiple first style templates based on the first trigger operation, and determine first display data of target display text associated with the first audio data based on the first target template.
[0113] Optionally, based on any of the optional technical solutions in the embodiments of this disclosure, the first style editing item further includes a template data setting item corresponding to the first style template. Based on this, the text style editing subunit can specifically be used to: display preset style data of the first target template through the style data setting item; and / or, receive data adjustment operations on the preset style data of the first target template through the style data setting item, and determine first style data of the target text data associated with the first audio data based on the data adjustment operations and the preset style data.
[0114] Optionally, based on any optional technical solution in the embodiments of this disclosure, the first display editing item includes a display position editing item. Further, the text display editing unit may include a display position editing subunit. The display position editing subunit can be used to receive a display position editing operation through the display position editing item, and determine, based on the display position editing operation, first display data of the target display text associated with the first audio data.
[0115] Based on any optional technical solution in the embodiments of this disclosure, optionally, the media processing logic data includes audio processing logic data of the first audio data and the second audio data in the media data to be processed input when applying the target media file, as well as text processing logic data of the target display text; the audio processing logic data is configured to play the first audio data and record the first audio data and the second audio data in the media data to be processed; the text processing logic data is configured to display the target display text.
[0116] Optionally, based on any optional technical solution in the embodiments of this disclosure, the text processing logic data is configured to: acquire preset text data corresponding to the first audio data, and determine target text data based on the preset text data; wherein, the target text data includes target display text and display time data for each character in the target display text; the display time data includes start display time, display duration, and end display time; and display the target display text based on the first audio data, the target text data, and the first display data.
[0117] Based on any optional technical solution in the embodiments of this disclosure, the media processing logic data may optionally include interactive processing logic data; the interactive processing logic data is configured to display target interactive data corresponding to the input media data when the target media file is applied.
[0118] Based on any optional technical solution in the embodiments of this disclosure, the interactive processing logic data is optionally configured to: generate target interactive data corresponding to the media data to be processed based on the first audio data and the input media data to be processed when applying the target media file, and display the target interactive data.
[0119] Based on any optional technical solution in the embodiments of this disclosure, optionally, the media data editing information includes an interactive data editing item corresponding to the target interactive data; the interactive data editing item includes a second display editing item. Further, the media data processing device may also include an interactive display editing module. The interactive display editing module is used to receive a second display editing operation through the second display editing item after the display media data editing information is provided, and to determine, based on the second display editing operation, second display data corresponding to the target interactive data input when applying the target media file. Further, the media file generation module 430 is specifically used to generate a target media file based on the first audio data, the first display data, the second display data, and the media processing logic data.
[0120] The media data processing apparatus provided in this disclosure can execute the media data processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the media data processing method.
[0121] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0122] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 500 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0123] like Figure 5As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0124] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0125] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0126] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0127] The electronic device provided in this disclosure and the media data processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this disclosure can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0128] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the media data processing method provided in the above embodiments.
[0129] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0130] According to one or more embodiments of this disclosure, [Example 1] provides a media data processing method, including: displaying media data editing information, wherein the media data editing information includes an audio data editing item; receiving an audio data editing operation through the audio data editing item, determining first audio data according to the audio data editing operation, and determining first display data of target display text associated with the first audio data; responding to a file generation request, obtaining pre-set media processing logic data corresponding to the first audio data, and generating a target media file according to the first audio data, the first display data, and the media processing logic data.
[0131] According to one or more embodiments of this disclosure, [Example 2] provides a media data processing method of Example 1, further comprising: optionally, the audio data editing item includes at least one of an audio data selection item, an audio data upload item, an audio data trimming item, and an audio data splicing item; the step of receiving an audio data editing operation through the audio data editing item includes at least one of the following operations: receiving an audio data selection operation for at least one segment of preset audio data through the audio data selection item; receiving an audio data upload operation through the audio data upload item; receiving an audio data trimming operation for at least one segment of preset audio data through the audio data trimming item; and receiving an audio data splicing operation for at least two segments of preset audio data through the audio data splicing item.
[0132] According to one or more embodiments of this disclosure, Example 3 provides a media data processing method of Example 1, which further includes: optionally, the media data editing information further includes a first display editing item; the first display data for determining the target display text associated with the first audio data includes: receiving a first display editing operation through the first display editing item, and determining the first display data for the target display text associated with the first audio data based on the first display editing operation.
[0133] According to one or more embodiments of this disclosure, [Example 4] provides a media data processing method similar to Example 3, further comprising: optionally, the first display editing item includes a first style editing item; the first style editing item includes template identifiers of a plurality of preset first style templates; the step of receiving a first display editing operation through the first display editing item and determining first display data of target display text associated with the first audio data based on the first display editing operation includes: receiving a first trigger operation for the template identifier of the preset first style template, determining a first target template among the plurality of first style templates based on the first trigger operation, and determining first display data of target display text associated with the first audio data based on the first target template.
[0134] According to one or more embodiments of this disclosure, [Example 5] provides a media data processing method of Example 4, further comprising: Optionally, the first style editing item further comprises the first style data of the style data setting item, which determines the target text data associated with the first audio data for display based on the first target template, including:
[0135] The preset style data of the first target template is displayed through the style data settings; and / or,
[0136] The style data setting item receives a data adjustment operation for the preset style data of the first target template, and determines the first style data of the target text data to be displayed in association with the first audio data based on the data adjustment operation and the preset style data.
[0137] According to one or more embodiments of this disclosure, Example Six provides a media data processing method similar to Example Three, further comprising: optionally, the first display editing item includes a display position editing item; the step of receiving a first display editing operation through the first display editing item and determining first display data of target display text associated with the first audio data based on the first display editing operation includes: receiving a display position editing operation through the display position editing item and determining first display data of target display text associated with the first audio data based on the display position editing operation.
[0138] According to one or more embodiments of this disclosure, [Example Seven] provides a media data processing method of Example One, further comprising: optionally, the media processing logic data includes audio processing logic data of first audio data and second audio data in the media data to be processed input when applying the target media file, and text processing logic data of the target display text; the audio processing logic data is configured to play the first audio data and record the first audio data and the second audio data in the media data to be processed; the text processing logic data is configured to display the target display text.
[0139] According to one or more embodiments of this disclosure, Example 8 provides a media data processing method of Example 7, further comprising: optionally, the text processing logic data is configured to: acquire preset text data corresponding to the first audio data, and determine target text data based on the preset text data; wherein, the target text data includes target display text and display time data for each character in the target display text; the display time data includes start display time, display duration, and end display time; and display the target display text based on the first audio data, the target text data, and the first display data.
[0140] According to one or more embodiments of this disclosure, Example 9 provides a media data processing method of Example 7, further comprising: optionally, the media processing logic data includes interactive processing logic data; the interactive processing logic data is configured to display target interactive data corresponding to the input media data to be processed when the target media file is applied.
[0141] According to one or more embodiments of this disclosure, Example 10 provides a media data processing method of Example 9, which further includes: optionally, the interactive processing logic data is configured to: generate target interactive data corresponding to the media data to be processed based on the first audio data and the input media data to be processed when the target media file is applied, and display the target interactive data.
[0142] According to one or more embodiments of this disclosure, Example 11 provides a media data processing method of Example 9, further comprising: optionally, the media data editing information includes an interactive data editing item corresponding to the target interactive data; the interactive data editing item includes a second display editing item; after displaying the media data editing information, the method further comprises: receiving a second display editing operation through the second display editing item, and determining second display data of the target interactive data corresponding to the media data to be processed input when applying the target media file according to the second display editing operation; further, generating the target media file according to the first audio data, the first display data and the media processing logic data includes: generating the target media file according to the first audio data, the first display data, the second display data and the media processing logic data.
[0143] According to one or more embodiments of this disclosure, [Example Twelve] provides a media data processing apparatus, comprising: an editing information display module for displaying media data editing information, wherein the media data editing information includes audio data editing items; a media data editing module for receiving audio data editing operations through the audio data editing items, determining first audio data based on the audio data editing operations, and determining first display data of target display text associated with the first audio data; and a media file generation module for, in response to a file generation request, acquiring pre-set media processing logic data corresponding to the first audio data, and generating a target media file based on the first audio data, the first display data, and the media processing logic data.
[0144] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0145] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0146] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the one or more programs, the electronic device causes the electronic device to: display media data editing information, wherein the media data editing information includes audio data editing items; receive audio data editing operations through the audio data editing items; determine first audio data based on the audio data editing operations; and determine first display data of target display text associated with the first audio data; and, in response to a file generation request, obtain pre-set media processing logic data corresponding to the first audio data, and generate a target media file based on the first audio data, the first display data, and the media processing logic data.
[0147] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0149] The modules and units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules and units do not necessarily limit the specific unit; for example, a media data editing module can also be described as "a module for editing first audio data and determining first display data of target display text associated with the first audio data."
[0150] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0151] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0152] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0153] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0154] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A media data processing method, characterized in that, include: Display media data editing information, wherein the media data editing information includes audio data editing items; The audio data editing item receives an audio data editing operation, determines the first audio data based on the audio data editing operation, and determines the first display data of the target display text associated with the first audio data. In response to a file generation request, pre-set media processing logic data corresponding to the first audio data is obtained, and a target media file is generated based on the first audio data, the first display data, and the media processing logic data.
2. The media data processing method according to claim 1, characterized in that, The audio data editing item includes at least one of the following: audio data selection item, audio data upload item, audio data trimming item, and audio data splicing item; receiving audio data editing operations through the audio data editing item includes at least one of the following operations: The audio data selection option receives an audio data selection operation for at least one segment of preset audio data. Receive audio data upload operations through the audio data upload item; The audio data trimming item receives an audio data trimming operation for at least one segment of preset audio data. The audio data splicing item receives audio data splicing operations for at least two preset audio data segments.
3. The media data processing method according to claim 1, characterized in that, The media data editing information further includes a first display editing item; the first display data for determining the target display text associated with the first audio data includes: The system receives a first display editing operation through the first display editing item and determines the first display data of the target display text associated with the first audio data based on the first display editing operation.
4. The media data processing method according to claim 3, characterized in that, The first display editing item includes a first style editing item; the first style editing item includes template identifiers for multiple preset first style templates; the step of receiving a first display editing operation through the first display editing item and determining the first display data of the target display text associated with the first audio data based on the first display editing operation includes: Receive a first trigger operation for a template identifier of a preset first style template, determine a first target template among a plurality of first style templates based on the first trigger operation, and determine first display data of target display text associated with the first audio data based on the first target template.
5. The media data processing method according to claim 4, characterized in that, The first style editing item further includes a style data setting item; the first style data for determining the target text data associated with the first audio data based on the first target template includes: The preset style data of the first target template is displayed through the style data settings. The style data setting item receives a data adjustment operation for the preset style data of the first target template, and determines the first style data of the target text data to be displayed in association with the first audio data based on the data adjustment operation and the preset style data.
6. The media data processing method according to claim 3, characterized in that, The first display editing item includes a display position editing item; the step of receiving a first display editing operation through the first display editing item and determining the first display data of the target display text associated with the first audio data based on the first display editing operation includes: The display position editing item receives a display position editing operation, and the first display data of the target display text associated with the first audio data is determined based on the display position editing operation.
7. The media data processing method according to claim 1, characterized in that, The media processing logic data includes audio processing logic data of first audio data and second audio data in the media data to be processed input when applying the target media file, as well as text processing logic data of the target displayed text; the audio processing logic data is configured to play the first audio data and record the first audio data and the second audio data in the media data to be processed. The text processing logic data is configured to display the target display text.
8. The media data processing method according to claim 7, characterized in that, The text processing logic data is configured as follows: Obtain preset text data corresponding to the first audio data, and determine target text data based on the preset text data; wherein, the target text data includes target display text and display time data for each character in the target display text; the display time data includes start display time, display duration, and end display time; The target display text is displayed based on the first audio data, the target text data, and the first display data.
9. The media data processing method according to claim 7, characterized in that, The media processing logic data includes interactive processing logic data; the interactive processing logic data is configured to display target interactive data corresponding to the input media data to be processed when the target media file is applied.
10. The media data processing method according to claim 9, characterized in that, The interactive processing logic data is configured as follows: When applying the target media file, target interactive data corresponding to the target media data is generated based on the first audio data and the input media data to be processed, and the target interactive data is displayed.
11. The media data processing method according to claim 9, characterized in that, The media data editing information includes interactive data editing items corresponding to the target interactive data; the interactive data editing items include a second display editing item; Following the display media data editing information, the following is also included: The second display editing operation is received through the second display editing item, and the second display data corresponding to the target interactive data to be processed is determined based on the second display editing operation. The step of generating a target media file based on the first audio data, the first display data, and the media processing logic data includes: The target media file is generated based on the first audio data, the first display data, the second display data, and the media processing logic data.
12. A media data processing device, characterized in that, include: An editing information display module is used to display media data editing information, wherein the media data editing information includes audio data editing items; The media data editing module is used to receive audio data editing operations through the audio data editing item, determine first audio data according to the audio data editing operations, and determine first display data of target display text associated with the first audio data; The media file generation module is used to respond to a file generation request, obtain pre-set media processing logic data corresponding to the first audio data, and generate a target media file based on the first audio data, the first display data, and the media processing logic data.
13. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the media data processing method as described in any one of claims 1-11.
14. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the media data processing method as described in any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the media data processing method as described in any one of claims 1-11.