Media data generation method, apparatus, device, medium, and product
Patent Information
- Application Number
- CN202510345017.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]现有的媒体数据制作软件多以“图生图、图生视频”等方式制作媒体数据,这种媒体制作方式所成的媒体数据内容相对单一,难以满足用户需求
[0019]本公开实施例所提供的技术方案,通过响应于接收到在媒体数据生成页面中所配置的至少一个第一媒体数据的操作,在显示页面中展示所述至少一个第一媒体数据;响应于接收到与所述第一媒体数据所对应的第二媒体数据的事件,依据预设形态在所述显示页面中展示所述第二媒体数据;其中,所述第二媒体数据包括文本信息以及与所述文本信息相对应的音频信息,所述文本信息与所述第一媒体数据的媒体内容相关,解决了现有技术中媒体数据制作内容单一,难以满足用户制作需求的问题,实现了通过在接收到媒体数据生成页面中配置的至少一个第一媒体数据后,在显示页面中展示至少一个第一媒体数据,实现第一媒体数据可视化的同时,便于用户及时确认所生成的第一媒体数据是否满足需求。进一步的,通过在接收到与第一媒体数据所对应的多模态的第二媒体数据时,依据预设形态在显示页面中展示第二媒体数据中的文本信息以及与文本信息相对应的音频信息,提高了制作媒体数据内容丰富性的同时,提高了媒体数据生成便捷性,达到满足用户个性化需求的技术效果。
Smart Images

Figure CN122802746A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer processing technology, and more particularly to a media data generation method, apparatus, device, medium, and product. Background Technology
[0002] With the development of artificial intelligence technology, more and more users are using related media data production software to create media data.
[0003] Existing media data production software mostly creates media data in the form of "image-to-image, image-to-video", which results in relatively simple media data content that is difficult to meet user needs. Summary of the Invention
[0004] This disclosure provides a media data generation method, apparatus, device, medium, and product to improve the richness of generated media data content while enhancing the convenience of media data generation, thereby meeting users' personalized needs.
[0005] In a first aspect, embodiments of this disclosure provide a media data generation method, the method comprising:
[0006] In response to receiving an operation that configures at least one first media data in the media data generation page, the at least one first media data is displayed in the display page;
[0007] In response to receiving an event that corresponds to the first media data, the second media data is displayed on the display page according to a preset format;
[0008] The second media data includes text information and audio information corresponding to the text information, wherein the text information is related to the media content of the first media data.
[0009] Secondly, embodiments of this disclosure also provide a media data generation apparatus, the apparatus comprising:
[0010] The first media data display module is used to respond to receiving an operation that configures at least one first media data in the media data generation page, and to display the at least one first media data in the display page;
[0011] The second media data display module is used to respond to an event of receiving second media data corresponding to the first media data and to display the second media data on the display page according to a preset format.
[0012] The second media data includes text information and audio information corresponding to the text information, wherein the text information is related to the media content of the first media data.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the media data generation method as described in any of the embodiments of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the media data generation method as described in any of the embodiments of this disclosure.
[0018] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the media data generation method as described in any of the embodiments of this disclosure.
[0019] The technical solution provided in this disclosure, in response to receiving at least one first media data configured on a media data generation page, displays the at least one first media data on a display page; in response to receiving an event of second media data corresponding to the first media data, displays the second media data on the display page according to a preset form; wherein the second media data includes text information and audio information corresponding to the text information, the text information being related to the media content of the first media data, solves the problem in the prior art that the content of media data production is monotonous and difficult to meet user production needs, and realizes that after receiving at least one first media data configured on the media data generation page, at least one first media data is displayed on the display page, realizing the visualization of the first media data and facilitating users to promptly confirm whether the generated first media data meets their needs. Furthermore, by displaying the text information and corresponding audio information in the second media data on the display page according to a preset form when receiving multimodal second media data corresponding to the first media data, the richness of media data production content is improved, while the convenience of media data generation is enhanced, achieving the technical effect of meeting users' personalized needs. Attached Figure Description
[0020] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 A schematic flowchart illustrating a media data generation method provided in an embodiment of this disclosure;
[0022] Figure 2 This is a schematic diagram for characterizing second media data provided in an embodiment of the present disclosure;
[0023] Figure 3 This is a schematic diagram illustrating a lyric component provided in an embodiment of the present disclosure;
[0024] Figure 4 This is a schematic diagram illustrating an audio playback component provided in an embodiment of this disclosure;
[0025] Figure 5 A schematic flowchart illustrating a media data generation method provided in an embodiment of this disclosure;
[0026] Figure 6 This is a schematic diagram of the media data generation method provided in the embodiments of this disclosure;
[0027] Figure 7 This is a schematic diagram for representing a first page according to an embodiment of the present disclosure;
[0028] Figure 8 This is a schematic diagram of an interface for characterizing audio parameter settings provided in an embodiment of this disclosure;
[0029] Figure 9 This is a schematic diagram of the media data generation method provided in the embodiments of this disclosure;
[0030] Figure 10 This is a timing diagram of the generation of media data according to embodiments of this disclosure;
[0031] Figure 11 This is a schematic diagram of the structure of a media data generation apparatus provided in an embodiment of the present disclosure.
[0032] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0033] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0034] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0035] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0036] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0037] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0038] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0039] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0040] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0041] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0042] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0043] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0044] Before introducing the technical solutions provided by the embodiments of this disclosure, the application scenarios can be illustrated first. The technical solutions provided by the embodiments of this disclosure can be applied to any scenario that requires the generation of media data. For example, during the development phase, developers can develop a media data production chain and integrate the media data production chain into a functional module or special effects toolkit. The functional module or special effects toolkit can be deployed in any existing platform, application software, or tool that supports media data generation, so that users can generate multimodal media data that meets their personalized needs based on the deployed functional module or special effects toolkit.
[0045] In practical applications, if a user wants to obtain media data that meets their needs, they can trigger the functional modules or special effects packages provided in the embodiments of this disclosure. After detecting the triggering of a functional module or special effects package, a media data generation page corresponding to the embodiments of this disclosure can be displayed to generate corresponding media data based on the editing operations on the media data generation page.
[0046] Figure 1 This is a flowchart illustrating a media data generation method provided in an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to any scenario that requires the generation of media data. The method can be executed by a media data generation device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.
[0047] like Figure 1 As shown, the method in this embodiment may specifically include:
[0048] S110. In response to receiving an operation that configures at least one first media data in the media data generation page, display at least one first media data in the display page.
[0049] The media data generation page can be a page used to configure or edit the first media data. For example, this page can be an interface in a media data generation tool, software, platform, or similar application. For instance, if the technical solution provided in this embodiment is integrated into a functional module or special effects package, when a trigger of the functional module or special effects package is detected, a main interface corresponding to the functional module or special effects package can be displayed. When an image upload or image capture function in this main interface is triggered, the displayed page is the media data generation page; or, when a trigger of the functional module or special effects package is detected, the displayed page is a capture page or an image upload page, and this page is used as the media data generation page. The first media data can refer to the initially generated, edited, or configured media content, and the first media data is related to the media data that the user wants to obtain. For example, the first media data can be media data captured in real time, or media data retrieved and uploaded from a media database. In this embodiment, the first media data can be one or a combination of multiple types of video, images, charts, audio, and text. The number of primary media data can be one or more. If there are multiple primary media data, the types of different primary media data can be the same or different. The display page is used to display the primary media data. For example, the display page can be a window page, an overlay page, or a pop-up page within the media data generation page; it can also be a new window page.
[0050] In this embodiment of the disclosure, at least one first media data can be edited on the media data generation page. After the configuration or editing operation of the first media data is completed, the system responds to this operation. At this time, the configured at least one first media data can be obtained and displayed on the display page, allowing the user to preview or further manipulate these media data to obtain the final desired media data. It should be noted that a range for the number of media data can also be preset. This ensures that when editing the first media data on the media data generation page, the number of edited first media data is within the specified range, such as more than 2 and less than 5 first media data. This setting can ensure the richness and diversity of the generated media data while avoiding information overload caused by excessive data and preventing any impact on the quality of the generated media data.
[0051] S120: In response to an event that a second media data corresponding to the first media data has been received, the second media data is displayed on the display page according to a preset format.
[0052] The second media data is media data processed from the first media data. For example, the second media data can be media data that extends, supplements, or is derived from the first media data. Optionally, the second media data is music data corresponding to the first media data. The music data includes text information and audio information corresponding to the text information. The audio information can be audio used to describe the text information, or music or related sound effects used to reflect the text content in the text information. The text information is related to the media content of the first media data. For example, the text information can be a description, explanation, or supplement to the first media data, or it can be lyrics corresponding to the media content. The preset format can be a pre-configured layout, style, format, or animation effect used to display the second media data. For example, the preset format is: the text information is displayed with a specific font, color, and layout; the audio information is displayed through player controls and may have volume control, pause / play buttons, etc.; the audio information is displayed next to, below, or in a pop-up window format.
[0053] In this embodiment, once at least one first media data is determined, the first media data can be processed based on either the client or the server. For example, text information related to the media content can be added to the first media data, or audio information corresponding to the text information can be added. The processed first media data is then used as second media data. It should be noted that when generating second media data based on the client, the second media data can be directly displayed on the client's display page according to a preset format. When generating second media data based on the server, the second media data can be sent to the client from the server. When the client receives the second media data, it can display the second media data on the display page according to a preset format, allowing users to intuitively see the display effect of the generated media data and choose to view or play the displayed second media data according to their needs.
[0054] For example, when a user interacts with the first media data (e.g., by clicking or hovering), the system triggers an event, sending the first media data to the server. The server then generates second media data corresponding to the first media data. When the system receives the second media data returned by the server, it can display the second media data on the display page according to a preset format. For example, text information in the second media data can be displayed as a pop-up, sidebar, or overlay, while audio information can be played automatically or triggered by the user clicking a play button.
[0055] Optionally, the preset formats include a preset cover format and a preset text display format. Correspondingly, the method of displaying the second media data on the display page according to the preset formats can be: displaying at least one first media data item in the second media data on the display page according to the preset cover format; displaying text information in the second media data on the display page according to the preset text display format; and when playing audio information corresponding to the text information in the second media data, displaying the first text corresponding to the played audio information in the text information in the first format.
[0056] The preset cover format can be the style and layout of the first media data (such as images, video clips, etc.) used to display the second media data, including but not limited to: the size, shape, and position of the cover; the style of the image or video on the cover (such as whether there is a border, whether filters are added, etc.); the interactive method of the cover (such as clicking the cover to play the second media data); and the display format (such as displaying it as a 3D model or a 2D model). The preset text display format can be the display format used to display text information, including but not limited to: text layout (such as displaying it next to, below, or as a pop-up window on the cover); style (such as font, font size, color, and typography); format; animation effects; and the interactive method of the text (such as clicking the text to play the second media data). The first format can be displayed through a player control and can include volume control, pause / play buttons, etc.
[0057] In this embodiment, at least one first media data item from the second media data can be displayed on the display page according to a preset cover format. For example, the first media data items can be rotated and displayed in a floating window, a card in a fixed position, or a 3D model. The text information from the second media data is displayed on the display page according to a preset text display format. For example, the text information can be a description of the audio information, lyrics, background story, etc. When an interactive operation is detected on the display page (such as clicking the cover or text, clicking the playback control, or hovering), the audio information corresponding to the text information in the second media data is played. During the playback of the audio information, the first text (such as lyrics or narration text) corresponding to the played audio information is displayed in the text information in a first format (such as highlighting, color changing, scrolling, etc.). For example, see [link to example]. Figure 2 The first media data is the cover image or video clip of the second media data. The text information in the second media data is lyrics or music introduction, and the audio information is the music file. The first media data in the second media data is displayed as the cover (displayed according to the preset cover format). The playback progress of the audio information can be adjusted by sliding the player. The lyrics (i.e., text information) scroll and play synchronously with the audio information and are highlighted.
[0058] It should be noted that after acquiring the second media data, the system can automatically play it. Alternatively, the second media data can be cached as a music resource, and the playback of audio information or the display of text information within the music resource can be controlled by manipulating the lyrics component or audio playback component. For example, the audio information can be played after the album art is triggered. Or, the audio information can be played when the text information or a line of text within the text information is triggered. Alternatively, the audio information can be played when the corresponding playback control is triggered, and the lyrics can be displayed on the display page in a preset text format. For example, when playing audio information, the currently playing lyrics (i.e., the first text) can be highlighted, enhancing the audiovisual experience and helping users better follow the music.
[0059] For example, the audio file of the second media data can be AI music, and a schematic diagram showing the association of the second media data with the lyrics component can be found here. Figure 3 The lyrics component allows control over parameters for playing lyrics (such as title height, title font size, title font, title font, first line height, first line lyric color, first line font size, first line font, second line height, second line lyric color, second line font size, second line font, etc.). These parameters may or may not be modifiable; there is no limitation on this. A diagram illustrating the association of secondary media data with the audio playback component can be found in [link to diagram]. Figure 4 The audio playback component can control parameters for playing audio information (such as function type, playback mode, whether to play automatically, whether to play silently, volume, etc.). These parameters may or may not be modifiable; there is no limitation on this.
[0060] The technical solution provided in this disclosure can enhance the expressive effect of the second media data content by preset cover and text display formats, improving the richness and professionalism of the media data content while allowing users to quickly understand the content of the second media data through the cover. Furthermore, the synchronized display of text and audio enriches the media data content and meets users' personalized needs. Simultaneously, after acquiring the second media data, it can be cached as a music resource. Playing the music resource to play the second media data effectively ensures smooth playback and avoids playback stuttering.
[0061] It should be noted that when displaying the second media data, the background information of the second media data can also be displayed based on preset background information; alternatively, at least one of the first media data in the second media data can be processed, such as blurring, adjusting transparency and brightness, and the processed first media data can be used as the background information of the second media data. Alternatively, text information and first media data can be merged to obtain the background information of the second media data; or text information and preset background information can be merged to obtain the background information of the second media data.
[0062] To further enhance the display effect of media data and make the display background match the music data, the background media data for displaying the second media data can be retrieved based on the first style type corresponding to the second media data, and the background information of the second media data can be rendered based on the background media data.
[0063] The first style type can be used to characterize the theme or visual style of the second media data, such as, but not limited to: retro, modern, romantic, technological, cheerful, sad, exciting, rock, jazz, electronic, minimalist, etc. Background media data refers to the media content used to render the background of the second media data. Background media data is related to the first style type of the second media data.
[0064] In this embodiment, the style type of the second media data can be identified as the first style type. Then, background media data matching the first style type can be retrieved from a server or database according to a preset mapping relationship or matching algorithm. For example, if the first style type of the second media data is "retro," a retro-style background image or video is selected as the background media data. Furthermore, the background media data can be used as the background layer of the second media data; that is, when the system displays the second media data, rendering technology can be used to render the background information of the second media data based on the retrieved background media data, allowing the background information to be combined with text and audio information. It should be noted that, to ensure that the display of the background media data does not interfere with the display of the second media data (text and audio information), the brightness and transparency of the background media data can be dynamically adjusted based on the text and audio information. By introducing background media data matching the first style type to render the background information of the second media data, the visual effect of the second media data can be enhanced, making the display of the second media data more vivid and improving the user experience. At the same time, different style types can correspond to different background media data, enabling personalized media data production and meeting the needs of different users.
[0065] The technical solution provided in this disclosure, in response to receiving at least one first media data configured on a media data generation page, displays the at least one first media data on a display page; in response to receiving an event of second media data corresponding to the first media data, displays the second media data on the display page according to a preset form; wherein, the second media data includes text information and audio information corresponding to the text information, the text information being related to the media content of the first media data, solves the problem in the prior art that the content of media data production is monotonous and difficult to meet user production needs, and realizes that by receiving at least one first media data configured on the media data generation page, at least one first media data is displayed on the display page, realizing the visualization of the first media data and facilitating users to promptly confirm whether the generated first media data meets their needs. Furthermore, by displaying the text information and corresponding audio information in the second media data on the display page according to a preset form when receiving the second media data corresponding to the first media data, the richness of the generated media data content is improved, while the convenience of media data generation is enhanced, achieving the technical effect of meeting users' personalized needs.
[0066] Figure 5 This is a schematic flowchart illustrating a media data generation method provided in an embodiment of this disclosure. Based on the above embodiments, the technical solution of this embodiment further includes sending at least one first media data point to a server after displaying the first media data, so that the server determines second media data based on audio parameter information and the at least one first media data point. For detailed implementation methods, please refer to the detailed description of the embodiments of this disclosure. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here.
[0067] like Figure 5 As shown, the method in this embodiment may specifically include:
[0068] S210. In response to receiving an operation that configures at least one first media data in the media data generation page, display the at least one first media data in the display page.
[0069] S220, Send at least one first media data to the server.
[0070] In this embodiment, first media data can be edited or uploaded on a client (such as a webpage or mobile application). Then, at least one piece of first media data can be sent to a server via an interface, remote transmission, HTTP / HTTPS, WebSocket, or other communication protocols. The server then determines second media data corresponding to the first media data based on the at least one piece of first media data and pre-configured audio parameter information. The server generates the second media data. The audio parameter information can be parameters used to guide how to generate or select matching audio content based on the first media data. For example, the audio parameter information can be music style (such as rock, jazz, classical), music rhythm (such as fast tempo, slow tempo), music emotion (such as cheerful, sad, exciting), audio format (such as MP3, WAV), audio length, etc. The second media data is the media data generated after processing based on the first media data and the audio parameter information.
[0071] Specifically, after receiving at least one piece of first media data, the server can retrieve matching second media data from the database based on the characteristics of the first media data (such as content type, theme, and emotion) and pre-configured audio parameters. Alternatively, it can use a machine learning model or algorithm engine to generate audio information based on at least one piece of first media data and audio parameters. Then, natural language generation technology can be used to generate a text description or lyrics that match the audio information, serving as text information, thereby obtaining the audio and text information in the second media data. Alternatively, a third-party audio service can be invoked to obtain adapted second media data. Furthermore, the server can send the second media data back to the client, allowing the client to receive the second media data and display it on a display page according to a preset format. For example, see [link to example]. Figure 6 Users can select special effects packages and configure at least one primary media data. After acquiring the user-inputted primary media data locally, the system saves it and uploads it to a resource storage server (such as the cloud). The resource storage server returns a URL link for the primary media data. The client then appends media data request parameters and sends it to the cloud editing server to request secondary media data. The cloud editing server requests both a text-to-image server (i.e., a text conversion server) and an audio server to obtain the URL link for the secondary media data. The client can then download and play music data based on the URL link of the secondary media data.
[0072] The advantage of this setup is that the client only needs to upload the first media data, while the server is responsible for generating the second media data that matches the first media data. This reduces the processing burden on the client while improving data processing efficiency and enhancing the user experience.
[0073] In this embodiment of the disclosure, after displaying at least one first media data, at least one first media data can be sent to the server to determine a first style type of the first media data based on the server's analysis and processing of the at least one first media data. The first style type is a type among at least one pre-set second style type. Based on the server retrieving audio parameter information corresponding to the first style type, the second media data can be determined based on the audio parameter information and at least one first media data.
[0074] It is understandable that after sending at least one piece of first media data to the server, the server can use algorithms to analyze and process the first piece of first media data. The analysis and processing methods can be related to the data type of the first media data. For example, for first media data containing images or videos, image recognition technology can be used to identify its style type; for first media data containing text, text analysis technology can be used to analyze its emotion or theme; for first media data containing audio content, audio analysis technology can be used to analyze its rhythm, pitch, and other features. Based on the analysis results, the first style type matching the first media data is determined from at least one second style type. Furthermore, the server can retrieve audio parameter information matching the first style type from a database or storage. For example, the audio parameter information corresponding to different style types can be pre-configured and have a mapping relationship. The audio parameter information mapped to the first style type can be retrieved by using the first style type as a key. This allows the server to determine the second media data based on the audio parameter information and at least one piece of first media data. The advantage of this setting is that by analyzing the first style type of the first media data, the style type of the generated second media data is the same as or similar to the style type of the first media data input by the user. This ensures that the style type of the second media data is adapted to the user's personalized needs, thereby satisfying the user's personalized needs and improving the user experience.
[0075] For example, the first media data consists of two images uploaded by the user, with the first style being jazz. The server generates background music and lyrics in the "jazz" style based on audio parameter information matching the "jazz" style type and the first media data.
[0076] In this embodiment of the disclosure, determining the first style type of the first media data based on the analysis and processing of at least one first media data by the server includes: performing style type analysis on the input at least one first media data based on the style type determination model deployed on the server, and outputting the first style type of the first media data; or, determining the similarity attributes between the third media data and at least one first media data based on the third media data corresponding to at least one second style type deployed on the server; and determining the first style type of the first media data based on the similarity attributes.
[0077] The style type determination model can be a pre-trained model used to determine the style type of media data. For example, it could be an image recognition model used to classify the style of images or video frames, a text analysis model used to classify the style of text content, or an audio analysis model used to classify the style of audio content. The third type of media data is media data pre-stored on the server that corresponds to the second style type. The similarity attribute is used to characterize the similarity between media data; for example, the larger the similarity attribute value, the more similar the two pieces of media data are, and vice versa.
[0078] In practical applications, the first style type of the first media data can be determined using two different methods, both deployed on the server side. One method involves pre-training a style type determination model using media data samples labeled with style types. The server inputs at least one piece of first media data into the pre-trained model and outputs its style type, thus obtaining the first style type of the first media data. Another method involves using similarity analysis algorithms (such as cosine similarity, Euclidean distance, Jaccard similarity, etc.) to perform similarity feature analysis on the third media data corresponding to at least one second style type deployed on the server and each piece of first media data. This includes analyzing the similarity between features such as color, texture, and composition of images, obtaining the similarity attributes between each piece of third media data and each piece of first media data. The second style type of the third media data corresponding to the maximum similarity attribute can be used as the first style type of the first media data. Alternatively, the number of third media data belonging to the same second style type can be determined, and the second style type corresponding to the maximum number of such third media data can be used as the first style type of the first media data. Alternatively, third-party media data with similar attributes exceeding a preset threshold can be selected. Based on this selected third-party media data, the number of third-party media data belonging to the same second-style type can be determined, and the second-style type corresponding to the largest number is taken as the first-style type of the first media data. Of course, the style types determined by both methods can also be combined to determine the first-style type of the first media data. By using different methods to determine the first-style type of the first media data, not only can the accuracy of style type determination be ensured, but it also facilitates the determination of second-party media data based on audio parameter information corresponding to the first-style type and at least one first-party media data, ensuring that the style type of the second-party media data is adapted to the user's personalized needs.
[0079] Optionally, the audio parameter information includes a first model for determining the text information corresponding to the first media data, fourth media data corresponding to the text information, and the genre of the second media data. The first model can be a pre-trained model used to determine the text information corresponding to the media data. The fourth media data can be pre-defined text (such as words or phrases) used as auxiliary media data to guide the generation of text information in the second media data. The fourth media data is the media data included in the text information. The genre is related to the first genre. The genre can be a predefined music genre, such as classical, rock, jazz, electronic, etc.
[0080] In this embodiment of the disclosure, determining second media data based on at least one first media data and audio parameter information includes: inputting at least one first media data and fourth media data into a first model for text conversion processing to obtain text information corresponding to at least one first media data; performing arranging processing on the text information based on the music genre to obtain audio information corresponding to the text information; and determining the second media data based on the text information, audio information, and at least one first media data.
[0081] Specifically, a first model can be pre-trained based on media data samples labeled with text content, enabling it to generate text information matching the input media data. This trained first model is then deployed on a server. The server can input at least one piece of first media data and a fourth piece of media data from audio parameter information into the first model. The first model performs text conversion processing on the input first and fourth media data, generating text information matching at least one piece of first media data. For example, the first model could be a graph-to-text model based on image recognition and natural language generation technology, processing first media data such as images or video frames into descriptive text or lyrics to obtain text information. Alternatively, the first model could be an audio-to-text model based on speech recognition technology, converting first media data such as audio content into text information. Or, the first model could be a text-to-text model based on natural language processing technology, rewriting or expanding the first media data to obtain text information. Furthermore, the music genre and text information can be input into a pre-trained audio generation model. The audio generation model processes the text information according to the music genre, outputting audio information corresponding to the text information. Alternatively, an audio service can be invoked, and the text information can be processed and arranged according to the music genre to output audio information. The text information, audio information, and at least one first media data are integrated to obtain second media data. The technical solution provided in this disclosure combines a first model for text generation and audio generation technology to generate rich multimodal media data content, thereby enhancing the richness of the second media data while meeting users' personalized needs.
[0082] In this embodiment of the disclosure, the audio parameter information may be determined by: in response to a triggering operation of the audio generation control on the first page, displaying an audio parameter setting interface; and in response to an editing operation in the audio parameter setting interface, determining the quantity range of at least one first media data, a first model, a second style type, and a fourth media data in the audio parameter information for configuring at least one first media data.
[0083] The first page includes an audio generation control. This control is a visual tool used to trigger access to the audio parameter settings interface. The audio parameter settings interface includes at least one selectable second style type, allowing users to configure parameters related to generating second media data. That is, the second style type is the style type that the user can select, while the first style type is the style type selected by the user. A quantity range is used to limit the number of first media data items that can be configured, such as "greater than 2 but less than 5".
[0084] Specifically, an audio generation control can be pre-developed on the first page. This way, when a user triggers the audio generation control or enters its shortcut key, it's considered a triggered operation. At this point, an audio parameter setting interface can be displayed, providing the ability to edit audio parameter information. For example, the audio parameter setting interface includes a quantity range input box supporting user input or selection of a quantity range, a model selection dropdown menu supporting user selection of a first model, a style type selector supporting user selection of a style type, and a component supporting user upload or selection of fourth media data, etc. Users can edit the quantity range of the first media data, select the first model for text conversion processing, select the first style type of the first media data from a preset second style type, and edit the audio parameter information such as the input of fourth media data in the audio parameter setting interface. After completing the editing operation, a "OK" or "Save" button can be triggered. The system responds to this operation by saving the configured audio parameter information, so that the server can retrieve the audio parameter information corresponding to the first style type to determine the second media data based on the audio parameter information and the at least one first media data.
[0085] For example, see Figure 7 You can click the "Music Generation" button (i.e., the audio generation control) on the first page to display the audio parameter settings interface. A diagram of the audio parameter settings interface can be found here. Figure 8 In the audio parameter settings interface, you can configure the underlying model (e.g., first model A), select the music genre (e.g., pop), and set the fourth media data (e.g., keyword B). See also Figure 9 After setting the audio parameters, the audio parameters are synchronized to the audio resources, requesting the audio server to generate the audio resources for the second media data. The generated second media data can be retrieved and displayed via a request. Furthermore, the audio resources can be encapsulated as script logic (i.e., the chain for generating the second media data), and these audio resources can be referenced by other components that use audio files (including audio playback components and lyrics components).
[0086] The technical solution provided in this disclosure supports custom audio parameter information, which can meet different media data production needs, and facilitates dynamic adjustment of audio parameter information according to different needs, that is, adjustment of audio generation and processing logic, thereby improving the convenience and efficiency of media data production.
[0087] S230. In response to receiving an event of receiving second media data corresponding to the first media data, the second media data is displayed on the display page according to a preset format.
[0088] The technical solution provided in this disclosure sends at least one first media data to the server after displaying at least one first media data, so that the server can generate multimodal second media data corresponding to the first media data based on at least one first media data and pre-configured audio parameter information. This enhances the richness of the second media data content, reduces the processing pressure on the client, improves data processing efficiency, and improves user experience.
[0089] As an optional embodiment of the above embodiments, Figure 10 This is a timing diagram of media data generation according to embodiments of this disclosure. To further clarify the technical solutions of the embodiments of this invention for those skilled in the art, specific application scenario examples are provided. For details, please refer to the following specific content.
[0090] Users can access the effects runtime (i.e., a functional module or effects prop package) on the client side and select an image (i.e., the first media data). After selecting the image, the effects runtime saves it locally and uploads it to the resource storage server for storage. The resource storage server returns the image's URL link. The effects runtime requests the client to generate music (i.e., requests to generate the second media data). The client calls the business server's interface, which in turn calls the cloud editing server's interface. The cloud editing server reviews the image's compliance; if it fails the review, an error is reported, and a failure and error message is sent to the business server. The business server then sends the failure and error message back to the client, which in turn returns the failure and error message to the effects runtime. If the review is successful, the image is sent to the text conversion server, which performs image understanding and generates image-understood text (i.e., information within the text information). The text conversion server then sends the image-understood text to the cloud editing server. The cloud editing server requests a title from the text conversion server. The text conversion server generates the title text (i.e., information within the text information) based on image-understanding text and returns it to the cloud editing server. The cloud editing server sends the image-understood text to the audio server, which generates music (i.e., audio information) based on the image-understood text. The audio server then sends the music and lyrics, along with other media data, to the cloud editing server. The cloud editing server reviews the compliance of the title and lyrics text. If the review is successful, it sends the music data (i.e., secondary media data, including the title text, lyrics, and music) to the business server, which then sends the music data to the client. The client sends the music data to the effects runtime, which downloads the music data to the client, plays the music, and displays the lyrics.
[0091] The technical solution provided in this disclosure, by receiving at least one first media data configured in the media data generation page and displaying it on a display page, visualizes the first media data and facilitates timely confirmation by the user whether the generated first media data meets their needs. Furthermore, by displaying the text information and corresponding audio information of the second media data on the display page according to a preset format upon receiving the second media data corresponding to the first media data, the richness of the generated media data content is improved, while the convenience of media data generation is enhanced, achieving the technical effect of meeting the user's personalized needs.
[0092] Figure 11 This is a schematic diagram of the structure of a media data generation apparatus provided in an embodiment of the present disclosure, as shown below. Figure 11 As shown, the device includes: a first media data display module 310 and a second media data display module 320.
[0093] The first media data display module 310 is used to display the at least one first media data on the display page in response to receiving an operation of at least one first media data configured in the media data generation page; the second media data display module 320 is used to display the second media data on the display page in response to receiving an event of receiving second media data corresponding to the first media data, according to a preset form; wherein the second media data includes text information and audio information corresponding to the text information, and the text information is related to the media content of the first media data.
[0094] Optionally, based on the above-described apparatus, the apparatus may further include:
[0095] The media data sending first unit is used to send the at least one first media data to the server so that the server can determine the second media data corresponding to the first media data based on the at least one first media data and pre-configured audio parameter information.
[0096] Optionally, based on the above-described apparatus, the apparatus may further include:
[0097] The media data sending second unit is used to send the at least one first media data to the server, so as to determine the first style type of the first media data based on the analysis and processing of the at least one first media data by the server, wherein the first style type is a type among at least one pre-set second style type; and to determine the second media data based on the audio parameter information corresponding to the first style type retrieved by the server and the audio parameter information and the at least one first media data.
[0098] Based on the above-mentioned device, optionally, the server is used to perform style type analysis on at least one input first media data based on a style type determination model deployed on the server, and output a first style type of the first media data; or, based on the third media data corresponding to at least one second style type deployed on the server, determine the similarity attribute between the third media data and at least one first media data respectively; and determine the first style type of the first media data based on the similarity attribute.
[0099] Based on the above-mentioned device, optionally, the audio parameter information includes a first model for determining text information corresponding to the first media data, a fourth media data corresponding to the text information, and the genre type of the second media data; wherein, the fourth media data is the media data included in the text information, and the genre type is related to the first style type.
[0100] Based on the above-mentioned device, optionally, the server is configured to input the at least one first media data and the fourth media data into the first model for text conversion processing to obtain text information corresponding to the at least one first media data; perform arranging processing on the text information based on the music style to obtain audio information corresponding to the text information; and determine the second media data based on the text information, the audio information, and the at least one first media data.
[0101] Based on the above-mentioned device, optionally, the preset format includes a preset cover format and a preset text display format, and the second media data display module 320 includes:
[0102] A cover display unit is used to display at least one first media data in the second media data according to a preset cover format on the display page;
[0103] A text information display unit is used to display text information in the second media data on the display page according to the preset text display format;
[0104] The playback unit is configured to, when playing audio information corresponding to the text information in the second media data, display the first text corresponding to the played audio information in the text information in a first form.
[0105] Optionally, based on the above-described apparatus, the apparatus may further include:
[0106] The background information display unit is used to retrieve background media data for displaying the second media data according to the first style type corresponding to the second media data, so as to render and display the background information of the second media data based on the background media data.
[0107] Optionally, based on the above-described apparatus, the apparatus may further include:
[0108] The audio parameter setting interface display unit is used to respond to the trigger operation of the audio generation control on the first page and display the audio parameter setting interface;
[0109] An audio parameter information determination unit is configured to, in response to an editing operation in the audio parameter setting interface, determine the quantity range of at least one first media data, a first model, a first style type, and a fourth media data in the audio parameter information for configuring at least one first media data; wherein the audio parameter setting interface includes at least one selectable second style type.
[0110] Optionally, based on the above-described device, the second media data may be music data corresponding to the first media data.
[0111] The technical solution of this disclosure, in response to receiving at least one first media data configured on a media data generation page, displays the at least one first media data on a display page; in response to receiving an event of receiving second media data corresponding to the first media data, displays the second media data on the display page according to a preset form; wherein the second media data includes text information and audio information corresponding to the text information, and the text information is related to the media content of the first media data, it solves the problem that the content of media data production in the prior art is monotonous and difficult to meet the user's production needs. It achieves visualization of the first media data by displaying at least one first media data on the display page after receiving at least one first media data configured on the media data generation page, while also allowing users to promptly confirm whether the generated first media data meets their needs. Furthermore, by displaying the text information and corresponding audio information in the second media data on the display page according to a preset form when receiving the second media data corresponding to the first media data, it improves the richness of the generated media data content and the convenience of media data generation, achieving the technical effect of meeting the user's personalized needs.
[0112] The media data generation apparatus provided in this disclosure can execute the media data generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0113] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0114] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Figure 12 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0115] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 12 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0116] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the methods of embodiments of this disclosure.
[0117] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0118] The electronic device provided in this disclosure and the media data generation method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this disclosure can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0119] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the media data generation method provided in the above embodiments.
[0120] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0121] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0122] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0123] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0124] In response to receiving an operation that configures at least one first media data in the media data generation page, the at least one first media data is displayed in the display page;
[0125] In response to receiving an event that corresponds to the first media data, the second media data is displayed on the display page according to a preset format;
[0126] The second media data includes text information and audio information corresponding to the text information, wherein the text information is related to the media content of the first media data.
[0127] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0129] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0130] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0131] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0132] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0133] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0134] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for generating media data, characterized in that, include: In response to receiving an operation that configures at least one first media data in the media data generation page, the at least one first media data is displayed in the display page; In response to receiving an event that corresponds to the first media data, the second media data is displayed on the display page according to a preset format; The second media data includes text information and audio information corresponding to the text information, wherein the text information is related to the media content of the first media data.
2. The method according to claim 1, characterized in that, After displaying the at least one first media data, the method further includes: The at least one first media data is sent to the server so that the server can determine the second media data corresponding to the first media data based on the at least one first media data and pre-configured audio parameter information.
3. The method according to claim 1, characterized in that, After displaying the at least one first media data, the method further includes: The at least one first media data is sent to the server so that, based on the server's analysis and processing of the at least one first media data, a first style type of the first media data is determined, wherein the first style type is a type among at least one pre-defined second style type; Based on the audio parameter information corresponding to the first style type retrieved by the server, the second media data is determined based on the audio parameter information and the at least one first media data.
4. The method according to claim 3, characterized in that, The step of determining the first style type of the first media data based on the analysis and processing of the at least one first media data by the server includes: Based on a style type determination model deployed on the server, at least one input first media data is analyzed for style type, and the first style type of the first media data is output; or, Based on the third media data corresponding to at least one second style type deployed on the server, determine the similarity attributes of the third media data with at least one first media data; Based on the similarity attributes, the first style type of the first media data is determined.
5. The method according to claim 2 or 3, characterized in that, The audio parameter information includes a first model for determining text information corresponding to the first media data, a fourth media data corresponding to the text information, and the genre type of the second media data; The fourth media data refers to the media data included in the text information, and the music style type is related to the first style type.
6. The method according to claim 5, characterized in that, The second media data is determined based on at least one first media data and audio parameter information, including: The at least one first media data and the fourth media data are input into the first model for text conversion processing to obtain the text information corresponding to the at least one first media data. Based on the musical style, the text information is arranged to obtain audio information corresponding to the text information; The second media data is determined based on the text information, the audio information, and the at least one first media data.
7. The method according to claim 6, characterized in that, The preset format includes a preset cover format and a preset text display format. Displaying the second media data on the display page according to the preset format includes: At least one of the first media data in the second media data is displayed on the display page according to a preset cover format; The text information in the second media data is displayed on the display page according to the preset text display format; When playing audio information corresponding to the text information in the second media data, the first text corresponding to the played audio information is displayed in the text information in a first form.
8. The method according to claim 1, characterized in that, The method further includes: Based on the first style type corresponding to the second media data, the background media data for displaying the second media data is retrieved, and the background information of the second media data is rendered and displayed based on the background media data.
9. The method according to claim 2 or 3, characterized in that, Also includes: In response to the triggering operation of the audio generation control on the first page, the audio parameter setting interface is displayed; In response to an editing operation in the audio parameter setting interface, the quantity range of at least one first media data, a first model, a first style type, and a fourth media data are determined in the audio parameter information; wherein the audio parameter setting interface includes at least one selectable second style type.
10. The method according to claim 1, characterized in that, The second media data is music data corresponding to the first media data.
11. A media data generation device, characterized in that, include: The first media data display module is used to respond to receiving an operation that configures at least one first media data in the media data generation page, and to display the at least one first media data in the display page; The second media data display module is used to respond to an event of receiving second media data corresponding to the first media data and to display the second media data on the display page according to a preset format. The second media data includes text information and audio information corresponding to the text information, wherein the text information is related to the media content of the first media data.
12. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the media data generation method as described in any one of claims 1-10.
13. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the media data generation method as described in any one of claims 1-10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the media data generation method as described in any one of claims 1-10.