Multimedia content generation method and device, electronic equipment, medium and program product
By understanding multiple attribute dimensions of the images input by the user, generating music description text and determining music, the problem of mismatch between music and images in the prior art is solved, efficient multimedia content generation is achieved, and presentation effect and user experience are improved.
Patent Information
- Application Number
- CN202510252607.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-16
AI Technical Summary
It is difficult to generate music videos that match the user input image. The existing music is added by default and has little correlation with the image content, and the user operation is complex and inefficient.
By understanding multiple attribute dimensions of the images input by the user, music description text is generated, and music is then determined, and multimedia content is generated based on the images and music.
The generated music matches the image, the presentation and audio-visual effects of multimedia content are improved, user needs are met, and generation efficiency is also improved.
Smart Images

Figure CN120017930A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and computer technology, and in particular to a method, device, electronic device, medium and program product for generating multimedia content. Background Art
[0002] With the development of artificial intelligence (AI) technology, generative models are constantly updated and iterated. Currently, users can easily transform creativity into generative content through some applications based on generative models. For example, users only need to enter a simple text description in the application to generate an image with one click, or users can enter an image in the application to generate a video. Summary of the invention
[0003] According to some embodiments of the present disclosure, a method for generating multimedia content is provided, comprising: receiving an image input by a user; understanding the image from multiple attribute dimensions and generating a music description text, wherein the music description text includes description information of generation characteristics of the music; determining music based on the music description text; generating multimedia content based on the image and the music, and displaying the multimedia content.
[0004] According to some other embodiments of the present disclosure, there is provided a device for generating multimedia content, comprising: a receiving module, configured to receive an image input by a user; a first generating module, configured to understand the image in multiple attribute dimensions and generate a music description text, wherein the music description text includes description information of generation features of the music; a determining module, configured to determine music based on the music description text; a second generating module, configured to generate multimedia content based on the image and the music; and a display module, configured to display the multimedia content.
[0005] According to some further embodiments of the present disclosure, an electronic device is provided, comprising: a processor; and a memory coupled to the processor, for storing instructions, which, when executed by the processor, causes the processor to execute a method for generating multimedia content as in any embodiment of the present disclosure.
[0006] According to some further embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for generating multimedia content of any embodiment of the present disclosure is implemented.
[0007] According to some further embodiments of the present disclosure, a computer program product is provided, comprising: instructions, which, when executed by a processor, cause the processor to execute a method for generating multimedia content as described in any embodiment of the present disclosure.
[0008] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The following is an explanation of the embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings:
[0010] Figure 1 A schematic diagram showing a flow chart of a method for generating multimedia content according to some embodiments of the present disclosure;
[0011] Figure 2 A schematic diagram showing an interactive interface of some embodiments of the present disclosure;
[0012] Figure 3 Schematic diagrams showing interaction interfaces of other embodiments of the present disclosure;
[0013] Figure 4 Schematic diagrams showing interaction interfaces of yet other embodiments of the present disclosure;
[0014] Figure 5 A schematic diagram showing the structure of a device for generating multimedia content according to some embodiments of the present disclosure;
[0015] Figure 6 A schematic diagram showing the structure of an electronic device according to some embodiments of the present disclosure;
[0016] Figure 7 A schematic diagram showing the structure of an electronic device according to some other embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein.
[0018] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of the steps set forth in these embodiments should be interpreted as being merely exemplary and not limiting the scope of the present disclosure.
[0019] The term “including” and its variations used in the present disclosure are intended to be open terms that include at least the following elements / features but do not exclude other elements / features, that is, “including but not limited to.” The term “based on” means “at least partly based on.”
[0020] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units. Unless otherwise specified, the concepts of "first", "second", etc. are not intended to imply that the objects described in this way must be in a given order in time, space, ranking, or any other manner.
[0021] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] It should be understood that the present disclosure does not limit how to obtain the image to be applied / processed. In some embodiments of the present disclosure, it can be obtained from a storage device, such as an internal memory or an external storage device. In other embodiments of the present disclosure, a photographic component can be mobilized to shoot. It should be noted that the acquired image can be a captured image or a frame of an image in a captured video, and is not particularly limited to this.
[0024] In the context of the present disclosure, an image may refer to any of a variety of images, such as a color image, a grayscale image, etc. It should be noted that in the context of the present specification, the type of image is not specifically limited. In addition, the image may be any appropriate image, such as an original image obtained by a camera device, or an image that has been subjected to specific processing, such as preliminary filtering, anti-aliasing, color adjustment, contrast adjustment, normalization, etc. It should be noted that the preprocessing operation may also include other types of preprocessing operations known in the art, which will not be described in detail here.
[0025] The embodiments of the present disclosure are described in detail below in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In addition, in one or more embodiments, specific features, structures or characteristics can be combined in any suitable manner that will be clear from the present disclosure by a person of ordinary skill in the art.
[0026] Currently, some applications or agents can provide the function of generating videos based on images input by users. The generated videos usually do not have music added, and even if music is added, it is some default background music.
[0027] Music in a video can enhance emotional resonance, create an atmosphere, and make the overall effect of the video better. In many cases, users want to generate videos with music, especially in some cases, users want to generate music videos based on images. However, the default background music added to the video is usually not closely related to the content of the image or video, and the effect is not good and does not meet the needs of users. If users add music and edit videos by themselves, the operation is complicated and inefficient.
[0028] Based on the above reasons, the present disclosure proposes a method for generating multimedia content, which can understand multiple attribute dimensions based on the image input by the user, generate a music description text, and then determine the music based on the music description text, and then generate multimedia content based on the image and the music. Since the music description text is generated based on the understanding of multiple attribute dimensions of the image, the generation characteristics of the music described in the music description text are associated with the multiple attribute dimensions of the image, so that the generated music matches the image, and the multimedia content finally generated is also matched with the music, which improves the presentation effect and audio-visual effect of the multimedia content as a whole, and better meets the needs of users. The method for generating multimedia content disclosed in the present disclosure can directly determine music and generate multimedia content based on images, without the need for users to select and add music, while improving the presentation effect and audio-visual effect of the multimedia content, and improving the efficiency of generating multimedia content.
[0029] Combine the following Figures 1 to 4 The method for generating multimedia content of the present disclosure is described. The method for generating multimedia content of the present disclosure can be executed by at least one of an application, an agent, a robot (Bot), a device for generating multimedia content, and an electronic device, but is not limited to the examples given.
[0030] Figure 1 Flowcharts of some embodiments of the method for generating multimedia content disclosed herein. Figure 1 As shown, the method of this embodiment includes: steps S102 to S108.
[0031] In step S102, an image input by a user is received.
[0032] For example, the user can input an image in the interactive interface, which may be an interactive interface provided by an application, an intelligent agent, a robot, a multimedia content generation device, an electronic device, etc. The image may be obtained from an internal storage device or an external storage device of an electronic device (such as a terminal) used by the user, or may be taken by calling a shooting component in the electronic device used by the user, without limitation to the examples given.
[0033] In step S104, the image is understood from multiple attribute dimensions to generate a music description text, wherein the music description text includes description information of the generated features of the music.
[0034] For example, a machine learning model can be used to understand an image in multiple attribute dimensions to generate a music description text. The attributes of an image can be various properties and characteristics that can describe and characterize the characteristics of the image. Understanding an image in multiple attribute dimensions can be to use a machine learning model to analyze and understand the image from multiple attribute dimensions, or to analyze and understand multiple attributes of the image. For example, the multiple attribute dimensions may include at least two of entity, scene, emotion, environment, theme, style, and effect elements, and are not limited to the examples given. The multiple attribute dimensions may be pre-configured, or they may be multiple attribute dimensions that need to be understood by the machine learning model based on the task of generating a music description text.
[0035] By understanding the image in multiple attribute dimensions, some features of the music can be generated to obtain the generated features of the music. Then, based on the generated features of the music, the description text of the music is determined, and the music description text can be used to describe the features (generated features) of the music to be determined later. For example, the generated features of music include features corresponding to at least one feature type of atmosphere, style, rhythm, melody, harmony, timbre, and vocals, and are not limited to the examples given. The feature type to which the generated features of music belong can be pre-set, or it can be the feature type required for generating music determined by the subsequent task of generating music by the machine learning model.
[0036] The music may have lyrics or not, which can be selected by the user. In the case of music with lyrics, the image is understood in multiple attribute dimensions to generate the lyrics of the music, and the music description text is generated according to the lyrics of the music and the generated features of the music, that is, the music description text can also include the lyrics of the music.
[0037] By understanding the image in multiple attribute dimensions and generating descriptive information and lyrics of the music's features, the subsequently determined music can be closely associated with the image and match each other, allowing the music to express the information and emotions that the image wants to convey, better meeting user needs.
[0038] In step S106, the music is determined according to the music description text.
[0039] Since the music description text includes description information of the generated features of the music, the music can be directly generated using a machine learning model based on the description information, or music that matches the generated features of the music can be selected from a music database based on the description information.
[0040] In step S108 , multimedia content is generated based on the image and the music, and the multimedia content is displayed.
[0041] For example, a machine learning model can be used to generate multiple video frames based on an image, and then the multiple video frames can be edited with music to align the multiple video frames and music in terms of time and content to generate the final multimedia content. The multimedia content can be displayed in the interactive interface between the user and the application or intelligent agent.
[0042] The method of the above embodiment can understand multiple attribute dimensions based on the image input by the user, generate a music description text, and then determine the music based on the music description text, and then generate multimedia content based on the image and the music. Since the music description text is generated based on the understanding of multiple attribute dimensions of the image, the generation characteristics of the music described in the music description text are associated with the multiple attribute dimensions of the image, so that the generated music matches the image, and the multimedia content finally generated also matches the music, which improves the presentation effect and audio-visual effect of the multimedia content as a whole and better meets the needs of users. The method of the above embodiment can directly determine the music and generate multimedia content based on the image, without the need for the user to select and add music, which improves the presentation effect and audio-visual effect of the multimedia content while improving the efficiency of generating multimedia content.
[0043] How to receive an image input by a user is described below in conjunction with some embodiments.
[0044] In some embodiments, receiving an image input by a user includes: displaying a camera control; in response to the user triggering the camera control, displaying a shooting preview interface, wherein the shooting preview interface includes a multimedia content generation option; in response to the user selecting the multimedia content generation option and triggering the capture of the image, receiving the image.
[0045] For example, a camera control is displayed in the interactive interface, and when a user triggers the camera control in the interactive interface, a shooting preview interface can be displayed, and the picture captured by the camera lens can be displayed in the shooting preview interface. Figure 2 As shown, it is a shooting preview interface, which can display the picture 201 captured by the camera lens, and can also display the multimedia content generation option 202. In addition, other options such as "take a photo" can also be displayed. For example, the user can select the multimedia content generation option 202 by sliding, clicking, etc. The shooting preview interface also includes a shooting control 203. In response to the user triggering the shooting control 203, an image is shot and the shot image is obtained.
[0046] In response to the user triggering the shooting control 203, a confirmation control may also be displayed, for example, the shooting control may be switched to display as a confirmation control. In response to the user triggering the confirmation control, an image may be received and a subsequent multimedia content generation process may be executed. In response to the user triggering the confirmation control, an image may be displayed in an interactive interface (or a new interface, window, floating layer, mask, etc.) and a multimedia content generation control may be displayed. In response to the user triggering the multimedia content generation control, a subsequent multimedia content generation process may be executed.
[0047] Of course, the multimedia content generation control can be displayed in the interactive interface, or the shooting preview interface can be displayed by the user inputting the multimedia content generation instruction. Only one option of the multimedia content generation option can be displayed in the shooting preview interface, or only the shooting control can be displayed. The user triggers the shooting control to trigger the generation of multimedia content based on the image. The specific control setting method and triggering method can be set according to the needs, and are not limited to the above examples.
[0048] Based on the method of the above embodiment, users can directly take photos and automatically generate multimedia content, and the multimedia content includes music generated based on the taken photos, realizing the photo-to-music conversion, improving the efficiency of generating multimedia content, and the images and music in the generated multimedia content are more matched, meeting the needs of users.
[0049] In some embodiments, receiving an image input by a user includes: displaying an image selection control; displaying multiple candidate images in response to the user triggering the image selection control; receiving and displaying the image and a multimedia content generation control in response to the user selecting an image from the multiple candidate images; and receiving a user triggering operation on the multimedia content generation control.
[0050] For example, you can display an image selection control in an interactive interface, or Figure 2 The preview interface shown shows an image selection control 204. When a user triggers the image selection control, multiple candidate images stored locally can be displayed, or there can be only one candidate image. The user can select an image by clicking or other operations, and the selected image can be displayed, and a multimedia content generation control can be displayed. Then, the user triggers the multimedia content generation control to execute a subsequent generation process of generating multimedia content.
[0051] In response to the user selecting an image, a confirmation control may also be displayed, for example, the shooting control may be switched to display as a confirmation control. In response to the user triggering the confirmation control, the image is received, the image is displayed in the interactive interface (or new interface, window, floating layer, mask, etc.), and the multimedia content generation control is displayed. In response to the user triggering the multimedia content generation control, the subsequent multimedia content generation process is executed. The specific control setting method and triggering method can be set according to the needs and are not limited to the above examples.
[0052] The image captured or selected by the user may be one or more images, that is, the image input by the user may be one or more images.
[0053] Based on the method of the above embodiment, the user can automatically generate multimedia content by selecting an existing image, and the multimedia content includes music generated based on the taken photos, which improves the efficiency of generating multimedia content, and the images and music in the generated multimedia content are more matched, meeting the needs of the user.
[0054] The following describes in detail how to understand images from multiple attribute dimensions and generate music description text.
[0055] In some embodiments, understanding an image from multiple attribute dimensions and generating a music description text includes: understanding an image from multiple attribute dimensions and generating a description text of the image, wherein the description text of the image includes information from multiple attribute dimensions; performing semantic understanding on the description text of the image and generating a music description text, wherein the generated features of the music are obtained based on the information from multiple attribute dimensions.
[0056] For example, the first machine learning model is used to understand the image in multiple attribute dimensions to generate a description text of the image. For example, the first machine learning model is a picture-to-text model, etc., not limited to the examples cited. For example, the second machine learning model is used to semantically understand the description text of the image to generate a music description text. For example, the second machine learning model is a large language model (LLM), etc., not limited to the examples cited. Since the generated features of music are obtained based on information of multiple attribute dimensions, the generated features of music are closely related to and match the information of multiple attribute dimensions of the image.
[0057] The method of the above embodiment has a deep understanding of the image and converts the information of multiple attribute dimensions of the image into the generation features of the music, thereby realizing the conversion from the image to the music description text, improving the accuracy of the generation features of the music, and thus improving the accuracy of the generated music and the degree of matching with the image.
[0058] In some embodiments, understanding an image from multiple attribute dimensions and generating a description text for the image includes: understanding the image according to multiple attribute dimensions of entities, scenes, emotions, environments, themes, styles, and effect elements, and determining information on multiple attribute dimensions of entities, scenes, emotions, environments, themes, styles, and effect elements in the image; and generating a description text for the image based on the information on multiple attribute dimensions.
[0059] The information of multiple attribute dimensions includes information of at least two attribute dimensions of entity, scene, emotion, environment, theme, style, and effect element. For example, it is determined that the entities in the image include the sun, high-rise buildings, and roads, the scene is a city street, the emotion is loneliness, the environment is evening, the theme is a city at sunset, the style is urban, and the effect element is cold colors, not limited to the above examples. The effect element can be used to represent the effects added to the image, and may include color-related effects, such as brightness, contrast, etc., may include light and shadow-related effects, such as shadows, glow, etc., may include clarity-related effects, such as blurring, sharpening, etc., may include filters, overlays, inversions, and other effects, not limited to the examples given.
[0060] The description text of the image can be generated by structured processing based on information of multiple attribute dimensions, or can be obtained by integrating information of multiple attribute dimensions based on a specific template. Information of multiple attribute dimensions can be extracted using a machine learning model. Therefore, multiple attribute dimensions can be determined by learning of a machine learning model and can include other attribute dimensions other than the above examples, such as the atmosphere of the image, the relationship between entities, etc., but are not limited to the above examples.
[0061] The method of the above embodiment can accurately extract information of multiple attribute dimensions of an image, thereby generating a more accurate description text of the image and improving the accuracy of subsequent music description text.
[0062] In some embodiments, semantic understanding is performed on the descriptive text of an image to generate a music description text, including: semantic understanding is performed on the descriptive text of the image to determine information of multiple attribute dimensions in entities, scenes, emotions, environments, themes, styles, and effect elements in the image; determining generated features of the music that match the information of multiple attribute dimensions in the entities, scenes, emotions, environments, themes, styles, and effect elements, wherein the generated features of the music include features corresponding to at least one feature type of the atmosphere, style, rhythm, melody, harmony, timbre, and vocals of the music; and generating a descriptive text for the music based on the generated features of the music.
[0063] The second machine learning model can be used to determine the generated features of the music that match the information of the multiple attribute dimensions based on the information of the multiple attribute dimensions of the image. For example, the features corresponding to at least one feature type of the atmosphere and style of the music can be first determined based on the information of the multiple attribute dimensions, for example, the atmosphere of the music is determined to be cheerful and the style is classical, and then the features corresponding to at least one feature type of the rhythm, melody, harmony, and vocals of the music can be determined based on the atmosphere and style.
[0064] After determining the generated features of the music, structured processing can be performed to generate music description text, or the generated features of the music can be filled into a corresponding template to obtain the music description text, or the music description text can be generated in the form of rich text.
[0065] The method of the above embodiment can determine the generation features of matching music based on multiple attribute dimensions of the image, improve the accuracy of the generation features of the music, and further improve the matching degree between the subsequently generated music and the image and multimedia content.
[0066] The generated music may include lyrics. The following describes how to generate a description text of the music when the music includes lyrics in conjunction with some embodiments.
[0067] In some embodiments, the music description text also includes the lyrics of the music. Generating the description text of the music based on the generation characteristics of the music includes: generating the lyrics of the music based on information of multiple attribute dimensions in the entities, scenes, emotions, environments, themes, styles, and effect elements in the image, as well as the generation characteristics of the music; generating the music description text based on the generation characteristics of the music and the lyrics of the music.
[0068] The machine learning model can be used to generate music lyrics based on the information of multiple attribute dimensions of the image and the generated features of the music. The generated music lyrics match the information of multiple attribute dimensions of the image and the generated features of the music, which can better express the characteristics of the image such as emotions and themes, and can also make the generated music more expressive and improve the auditory effect. The generated features of the music and the lyrics of the music can be structured or filled in with corresponding templates to obtain music description text.
[0069] In some embodiments, generating lyrics of music based on information of multiple attribute dimensions in entities, scenes, emotions, environments, themes, styles, and effect elements in an image, as well as generation characteristics of the music, includes: determining the theme, emotion, and keywords of the lyrics based on information of multiple attribute dimensions in entities, scenes, emotions, environments, themes, styles, and effect elements in the image, as well as generation characteristics of the music; generating lyrics of the music based on the theme, emotion, and keywords of the lyrics.
[0070] The theme, emotion and keywords of the lyrics can be determined based on the information of multiple attribute dimensions of the image and the generated features of the music. Furthermore, the lyrics of the music can be generated based on the theme, emotion and keywords of the lyrics and the generated features of the music. The generated lyrics of the music can match the rhythm, melody, style, duration, structure, etc. of the music.
[0071] The method of the above embodiment can improve the accuracy of the generated lyrics and can make the generated lyrics more closely match the information of multiple attribute dimensions of the image and the generated features of the music, thereby improving the audio-visual effect of the subsequently generated multimedia content.
[0072] In the case where the music includes lyrics, the lyrics of the music may be generated first, and then the generation features of the music may be generated to generate the music description text. In some embodiments, the lyrics of the music are generated based on information of multiple attribute dimensions of the image, the generation features of the music are generated based on information of multiple attribute dimensions of the image and the lyrics of the music, and the music description text is generated based on the lyrics of the music and the generation features.
[0073] For example, the theme, emotion and keywords of the lyrics are determined based on the information of multiple attribute dimensions of the image, the lyrics are generated based on the theme, emotion and keywords of the lyrics, and the generation features of the music are generated based on the information of multiple attribute dimensions of the image and the lyrics.
[0074] Whether the music includes lyrics can be configured by the user. In some embodiments, a lyrics generation control is displayed; in response to the user triggering the lyrics generation control, the image is understood in multiple attribute dimensions to generate lyrics for the music.
[0075] After receiving the image, the image and the lyrics generation control (or only the lyrics generation control) can be displayed in the interactive interface (or new interface, window, floating layer, mask, etc.). The user triggers the lyrics generation control to trigger the function of generating the lyrics of the music. For specific information on how to generate the lyrics, please refer to the aforementioned embodiment and will not be repeated here. The user can also directly input the prompt information for generating the lyrics to trigger the function of generating the lyrics of the music. For example, the image is displayed in the interactive interface, and the input area can also be displayed. In response to the prompt information for generating the lyrics input by the user, the image is understood in multiple attribute dimensions to generate the lyrics of the music.
[0076] According to the method of the above embodiment, the user can trigger the generation of music lyrics by simply triggering the lyrics generation control, without the need for the user to input lyrics prompt information, thereby improving the efficiency of lyrics generation and the convenience of operation.
[0077] Whether the music includes lyrics can be configured by the user, and other features of the music can also be configured by the user. The following is described in conjunction with some embodiments.
[0078] In some embodiments, understanding an image from multiple attribute dimensions and generating music description text includes: displaying controls corresponding to one or more candidate features, wherein, when displaying controls corresponding to multiple candidate features, the multiple candidate features correspond to one or more feature types; in response to a user triggering an operation on a control of one or more target candidate features among the one or more candidate features, generating music description text based on the one or more target candidate features and understanding the image from multiple attribute dimensions.
[0079] For example, controls for displaying one or more candidate features in various forms of interfaces such as interactive interfaces, new interfaces, windows, floating layers, and masks are not limited to the examples given. For example, after receiving an image, the image and controls corresponding to one or more candidate features are displayed on a new interface. For another example, after receiving an image, the image and controls corresponding to one or more feature types are displayed on a new interface, and in response to a user triggering a control corresponding to a target feature type, one or more candidate features corresponding to the target feature type or controls corresponding to one or more candidate features are displayed. In response to a user's selection operation on one or more target candidate features, or a triggering operation on controls of one or more target candidate features, a music description text is generated based on one or more target candidate features and an understanding of multiple attribute dimensions of the image.
[0080] like Figure 3 As shown, an image 308 is displayed in the interface, and controls 301-303 of multiple candidate features can also be displayed. The controls 301-303 of the candidate features can correspond to the same feature type, for example, style. The user can select the style of music as rock, folk or jazz by triggering controls 301-303. The controls 301-303 of the candidate features can be controls of some candidate features. For example, in response to a preset operation of the user (for example, a sliding operation along a specific direction), controls of other candidate features can be displayed.
[0081] like Figure 3 As shown, multiple feature type controls 305-306 may be displayed in the interface, and the currently selected target candidate feature may be displayed in the control of each feature type. For example, the currently selected target candidate feature is displayed as "Cheerful" in the control 305 corresponding to the feature type of atmosphere. In response to the user triggering the control of the target feature type, one or more candidate features under the target feature type may be displayed in the form of a drop-down menu, and the user may select the target candidate feature. Multiple feature type controls 305-306 may only be controls of some feature types. For example, the feature type controls may be displayed in response to a preset operation of the user (for example, a sliding operation along a specific direction).
[0082] like Figure 3As shown, the lyrics generation control 307 may be displayed in the interface, and in response to the user triggering the lyrics generation control 307, the options of having lyrics and not having lyrics may be displayed, and in response to the user selecting the option of having lyrics, "having lyrics" may be displayed in the lyrics generation control 307. In addition, the lyrics generation control may be in the form of a switch or other forms, and is not limited to the examples given.
[0083] like Figure 3 As shown, a template control 309 may also be displayed in the interface. In response to the user triggering the template control 309, one or more templates may be displayed, each template may include one or more features that have been determined. In response to the user selecting a target template, a music description text is generated based on the one or more features corresponding to the target template and the understanding of multiple attribute dimensions of the image.
[0084] like Figure 3 As shown, a delete control 310 may also be displayed in the image display area. In response to the user triggering the delete control 310, an image adding control may be displayed. In response to the user triggering the image adding control, one or more candidate images may be displayed or a camera control may be displayed, etc. The user may re-upload or take an image.
[0085] In the above embodiments, the display method of the controls corresponding to the candidate features and feature types is only an example, and other settings may be adopted in actual use.
[0086] For example, one or more reference features of music can be determined based on the understanding of multiple attribute dimensions of the image, and the one or more reference features can be displayed in the interface as one or more candidate features. One or more reference features can also be displayed in a specific position. For example, in the control of the feature type, the reference feature corresponding to the feature type is displayed by default, and the reference feature is displayed at the first place in the sorting, etc.
[0087] Since one or more reference features obtained based on the understanding of multiple attribute dimensions of the image are more in line with user needs, displaying these reference features at specific locations can enable users to determine target candidate features more quickly and efficiently, thereby improving the efficiency of multimedia content generation.
[0088] One or more target candidate features selected by the user may be only partial features for generating the music description text. Combined with the understanding of multiple attribute dimensions of the image, another partial feature may be generated to further generate the music description text.
[0089] The method of the above embodiment allows the user to select one or more target candidate features, combined with the understanding of multiple attribute dimensions of the image, to generate a music description text, so that the generated features of the music in the music description text are more in line with the user's needs, so that the subsequently generated multimedia content can bring a better audio-visual experience to the user.
[0090] One or more target candidate features selected by the user may be inconsistent or conflicting with the features of the music obtained by understanding the image, and judgment and processing are required for this situation.
[0091] In some embodiments, generating music description text based on one or more target candidate features and understanding of multiple attribute dimensions of an image includes: determining one or more generated features of the music based on understanding of multiple attribute dimensions of the image; matching the one or more target candidate features with the one or more generated features to determine whether there are generated features that conflict with the one or more target candidate features; in response to the existence of generated features that conflict with the one or more target candidate features, deleting the generated features that conflict with the one or more target candidate features, and generating music description text based on the one or more target candidate features and the remaining generated features.
[0092] The target candidate feature and the generated feature of the same feature type can be matched. If they are inconsistent, the target candidate feature shall prevail. In the case where the target candidate feature and the generated feature are of different feature types, if there is a contradiction between the two, the generated feature shall be deleted and regenerated based on the target candidate feature.
[0093] Based on the understanding of multiple attribute dimensions of the image, one or more generated features of the music can be determined by referring to the aforementioned embodiments, which will not be described in detail here.
[0094] If the target candidate features selected by the user meet the features required for generating music description text or generating music, it is no longer necessary to understand the image in multiple attribute dimensions.
[0095] The method of the above embodiment can refer to one or more target candidate features selected by the user and understand the image in multiple attribute dimensions, while avoiding contradictions between the two, which may cause the generated music or multimedia content to be inaccurate or confused, thereby improving the accuracy of the generated music or multimedia content and enhancing the audio-visual effect.
[0096] In addition to the above embodiments, some controls may be configured for the user to operate and select target candidate features, the user may also express the need for the generated music or multimedia content by inputting prompt information.
[0097] In some embodiments, prompt information of multimedia content input by a user is received; and music description text is generated based on the understanding of multiple attribute dimensions of the image and the prompt information.
[0098] For example, after receiving the image, the image can be displayed in the interactive interface, and an input area can also be displayed in the interactive interface, and the user can enter prompt information in the input area. Figure 3 As shown, an image is displayed in the new interface, and an input area 311 may also be displayed, and the user may input prompt information in the input area 311 in the form of voice, text, etc.
[0099] The method of the above embodiment can generate a music description text by the user inputting prompt information, combined with the understanding of multiple attribute dimensions of the image, so that the generation characteristics of the music in the music description text are more in line with the needs of the user, so that the subsequently generated multimedia content can bring a better audio-visual experience to the user.
[0100] In some embodiments, an image is understood from multiple attribute dimensions to generate a description text of the image, wherein the description text of the image includes information from multiple attribute dimensions; prompt information is semantically understood to determine whether the prompt information includes a theme of the multimedia content; in response to the prompt information including the theme of the multimedia content, a music description text is generated based on the description text of the image and the theme of the multimedia content.
[0101] For example, based on the information of multiple attribute dimensions in the entities, scenes, emotions, environments, themes, styles, and effect elements in the image and the theme of the multimedia content, the features corresponding to at least one feature type of the music's atmosphere, style, rhythm, melody, harmony, timbre, and vocals are determined, and then the music description text is generated.
[0102] The method of the above embodiment, when the user inputs prompt information, combines the understanding of multiple attribute dimensions of the theme of the multimedia content and the image in the prompt information to generate a music description text, so that the generation characteristics of the music in the music description text are more in line with the user's needs, so that the subsequently generated multimedia content can bring a better audio-visual experience to the user.
[0103] In some embodiments, understanding the image in multiple attribute dimensions and providing prompt information to generate music description text also includes: semantically understanding the prompt information to determine whether the prompt information includes target features of the music; in response to the prompt information including the target features of the music, generating the music description text based on the descriptive text of the image, the theme of the multimedia content and the target features of the music.
[0104] If the prompt information input by the user includes the target feature of the music, the music description text is generated by combining the description text of the image, the theme of the multimedia content and the target feature of the music. For example, the generated features of the music are determined based on the description text of the image and the theme of the multimedia content, the generated features of the music are matched with the target features, and it is determined whether there are generated features that conflict with the target features. In response to the presence of generated features that conflict with the target features, the generated features that conflict with the target features are deleted, and the music description text is generated based on the target features and the remaining generated features. If the prompt information only includes the target feature of the music, the music description text can be generated based on the description text of the image and the target feature.
[0105] In the above embodiment, the target features of the music in the prompt information, the theme of the multimedia content and the descriptive text of the image are combined to generate a music description text, so that the generation features of the music in the music description text are more in line with the needs of the user, so that the subsequently generated multimedia content can bring a better audio-visual experience to the user.
[0106] The following describes how to generate multimedia content from images and music.
[0107] In some embodiments, generating multimedia content based on images and music includes: understanding the image in multiple attribute dimensions and determining the camera movement method; generating multiple frames of video images based on the image and the camera movement method; generating multimedia content based on the multiple frames of video images and music.
[0108] A machine learning model can be used to generate multiple video frames. For example, the machine learning model can be an image-generated video model. In addition to determining the camera movement mode, multiple scene types can also be determined, such as close-up, long shot, medium shot, etc. Multiple video frames can be generated based on the image, scene type, and camera movement mode. The multiple video frames and music are further edited and integrated to obtain multimedia content.
[0109] The method of the above embodiment determines the camera movement mode based on the understanding of multiple attribute dimensions of the image, and then generates multiple frames of video images and multimedia content, which can make the multimedia content more fluent and rich in expression and improve the audio-visual effect of the multimedia content.
[0110] In some embodiments, generating multimedia content based on images and music includes: generating descriptive text for the image based on understanding multiple attribute dimensions of the image; expanding the descriptive text for the image to generate descriptive text for multiple frames of video images; generating multiple frames of video images based on the descriptive text for the multiple frames of video images; generating multimedia content based on multiple frames of video images and music.
[0111] The description text of the image can be expanded using a machine learning model (e.g., LLM) to generate description text of multiple video frames, and then multiple video frames can be generated. The description text of the multiple video frames can include the content, scene, and camera movement of each frame.
[0112] The method of the above embodiment can generate a description text of a multi-frame video image with richer content by expanding the description text of the image, thereby enriching the content of the generated multimedia content and improving the audio-visual effect.
[0113] When the user inputs prompt information including the theme of multimedia content or other requirements, multiple frames of video images are generated based on the prompt information and the image. For example, the image is understood in multiple attribute dimensions and the prompt information is used to determine the camera movement method; multiple frames of video images are generated based on the image, the camera movement method and the prompt information; multimedia content is generated based on the multiple frames of video images and music. For another example, the image description text is generated based on the understanding of the image in multiple attribute dimensions and the prompt information; the image description text is expanded to generate the description text of the multiple frames of video images; the multiple frames of video images are generated based on the description text of the multiple frames of video images; multimedia content is generated based on the multiple frames of video images and music.
[0114] In the process of generating multimedia content based on multi-frame video images and music, it is necessary to match the duration of the multi-frame video images with the music, and to match the content of the multi-frame video images with the content of the lyrics of the music to achieve better audio-visual effects of the multimedia content.
[0115] like Figure 4 As shown, after the multimedia content is generated, a preview interface of the multimedia content can be displayed. A play control 401 can be displayed in the preview interface. In response to the user triggering the play control 401, the multimedia content is played. A publishing control 402 can also be displayed in the preview interface. In response to the user triggering the publishing control 402, one or more publishing channels can be displayed. After the user selects a publishing channel, the multimedia content can be published through the selected publishing channel. A download control 403 and a share control 404 can also be displayed in the preview interface. The user can download and store the multimedia content and share the multimedia content.
[0116] The present disclosure also provides a device for generating multimedia content. Figure 5 Give a description.
[0117] Figure 5 FIG. 1 is a structural diagram of some embodiments of the apparatus for generating multimedia content disclosed in the present invention. Figure 5As shown, the multimedia content generation device 50 of this embodiment includes: a receiving module 510 , a first generation module 520 , a determination module 530 , a second generation module 540 , and a display module 550 .
[0118] The receiving module 510 is configured to receive an image input by a user.
[0119] The first generation module 520 is configured to understand the image from multiple attribute dimensions and generate a music description text, wherein the music description text includes description information of the generated features of the music.
[0120] The determination module 530 is configured to determine music according to the music description text.
[0121] The second generating module 540 is configured to generate multimedia contents according to the images and music.
[0122] The display module 550 is configured to display multimedia contents.
[0123] The multimedia content generation device of the above embodiment can understand multiple attribute dimensions based on the image input by the user, generate a music description text, and then determine the music based on the music description text, and then generate multimedia content based on the image and the music. Since the music description text is generated based on the understanding of multiple attribute dimensions of the image, the generation characteristics of the music described in the music description text are associated with the multiple attribute dimensions of the image, so that the generated music matches the image, and the final generated multimedia content matches the music, which improves the presentation effect and audio-visual effect of the multimedia content as a whole and better meets the needs of users. The generation device of the above embodiment can directly determine the music and generate multimedia content based on the image, without the need for the user to select and add music, which improves the presentation effect and audio-visual effect of the multimedia content while improving the efficiency of generating multimedia content.
[0124] In some embodiments, the first generation module 520 is configured to understand the image from multiple attribute dimensions and generate a description text of the image, wherein the description text of the image includes information from multiple attribute dimensions; perform semantic understanding on the description text of the image and generate a music description text, wherein the generation features of the music are obtained based on the information from multiple attribute dimensions.
[0125] In some embodiments, the first generation module 520 is configured to understand the image according to multiple attribute dimensions of entities, scenes, emotions, environments, themes, styles, and effect elements, determine information of multiple attribute dimensions of entities, scenes, emotions, environments, themes, styles, and effect elements in the image; and generate a description text of the image based on the information of multiple attribute dimensions.
[0126] In some embodiments, the first generation module 520 is configured to perform semantic understanding on the descriptive text of the image, determine information on multiple attribute dimensions of entities, scenes, emotions, environments, themes, styles, and effect elements in the image; determine generation features of music that match the information on multiple attribute dimensions of entities, scenes, emotions, environments, themes, styles, and effect elements, wherein the generation features of music include features corresponding to at least one feature type of the atmosphere, style, rhythm, melody, harmony, timbre, and vocals of the music; and generate a descriptive text for the music based on the generation features of the music.
[0127] In some embodiments, the music description text also includes the lyrics of the music, and the first generation module 520 is configured to generate the lyrics of the music based on the information of multiple attribute dimensions in the entities, scenes, emotions, environments, themes, styles, and effect elements in the image, as well as the generation characteristics of the music; and generate the music description text based on the generation characteristics of the music and the lyrics of the music.
[0128] In some embodiments, the first generation module 520 is configured to determine the theme, emotion and keywords of the lyrics based on information of multiple attribute dimensions in the entities, scenes, emotions, environments, themes, styles, effect elements in the image, and the generation characteristics of the music; and generate lyrics of the music based on the theme, emotion and keywords of the lyrics.
[0129] In some embodiments, the second generation module 540 is configured to understand the image in multiple attribute dimensions and determine the camera movement method; generate multiple frames of video images based on the image and the camera movement method; and generate multimedia content based on the multiple frames of video images and music.
[0130] In some embodiments, the second generation module 540 is configured to generate a description text of the image based on an understanding of multiple attribute dimensions of the image; expand the description text of the image to generate a description text of multiple frames of video images; generate multiple frames of video images based on the description text of the multiple frames of video images; and generate multimedia content based on the multiple frames of video images and music.
[0131] In some embodiments, the display module 550 is configured to display controls corresponding to one or more candidate features, wherein, in the case of displaying controls corresponding to multiple candidate features, the multiple candidate features correspond to one or more feature types; the first generation module 520 is configured to generate music description text in response to a user triggering operation on a control of one or more target candidate features among the one or more candidate features, based on the one or more target candidate features and an understanding of multiple attribute dimensions of the image.
[0132] In some embodiments, the first generation module 520 is configured to determine one or more generation features of the music based on an understanding of multiple attribute dimensions of the image; match one or more target candidate features with one or more generation features to determine whether there are generation features that conflict with the one or more target candidate features; in response to the existence of generation features that conflict with the one or more target candidate features, delete the generation features that conflict with the one or more target candidate features, and generate music description text based on the one or more target candidate features and the remaining generation features.
[0133] In some embodiments, the music description text also includes the lyrics of the music, and the display module 550 is configured to display the lyrics generation control; the first generation module 520 is configured to respond to the user triggering the lyrics generation control, understand the image in multiple attribute dimensions, and generate the lyrics of the music.
[0134] In some embodiments, the receiving module 510 is further configured to receive prompt information of multimedia content input by a user; the first generating module 520 is configured to generate music description text based on the understanding of multiple attribute dimensions of the image and the prompt information.
[0135] In some embodiments, the first generation module 520 is configured to understand the image in multiple attribute dimensions and generate a description text of the image, wherein the description text of the image includes information of multiple attribute dimensions; perform semantic understanding on the prompt information to determine whether the prompt information includes the theme of the multimedia content; in response to the prompt information including the theme of the multimedia content, generate a music description text based on the description text of the image and the theme of the multimedia content.
[0136] In some embodiments, the first generation module 520 is configured to perform semantic understanding on the prompt information to determine whether the prompt information includes the target features of the music; in response to the prompt information including the target features of the music, generate a music description text based on the descriptive text of the image, the theme of the multimedia content and the target features of the music.
[0137] In some embodiments, the display module 550 is configured to display an image selection control; in response to a user triggering the image selection control, multiple candidate images are displayed; the receiving module 510 is configured to receive and display an image and a multimedia content generation control in response to a user's selection operation of an image from multiple candidate images; and receive a user's trigger operation on the multimedia content generation control.
[0138] In some embodiments, the display module 550 is configured to display camera controls; in response to a user triggering the camera controls, a shooting preview interface is displayed, wherein the shooting preview interface includes a multimedia content generation option; the receiving module 510 is configured to receive an image in response to a user selecting a multimedia content generation option and triggering the capture of an image.
[0139] The present disclosure also provides an electronic device, including: a processor; and a memory coupled to the processor, for storing instructions, which, when executed by the processor, causes the processor to execute a method for generating multimedia content as in any embodiment of the present disclosure.
[0140] The electronic device disclosed in the present invention can understand multiple attribute dimensions based on the image input by the user, generate a music description text, and then determine the music based on the music description text, and then generate multimedia content based on the image and the music. Since the music description text is generated based on the understanding of multiple attribute dimensions of the image, the generation characteristics of the music described in the music description text are associated with the multiple attribute dimensions of the image, so that the generated music matches the image, and the multimedia content finally generated also matches the music, which improves the presentation effect and audio-visual effect of the multimedia content as a whole and better meets the needs of the user. The electronic device disclosed in the present invention can directly determine the music and generate multimedia content based on the image, without the user selecting and adding music, which improves the presentation effect and audio-visual effect of the multimedia content while improving the efficiency of generating multimedia content.
[0141] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for generating multimedia content as in any embodiment of the present disclosure is implemented.
[0142] The computer-readable storage medium disclosed in the present invention can understand multiple attribute dimensions based on the image input by the user, generate a music description text, and then determine the music based on the music description text, and then generate multimedia content based on the image and the music. Since the music description text is generated based on the understanding of multiple attribute dimensions of the image, the generation characteristics of the music described in the music description text are associated with the multiple attribute dimensions of the image, so that the generated music matches the image, and the multimedia content finally generated is also matched with the music, which improves the presentation effect and audio-visual effect of the multimedia content as a whole, and better meets the needs of users. The computer-readable storage medium disclosed in the present invention can directly determine the music and generate multimedia content based on the image, without the need for the user to select and add music, while improving the presentation effect and audio-visual effect of the multimedia content, the generation efficiency of the multimedia content is improved.
[0143] The present disclosure further provides a computer program product, including: instructions, which, when executed by a processor, enable the processor to execute a method for generating multimedia content as described in any embodiment of the present disclosure.
[0144] The computer program product disclosed in the present invention can understand multiple attribute dimensions based on the image input by the user, generate a music description text, and then determine the music based on the music description text, and then generate multimedia content based on the image and the music. Since the music description text is generated based on the understanding of multiple attribute dimensions of the image, the generation characteristics of the music described in the music description text are associated with the multiple attribute dimensions of the image, so that the generated music matches the image, and the multimedia content finally generated also matches the music, which improves the presentation effect and audio-visual effect of the multimedia content as a whole and better meets the needs of users. The computer program product disclosed in the present invention can directly determine the music and generate multimedia content based on the image, without the need for the user to select and add music, which improves the presentation effect and audio-visual effect of the multimedia content while improving the efficiency of generating multimedia content.
[0145] Combine the following Figure 6 and 7 An electronic device of the present disclosure is described. Figure 6 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0146] The memory 61 is used to store one or more computer-readable instructions. The memory 61 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. The memory 61 may store, for example, an operating system, an application, a boot loader (BootLoader), a database, and other programs, and may also store various applications and various data.
[0147] The processor 62 is used to run computer-readable instructions to implement the method described in any of the above embodiments. The specific implementation of each step of the method can be referred to the above embodiments, and the repeated parts will not be repeated here.
[0148] The processor 62 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) may be an X86 or ARM architecture, etc.
[0149] The processor 62 and the memory 61 may communicate with each other directly or indirectly. For example, the processor 62 and the memory 61 may communicate with each other through a network. The network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 62 and the memory 61 may also communicate with each other through a system bus, which is not limited in the present disclosure.
[0150] It should be noted that Figure 6 The components of the electronic device 6 shown are only exemplary and non-restrictive. The electronic device 6 may also have other components according to actual application requirements. The processor 62 may control other components in the electronic device 6 to perform desired functions.
[0151] The electronic device 6 may be implemented by software, firmware and / or hardware, and may be integrated into a device installed with relevant application programs.
[0152] Figure 7 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.
[0153] Figure 7 The electronic device 7 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.
[0154] Electronic devices include, but are not limited to, mobile terminals such as smart phones, laptops, personal digital assistants (PDA), tablet personal computers (Tablet PC), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital televisions, desktop computers, etc.
[0155] like Figure 7 As shown, a central processing unit (CPU) 71 performs various processes according to a program stored in a read-only memory (ROM) 72 or a program loaded from a storage section 78 to a random access memory (RAM) 73. In the RAM 73, data required when the CPU 71 performs various processes is stored as needed. The central processing unit is merely exemplary, and it may also be other types of processors, such as the various processors described above. The ROM 72, the RAM 73, and the storage section 78 may be various forms of computer-readable storage media. It should be noted that although Figure 7 ROM 72, RAM 73 and storage section 78 are shown separately in FIG. 1 , but one or more of them may be combined or located in the same or different memory or storage modules.
[0156] The CPU 71, the ROM 72, and the RAM 73 are connected to one another via a bus 74. To the bus 74, an input / output interface 75 is also connected.
[0157] The following components are connected to the input / output interface 75: an input portion 76, such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output portion 77, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage portion 78, including a hard disk, a magnetic tape, etc.; and a communication portion 79, including a network interface card such as a LAN card, a modem, etc. The communication portion 79 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 7 Some of the electronic devices 7 are shown to communicate via a bus 74, but they may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0158] A drive 710 is also connected to the input / output interface 75 as needed. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory or the like is mounted on the drive 710 as needed so that a computer program read therefrom is installed into the storage section 78 as needed.
[0159] When the above-described series of processing is realized by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 711 .
[0160] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which, when the computer program product is run on a computer, enables the computer to implement the method described in any of the aforementioned embodiments. The computer program product includes a computer instruction carried on a computer-readable medium, containing a program code for executing the method shown in the flowchart. In such an embodiment, the computer instruction can be downloaded and installed from the network through the communication part 79, or installed from the storage part 78, or installed from the ROM 72. When the computer program is executed by the CPU 71, the method of the embodiment of the present disclosure is executed.
[0161] It should be noted that, in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, apparatus, or device or for use in conjunction with an instruction execution system, apparatus, or device.
[0162] The computer readable medium may be a computer readable storage medium, or a computer readable signal medium, or any combination of the two.
[0163] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections with one or more wires, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device, or device. Computer instructions are stored on a computer-readable storage medium, and when the instructions are executed by a processor, the method described in any of the foregoing embodiments is implemented.
[0164] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer readable program codes. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than a computer readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0165] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0166] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to execute the method described in any of the above embodiments. For example, the instructions may be embodied as computer program codes.
[0167] In embodiments of the present disclosure, computer program codes for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In situations involving a remote computer, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).
[0168] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0169] The functions described above may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0170] Although some specific embodiments of the present disclosure have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. It should be understood by those skilled in the art that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for generating multimedia content, comprising: Receive an image input by a user; Understanding the image from multiple attribute dimensions to generate a music description text, wherein the music description text includes description information of generated features of the music; Determine the music according to the music description text; The multimedia content is generated according to the image and the music, and the multimedia content is displayed.
2. The generation method according to claim 1, wherein: The step of understanding the image in multiple attribute dimensions and generating a music description text comprises: Understanding the image from multiple attribute dimensions and generating a description text of the image, wherein the description text of the image includes information of the multiple attribute dimensions; The description text of the image is semantically understood to generate the music description text, wherein the generated features of the music are obtained based on the information of the multiple attribute dimensions.
3. The generation method according to claim 2, wherein: The understanding of the image in multiple attribute dimensions and generating a description text of the image includes: Understanding the image according to multiple attribute dimensions of entity, scene, emotion, environment, theme, style, and effect elements, and determining information of multiple attribute dimensions of entity, scene, emotion, environment, theme, style, and effect elements in the image; Generate a description text of the image according to the information of the multiple attribute dimensions.
4. The generation method according to claim 3, wherein: The performing semantic understanding on the description text of the image to generate the music description text comprises: Performing semantic understanding on the description text of the image to determine information on multiple attribute dimensions of entities, scenes, emotions, environments, themes, styles, and effect elements in the image; Determine the generated features of the music that match the information of multiple attribute dimensions in the entity, scene, emotion, environment, theme, style, and effect elements, wherein the generated features of the music include features corresponding to at least one feature type of the atmosphere, style, rhythm, melody, harmony, timbre, and vocals of the music; Generate a description text of the music according to the generation features of the music.
5. The generation method according to claim 4, wherein: The music description text also includes lyrics of the music, and the step of generating the music description text according to the generated features of the music includes: Generate lyrics of the music according to information of multiple attribute dimensions in entities, scenes, emotions, environments, themes, styles, and effect elements in the image and generation features of the music; The music description text is generated according to the generation characteristics of the music and the lyrics of the music.
6. The generation method according to claim 5, wherein: Generating lyrics of the music according to information of multiple attribute dimensions in entities, scenes, emotions, environments, themes, styles, and effect elements in the image and the generation features of the music includes: Determine the theme, emotion and keywords of the lyrics according to information of multiple attribute dimensions in the entities, scenes, emotions, environments, themes, styles and effect elements in the image and the generation characteristics of the music; The lyrics of the music are generated according to the theme, emotion and keywords of the lyrics.
7. The generation method according to claim 1, wherein: Generating the multimedia content according to the image and the music includes: Understanding the image from multiple attribute dimensions and determining the camera movement method; Generate multiple frames of video images according to the image and the camera movement mode; The multimedia content is generated according to the multiple frames of video images and the music.
8. The generation method according to claim 1, wherein: Generating the multimedia content according to the image and the music includes: Generate a description text of the image based on the understanding of multiple attribute dimensions of the image; Expanding the description text of the image to generate description text of multiple frames of video images; Generate the multiple frames of video images according to the description texts of the multiple frames of video images; The multimedia content is generated according to the multiple frames of video images and the music.
9. The generation method according to claim 1, wherein: The step of understanding the image in multiple attribute dimensions and generating a music description text comprises: Displaying controls corresponding to one or more candidate features, wherein, in the case of displaying controls corresponding to multiple candidate features, the multiple candidate features correspond to one or more feature types; In response to the user triggering an operation on a control of one or more target candidate features among the one or more candidate features, the music description text is generated based on the one or more target candidate features and an understanding of multiple attribute dimensions of the image.
10. The generation method according to claim 9, wherein: The step of generating the music description text according to the one or more target candidate features and understanding the image in multiple attribute dimensions includes: Determining one or more generative features of the music based on understanding the image in multiple attribute dimensions; Matching the one or more target candidate features with the one or more generated features to determine whether there are generated features that conflict with the one or more target candidate features; In response to the existence of generated features that contradict the one or more target candidate features, the generated features that contradict the one or more target candidate features are deleted, and the music description text is generated based on the one or more target candidate features and the remaining generated features.
11. The generation method according to any one of claims 1 to 10, wherein: The music description text also includes lyrics of the music, and the understanding of the image in multiple attribute dimensions to generate the music description text includes: Display lyrics generation controls; In response to the user triggering the lyrics generation control, the image is understood from multiple attribute dimensions to generate lyrics for the music.
12. The generation method according to any one of claims 1 to 10, wherein: The understanding of the image in multiple attribute dimensions to generate a music description text includes: receiving prompt information of the multimedia content input by the user; The music description text is generated based on the understanding of multiple attribute dimensions of the image and the prompt information.
13. The generation method according to claim 12, wherein: The understanding of the image in multiple attribute dimensions and the prompt information to generate the music description text includes: Understanding the image from multiple attribute dimensions and generating a description text of the image, wherein the description text of the image includes information of the multiple attribute dimensions; Performing semantic understanding on the prompt information to determine whether the prompt information includes the theme of the multimedia content; In response to the prompt information including the theme of the multimedia content, the music description text is generated according to the description text of the image and the theme of the multimedia content.
14. The generation method according to claim 13, wherein: The step of understanding the image in multiple attribute dimensions and the prompt information to generate the music description text further includes: Performing semantic understanding on the prompt information to determine whether the prompt information includes the target feature of the music; In response to the prompt information including the target feature of the music, the music description text is generated according to the description text of the image, the theme of the multimedia content and the target feature of the music.
15. The generation method according to any one of claims 1 to 10, wherein: The receiving of an image input by a user comprises: Display image selection controls; In response to the user triggering the image selection control, displaying a plurality of candidate images; In response to the user's selection operation of the image from the plurality of candidate images, receiving and displaying the image and a multimedia content generation control; A triggering operation of the user on the multimedia content generation control is received.
16. The generation method according to any one of claims 1 to 10, wherein: The receiving of an image input by a user comprises: Display camera controls; In response to the user triggering the camera control, displaying a shooting preview interface, wherein the shooting preview interface includes a multimedia content generation option; In response to the user selecting the multimedia content generation option and triggering the capturing of the image, the image is received.
17. A device for generating multimedia content, comprising: A receiving module, configured to receive an image input by a user; A first generating module is configured to understand the image from multiple attribute dimensions and generate a music description text, wherein the music description text includes description information of generated features of the music; A determination module, configured to determine the music according to the music description text; A second generating module is configured to generate the multimedia content according to the image and the music; The display module is configured to display the multimedia content.
18. An electronic device comprising: processor; as well as A memory coupled to the processor, for storing instructions, wherein when the instructions are executed by the processor, the processor executes the method for generating multimedia content according to any one of claims 1 to 16.
19. A computer-readable storage medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method for generating multimedia contents according to any one of claims 1 to 16 is implemented.
20. A computer program product comprising: The instructions, when executed by a processor, cause the processor to execute the method for generating multimedia content according to any one of claims 1 to 16.