Method and device for generating media object album, storage medium and equipment
By using generative artificial intelligence to generate media object collections, the problems of insufficient or excessive personalization caused by manual participation in existing technologies are solved, and the generation of high-quality and diverse media object collections is achieved.
Patent Information
- Application Number
- CN202410303240.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies for generating media object collections require manual participation, resulting in insufficient or excessive personalization, failure to meet user needs, and lack of diversity.
Generative artificial intelligence is used to generate a media object collection based on the title, introduction, cover and media object keywords, including selecting a collection title prompt template and keywords, generating a collection title and introduction, determining the collection cover, and recalling target media objects from the media object library to assemble them into a collection.
It reduces labor costs, meets users' personalized needs while maintaining diversity, and the generated collection is of higher quality.
Smart Images

Figure CN120653790A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of Internet technology. More specifically, the embodiments of the present disclosure relate to a method for generating a media object collection, an apparatus for generating a media object collection, a computer-readable storage medium, and an electronic device. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims, and no statement herein is admitted to be prior art by inclusion in this section.
[0003] With the development of the digital media object (such as music, audiobooks, podcasts, etc.) industry, media object collections (such as playlists, podcast collections, audiobook collections, etc.) play an increasingly important role in third-party platform consumption. Summary of the Invention
[0004] However, related technologies require manual participation to generate media object collections, and the generated media object collections are either insufficiently personalized and unable to meet user needs, or overly personalized and lack diversity.
[0005] Therefore, a method for generating a collection of media objects is highly needed to reduce labor costs while meeting the personalized needs of users and maintaining diversity.
[0006] In this context, embodiments of the present disclosure are intended to provide a method for generating a media object collection, an apparatus for generating a media object collection, a computer-readable storage medium, and an electronic device.
[0007] According to a first aspect of the present disclosure, a method for generating a media object collection is provided, the method comprising: generating a target collection title through generative artificial intelligence based on title keywords; generating a target collection introduction through generative artificial intelligence based on the introduction keywords; determining a target collection cover based on the cover keywords; recalling media objects from a media object library based on the media object keywords to obtain a target media object; and assembling the target collection title, the target collection introduction, the target collection cover, and the target media object to obtain a target media object collection.
[0008] In one embodiment, generating a target collection title through generative artificial intelligence based on title keywords includes: determining a target collection title prompt template selected by the user from a collection title prompt template library; obtaining a title prompt and title keywords input by the user based on the target collection title prompt template; and generating the target collection title based on the title prompt and the title keywords through a text-based large model.
[0009] In one embodiment, generating a target collection introduction based on the introduction keywords through generative artificial intelligence includes: determining the target collection introduction prompt template selected by the user from a collection introduction prompt template library; obtaining the introduction prompt and introduction keywords input by the user based on the target collection introduction prompt template; and generating the target collection introduction based on the introduction prompt and the introduction keywords through a text-based large model.
[0010] In one embodiment, determining the target collection cover based on the cover keywords includes: determining the target collection cover prompt template selected by the user from a collection cover prompt template library; obtaining the cover prompt and cover keywords input by the user based on the target collection cover prompt template; and generating the target collection cover based on the target collection title, the target collection introduction, the cover prompt and the cover keywords through a large picture model.
[0011] In one embodiment, determining the target collection cover based on the cover keyword includes: obtaining the cover keyword; matching the cover keyword with words corresponding to the labels of the collection cover in the collection cover database; and determining the collection cover to which the label corresponding to the word with the highest similarity to the cover keyword belongs as the target collection cover.
[0012] In one embodiment, the recalling of media objects from a media object library based on the media object keywords to obtain a target media object includes: obtaining the media object keywords; and performing intersection and difference selection on the media objects in the media object library based on the media object keywords to obtain the target media object.
[0013] In one embodiment, performing intersection-and-subtraction selection on media objects in a media object library based on the media object keywords to obtain the target media object includes: determining a selection condition based on the media object keywords; and performing intersection-and-subtraction selection on media objects in the media object library based on the selection condition to obtain the target media object.
[0014] In one embodiment, the recalling of media objects from a media object library based on the media object keywords to obtain the target media object includes: obtaining the media object keywords; and recalling media objects from a media object library based on the media object keywords to obtain the target media object.
[0015] In one embodiment, the method further includes: determining whether the media object keywords meet preset requirements; if the media object keywords do not meet the preset requirements, performing label mapping on the media object keywords in the media object keywords to obtain media object keywords that meet the preset requirements.
[0016] In one embodiment, the label mapping of the media object keywords in the media object keywords to obtain the media object keywords that meet the preset requirements includes: determining corresponding labels from label databases of different dimensions based on the media object keywords; and determining the words corresponding to the determined labels as media object keywords that meet the preset requirements.
[0017] According to a second aspect of the present disclosure, a device for generating a media object collection is provided, the device comprising: a title generation module, configured to generate a target collection title through generative artificial intelligence based on title keywords; a description generation module, configured to generate a target collection description through generative artificial intelligence based on the description keywords; a cover determination module, configured to determine a target collection cover based on the cover keywords; a media object determination module, configured to recall media objects from a media object library based on the media object keywords to obtain a target media object; and an assembly module, configured to assemble the target collection title, the target collection description, the target collection cover and the target media object to obtain a target media object collection.
[0018] In one embodiment, the title generation module is configured to: determine the target collection title prompt template selected by the user from the collection title prompt template library; obtain the title prompt and title keywords input by the user based on the target collection title prompt template; and generate the target collection title based on the title prompt and the title keywords through a text-based large model.
[0019] In one embodiment, the introduction generation module is configured to: determine the target collection introduction prompt template selected by the user from the collection introduction prompt template library; obtain the introduction prompt and introduction keywords input by the user based on the target collection introduction prompt template; and generate the target collection introduction based on the introduction prompt and the introduction keywords through a text-based large model.
[0020] In one embodiment, the cover determination module is configured to: determine the target collection cover prompt template selected by the user from the collection cover prompt template library; obtain the cover prompt and cover keywords input by the user based on the target collection cover prompt template; and generate the target collection cover through a picture-type large model based on the target collection title, the target collection introduction, the cover prompt and the cover keywords.
[0021] In one embodiment, the cover determination module is configured to: obtain the cover keyword; match the cover keyword with the words corresponding to the labels of the collection cover in the collection cover database; and determine the collection cover to which the label corresponding to the word with the highest similarity to the cover keyword belongs as the target collection cover.
[0022] In one embodiment, the media object determination module is configured to: obtain the media object keyword; and perform intersection-and-subtraction selection on media objects in a media object library based on the media object keyword to obtain the target media object.
[0023] In one embodiment, the media object determination module is configured to: determine a selection condition according to the media object keyword; and perform intersection-and-subtraction selection on the media objects in the media object library according to the selection condition to obtain the target media object.
[0024] In one embodiment, the media object determination module is configured to: obtain the media object keyword; and retrieve media objects from a media object library based on the media object keyword to obtain the target media object.
[0025] In one embodiment, the device also includes a label mapping module, which is configured to: determine whether the media object keywords meet the preset requirements; if the media object keywords do not meet the preset requirements, perform label mapping on the media object keywords in the media object keywords to obtain media object keywords that meet the preset requirements.
[0026] In one embodiment, the tag mapping module is configured to: determine corresponding tags from tag databases of different dimensions based on the media object keywords; and determine the words corresponding to the determined tags as media object keywords that meet the preset requirements.
[0027] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any one of the above methods is implemented.
[0028] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform any one of the above methods by executing the executable instructions.
[0029] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements any one of the above methods when executed by a processor.
[0030] According to the method for generating a media object collection, the device for generating a media object collection, the computer-readable storage medium, and the electronic device of the embodiment of the present disclosure, a target collection title is generated by generative artificial intelligence based on the title keywords; a target collection introduction is generated by generative artificial intelligence based on the introduction keywords; a target collection cover is determined based on the cover keywords; media objects are recalled from a media object library based on the media object keywords to obtain a target media object; and the target collection title, the target collection introduction, the target collection cover, and the target media object are assembled to obtain a target media object collection. The ability to generate a media object collection based on keywords through generative artificial intelligence reduces labor costs while meeting the personalized needs of users without sacrificing diversity. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0032] Figure 1 A schematic diagram of a process architecture for generating a media object collection in an embodiment of the present disclosure is shown;
[0033] Figure 2 A flow chart showing a method for generating a media object collection according to an embodiment of the present disclosure is shown;
[0034] Figure 3 A flowchart showing a method for generating a media object collection in an embodiment of the present disclosure for generating a target collection title is shown;
[0035] Figure 4 A schematic diagram illustrating a collection title prompt template in a method for generating a media object collection in an embodiment of the present disclosure;
[0036] Figure 5 A flowchart showing a method for generating a media object collection in an embodiment of the present disclosure for generating a target collection introduction is shown;
[0037] Figure 6A schematic diagram illustrating a collection introduction prompt template in a method for generating a media object collection in an embodiment of the present disclosure;
[0038] Figure 7 A flowchart of generating a target album cover in a method for generating a media object album in an embodiment of the present disclosure is shown;
[0039] Figure 8 A schematic diagram illustrating a collection cover prompt template in a method for generating a media object collection in an embodiment of the present disclosure;
[0040] Figure 9 A schematic diagram showing an album cover generated in a method for generating a media object album in an embodiment of the present disclosure;
[0041] Figure 10 A flowchart illustrating another method for generating a target collection cover in a method for generating a media collection in an embodiment of the present disclosure is shown;
[0042] Figure 11 A flowchart of determining a target media object in a method for generating a media object collection in an embodiment of the present disclosure is shown;
[0043] Figure 12 A flowchart illustrating the method of determining a target media object by intersection and difference in a method for generating a media object collection according to an embodiment of the present disclosure is shown;
[0044] Figure 13 A schematic diagram illustrating a selection condition in a method for generating a media object collection in an embodiment of the present disclosure;
[0045] Figure 14 A flowchart illustrating another method for determining a target media object in a method for generating a media object collection according to an embodiment of the present disclosure is shown;
[0046] Figure 15 A flowchart illustrating whether a keyword is qualified in a method for generating a media object collection in an embodiment of the present disclosure is shown;
[0047] Figure 16 A flowchart showing label mapping in a method for generating a media object collection in an embodiment of the present disclosure is shown;
[0048] Figure 17 A schematic structural diagram of a device for generating a media object collection according to an embodiment of the present disclosure is shown;
[0049] Figure 18 A schematic structural diagram of an electronic device in an embodiment of the present disclosure is shown.
[0050] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION
[0051] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. Rather, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0052] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, apparatus, device, method, or computer program product. Therefore, the present disclosure may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software.
[0053] According to an embodiment of the present disclosure, a method for generating a media object collection, an apparatus for generating a media object collection, a computer-readable storage medium, and an electronic device are provided.
[0054] In this document, any number of elements in the drawings is for illustration and not for limitation, and any naming is for distinction only and does not have any limiting meaning.
[0055] The principles and spirit of the present disclosure are described in detail below with reference to several representative embodiments of the present disclosure. SUMMARY OF THE INVENTION
[0057] With the development of the digital media object (such as music, audiobooks, podcasts, etc.) industry, media object collections (such as playlists, podcast collections, audiobook collections, etc.) play an increasingly important role in third-party platform consumption.
[0058] In the first related technology, information about media objects, including their names, creators, and playback duration, is collected through tools such as databases or APIs. Media objects are then screened based on factors such as style, genre, and sentiment. These media objects are then sorted based on specific metrics such as playback volume, likes, and comments. Finally, the top-ranked media objects are added to a media collection, and information such as covers, tags, and descriptions are set. This method has the following drawbacks:
[0059] (1) Human participation is required in the song screening and sorting process. The subjectivity and limitations of human intervention may lead to insufficient quality and diversity of the playlist.
[0060] (2) It does not take into account the user's personalized needs and interests and cannot make personalized recommendations based on the user's historical media listening records, preferences, comments, and other information. Therefore, it may not meet the user's needs;
[0061] (3) Reliance on tools such as databases or APIs. If these tools lack relevant data or information, the quality and diversity of the media collection will be affected. At the same time, the use of these tools may require payment or be subject to restrictions on use.
[0062] In the second related technique, a target user's preferred media object set is obtained, from which a seed media object is determined. A set of pictures associated with the seed media object is then obtained, and the target user's predicted preference scores for these pictures are calculated. A seed media object picture is then selected from the set. Next, similar media objects are retrieved based on each preferred media object in the preferred media object set, and the other preferred media objects and similar media objects are identified as candidate media objects. The target user's predicted preference scores for the candidate media objects are then calculated, and a media object list is generated based on the predicted scores. Finally, a media object collection is generated based on the seed media object pictures and the media object list. This method has the following disadvantages:
[0063] (1) It is necessary to generate a media object collection based on the user's historical preference set of media objects and the predicted preference score for pictures. For new users or users with less data, it may not be possible to generate an accurate media object collection;
[0064] (2) Generating a media object collection based on a user's historical preferences may result in the generated media object collection being overly personalized and lacking in diversity.
[0065] In view of the above, the present disclosure provides a method for generating a media object collection, an apparatus for generating a media object collection, a computer-readable storage medium, and an electronic device. Generative artificial intelligence is used to generate a target collection title based on a title keyword set; a target collection introduction is generated based on a description keyword set; a target collection cover is determined based on a cover keyword set; media objects are recalled from a media object library based on a media object keyword set to obtain a target media object; and the target collection title, target collection introduction, target collection cover, and target media object are assembled to obtain a target media object collection. Generative artificial intelligence can generate a media object collection based on keywords, reducing labor costs while meeting the personalized needs of users without sacrificing diversity.
[0066] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.
[0067] Application Scenario Overview
[0068] It should be noted that the following application scenarios are only provided to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0069] The present disclosure can be applied to any scenario of generating a media object collection. The server generates a target collection title through generative artificial intelligence based on title keywords; generates a target collection introduction through generative artificial intelligence based on introduction keywords; determines the target collection cover based on cover keywords; recalls media objects from a media object library based on media object keywords to obtain target media objects; and assembles the target collection title, target collection introduction, target collection cover, and target media objects to obtain a target media object collection.
[0070] Exemplary Methods
[0071] The following combination Figure 1 The system architecture and application scenarios of the operating environment of this exemplary embodiment are exemplarily described.
[0072] Figure 1 A schematic diagram of a system architecture is shown, and the system architecture 100 may include a terminal 110 and a server 120. The terminal 110 may be a smart phone, a tablet computer, a personal computer, etc. The server 120 may generally refer to a background system (such as a media object collection generation system) that provides related services for generating a media object collection, and may be a server or a cluster formed by multiple servers. The terminal 110 may receive title keywords, introduction keywords, cover keywords, and media object keywords, and send them to the server. The server 120 generates a target collection title through generative artificial intelligence based on the title keywords; generates a target collection introduction through generative artificial intelligence based on the introduction keywords; determines the target collection cover based on the cover keywords; recalls media objects from the media object library based on the media object keywords to obtain the target media object; and assembles the target collection title, target collection introduction, target collection cover, and target media object to obtain the target media object collection. The terminal 110 and the server 120 may be connected via a wired or wireless communication link for data exchange.
[0073] An exemplary embodiment of the present disclosure first provides a method for generating a media object collection, which may include:
[0074] Generate target collection titles using generative AI based on title keywords;
[0075] Generate a target collection profile using generative AI based on the profile keywords;
[0076] Determine the target album cover based on cover keywords;
[0077] According to the media object keywords, media objects are recalled from the media object library to obtain the target media object;
[0078] The target collection title, the target collection introduction, the target collection cover and the target media object are assembled to obtain a target media object collection.
[0079] Figure 2 The exemplary process of the method for generating a media object collection is shown below. Figure 2 Each step is described in detail.
[0080] refer to Figure 2 ,In step S210, the target collection title is generated by generative artificial intelligence based on the title keywords.
[0081] Generative AI (Artificial Intelligence Generated Content, AIGC) refers to technologies based on AI techniques such as generative adversarial networks and large pre-trained models that generate relevant content with appropriate generalization capabilities by learning from and identifying existing data. The core concept of AIGC technology is to use AI algorithms to generate content with a certain degree of creativity and quality. By training models and learning from large amounts of data, AIGC can generate relevant content based on input conditions or guidance. For example, by inputting keywords, descriptions, or samples, AIGC can generate matching articles, images, audio, and more.
[0082] In one embodiment, specifically, reference Figure 3 , the above step S210 may further include the following steps S310 to S330:
[0083] Step S310: Determine the target album title prompt template selected by the user from the album title prompt template library.
[0084] Among them, the prompt (prompt) template is a concept in natural language processing, which aims to guide the model to better understand and respond to input by designing specific templates and filling methods. Here, template: is a fixed text structure used to generate prompts. Common template forms include "{Task}:{Input}" and "How to{Task}:{Input}", which provide consistent input methods between different tasks. Filling method: In the text generation process, the filling method is used to combine the prompt keywords with the input text. Common methods include direct insertion, splicing encoding vectors and embedding, etc., to effectively integrate the prompt keywords with the input text.
[0085] The collection title hint template is at least used to determine the title hint and title keywords related to the title. Then, the model (such as GPT model, Midjourney model, etc.) will generate text, pictures or videos based on the title hint and title keywords; for example: Figure 4 As shown, this is a collection title prompt template.
[0086] Step S320: Obtain the title prompt and title keywords input by the user according to the target collection title prompt template.
[0087] Among them, the title prompt is used to describe the conditions that the generated title must meet. It can be input by the user or it can be included in the collection title prompt template. There is no limitation here. Title keywords are words that describe the characteristics of the title. Title keywords can be words, or phrases and sentences with similar semantics to words, etc., which are generally input by the user. For example: Figure 4 As shown, in the collection title prompt template, the user selected the ChatGpt model in the "Select Model List". The ChatGpt model is used to generate text. The title prompt is the gray font in the box below "Enter Prompt Words". "Suppose you (that is, the ChatGpt model) are a playlist expert on a music streaming platform. You are editing your own playlist title. Please write a playlist title for the collection of ${keyword} in 20 Chinese characters. It should be clear in key points, literary and appealing. It is required not to include specific singer names, double quotes and other special symbols." The title keyword is the "keyword" below the "variable". The title prompt can be provided by the template or entered by the user. The title keyword is generally entered by the user; specifically, when the title prompt is blank, it is entered by the user. When the title prompt is the gray font provided by the template (which can be regarded as a reference example), the user can refer to the gray font to enter the title prompt, or directly use the gray font as the title prompt.
[0088] Step S330: Generate the target collection title based on the title prompt and title keywords through the text-based large model.
[0089] Large text models are mainly used to implement dialogue, text generation, code generation, etc.; for example: GPT, PaLM, Llama and other models.
[0090] For example, if Figure 4 As shown, the ChatGpt model generates a collection title (Music Painting: Play the heartstrings, stir the soul, stir the emotions, bloom the love, and show the unique charm of ${keyword}) based on the title prompt and title keywords (such as: drawing, picture, imagination).
[0091] Continue to refer Figure 2,In step S220, a target collection introduction is generated through generative artificial intelligence based on the introduction keywords.
[0092] Introduction keywords are words that describe the characteristics of the introduction. For details, refer to Figure 5 , the above step S220 may further include the following steps S510 to S530:
[0093] Step S510: Determine the target album introduction prompt template selected by the user from the album introduction prompt template library.
[0094] The collection introduction prompt template is at least used to determine the introduction prompts and introduction keywords related to the introduction, and then the model (such as: GPT model, Midjourney model, etc.) will generate text, pictures or videos based on the introduction prompts and introduction keywords; for example: Figure 6 The following is a template for a collection introduction prompt.
[0095] Step S520: Obtain the introduction prompt and introduction keywords input by the user according to the target collection introduction prompt template.
[0096] Among them, the introduction keywords are words used to describe the characteristics of the introduction. Introduction keywords can be words, or phrases or sentences with similar semantics to words, and are generally entered by the user; for example: Figure 6 As shown, in the collection introduction prompt template, the user selected the ChatGpt model in the "Select Model List". The ChatGpt model is used to generate text. The introduction prompt is the gray font in the box below "Enter prompt words". "Suppose you (that is, the ChatGpt model) are a playlist expert on a music streaming platform, and you are editing your own playlist introduction. Please follow the template of [History of the genre] + [Features of the performance of the genre] + [Features of the melody of the genre] + [2 representative artists and their representative works] + [A sentence describing the scene corresponding to the genre] to generate a complete, fluent, non-repetitive and literary playlist introduction describing the ${keyword} genre". The introduction keyword is the "keyword" below the "variable". The introduction prompt can be provided by the template or entered by the user. The introduction keyword is generally entered by the user; specifically, when the introduction prompt is blank, it is entered by the user. When the introduction prompt is in the gray font provided by the template (which can be regarded as a reference example), the user can refer to the gray font to enter the introduction prompt, or directly use the gray font as the introduction prompt.
[0097] Step S530: Generate a target collection introduction based on the introduction prompts and introduction keywords through the text-based large model.
[0098] For example, if Figure 6As shown, the ChatGpt model generates an album introduction based on the introduction prompts and introduction keywords (Welcome to my happy music playlist! Let us immerse ourselves in this vibrant and joyful music world together!...).
[0099] Continue to refer Figure 2 In step S230, the target album cover is determined based on the cover keywords.
[0100] The album cover can be generated by generative method or determined by matching method. Figure 7 The process of generating the album cover is as follows, that is, the above step S230 can further include the following steps S710 to S730:
[0101] Step S710: Determine the target album cover prompt template selected by the user from the album cover prompt template library.
[0102] The album cover prompt template is at least used to determine the cover prompts and cover keywords related to the cover, and then the model (such as: GPT model, Midjourney model, etc.) will generate text, pictures or videos based on the cover prompts and cover keywords; for example: Figure 8 The following is a collection cover prompt template.
[0103] Step S720: Obtain the cover hint and cover keywords input by the user according to the target album cover hint template.
[0104] Cover keywords are words that describe the characteristics of the cover. They can be words, phrases, or sentences with similar semantics, and are generally entered by the user. Examples include minimalist, concise and elegant, and simple and elegant design styles.
[0105] For example, if Figure 8As shown, in the album cover prompt template, the cover prompt is the gray font "Now you are an image prompt generator, you can generate prompts to describe the image. The prompt framework is: subject (Subject) + theme (Theme) + style (Medium) + scene (Environment) + composition (Composition) + lighting effect (Lighting) + hue (Color) + mood (Mood), etc. The subject can be a person, an object, an animal, etc.; the theme can be music, a festival, a hot event, etc.; style refers to the style of the entire picture, such as graffiti, illustration, 3D, minimalism, etc.; scene refers to the environment where the subject is located, which can be indoors or outdoors, etc.; composition refers to where the focus of the lens is and where the subject is facing; hue refers to the color of the entire picture, such as vivid, bright, etc.; mood refers to the image through the picture. The characteristics of the entire image, such as calm or noisy, can be seen. Generate prompts according to this framework. The prompts should be as rich as possible, but the number of words should not exceed 60, and they should be generated in the order of the framework. Do not add explanatory words before the parameters. ...", the cover keywords are the subject (Subject), theme (Theme), style (Medium), scene (Environment), composition (Composition), lighting effect (Lighting), color (Color), and style (Mood) below the "variable". The cover prompt can be provided by the template or entered by the user. The cover keyword is generally entered by the user; specifically, when the cover prompt is blank, it is entered by the user. When the cover prompt is the gray font provided by the template (which can be regarded as a reference example), the user can refer to the gray font to enter the cover prompt, or directly use the gray font as the cover prompt.
[0106] Step S730: Generate a target album cover based on the target album title, target album introduction, cover prompts and cover keywords through the image-based large model.
[0107] Image-based large models are large models used to generate images. They can realize text-to-image generation and image-to-image generation; for example, the Midjourney, Stable Diffusion, and DALL.3 models.
[0108] For example, if Figure 9 As shown in the figure, the Midjourney model generates album covers based on cover prompts and cover keywords.
[0109] refer to Figure 10The process of matching and determining the album cover is as follows, that is, the above step S230 may further include the following steps S1010 to S1030:
[0110] Step S1010: Obtain cover keywords.
[0111] Among them, the cover keywords are the same as the cover keywords in the above step S720, and will not be repeated here.
[0112] Step S1020: Match the cover keywords with words corresponding to the labels of the album covers in the album cover database.
[0113] Among them, the collection cover database includes a large number of collection covers. Each collection cover is marked with a label to describe the characteristics of the collection cover. The label can be obtained by industry professionals by marking the collection cover. A collection cover can have one label or multiple labels. The label can be a word, or a phrase or sentence with similar semantics to the word, etc., which is not limited here.
[0114] In actual operation, a threshold can be set based on experience (for example, a value within 85%-100%). When the semantic similarity between the cover keyword and the word corresponding to the label of the album cover exceeds the threshold, the two are considered to match; otherwise, the two are considered to not match.
[0115] Step S1030: Determine the album cover to which the tag corresponding to the word with the highest similarity to the cover keyword belongs as the target album cover.
[0116] Similarity refers to the similarity between two objects. It is generally determined by calculating the distance between the features of the objects. If the distance is small, the similarity is high; conversely, the similarity is low. Similarity can be calculated using methods such as Manhattan distance, Euclidean distance, cosine similarity, and Pearson correlation coefficient.
[0117] Continue to refer Figure 2 In step S240, media objects are retrieved from the media object library according to the media object keywords to obtain the target media object.
[0118] The target media object can be determined by offline recall or online recall; for details, refer to Figure 11 The offline recall process is as follows, that is, the above step S240 may further include the following steps S1110 and S1120:
[0119] Step S1110: Acquire media object keywords.
[0120] Media object keywords are words that describe the characteristics of a media object. These can be words, phrases, or sentences with semantic similarities to words, and are typically user-entered. Examples include words that describe the genre, the setting, the theme, and the emotion of the media object.
[0121] Step S1120: Based on the media object keywords, perform intersection and difference selection on the media objects in the media object library to obtain the target media object.
[0122] Among them, the media object library includes a large number of media objects, each of which is marked with a label to describe the characteristics of the media object. The label can be obtained by industry professionals by labeling the media object. A media object can have one label or multiple labels. The label can be a word, or a phrase or sentence with similar semantics to the word, etc., which is not limited here.
[0123] In actual operation, a threshold value can be set based on experience (for example, a value within 85%-100%). When the semantic similarity between the media object keyword and the word corresponding to the media object tag exceeds the threshold, the two are considered to match; otherwise, they are considered to be mismatched. Figure 12 , the above step S1120 may further include the following steps S1210 and S1220:
[0124] Step S1210: Determine the selection criteria based on the media object keywords.
[0125] Among them, the keywords in the media object keywords can be determined as the selection conditions, or semantic expansion can be performed based on the keywords in the media object keywords, and the expanded words, phrases or sentences can be determined as the selection conditions, which is not limited here. Figure 13 As shown, the music style is rock and the emotion is happy, which are the selection criteria.
[0126] Step S1220: Perform intersection and difference selection on the media objects in the media object library according to the selection criteria to obtain the target media object.
[0127] For example, if Figure 13 As shown, according to the selection criteria, more than 7.8 million songs with rock genre and happy emotion can be selected.
[0128] refer to Figure 14 The online recall process is as follows, that is, the above step S240 may further include the following steps S1410 and S1420:
[0129] Step S1410: Acquire media object keywords.
[0130] The media object keywords are the same as those in the above step S1110 and are not described again here.
[0131] Step S1420: Recall media objects from the media object library according to the media object keywords to obtain target media objects.
[0132] Among them, the media object library includes a large number of media objects, each of which is marked with a label to describe the characteristics of the media object. The label can be obtained by industry professionals by labeling the media object. A media object can have one label or multiple labels. The label can be a word, or a phrase or sentence with similar semantics to the word, etc., which is not limited here.
[0133] In actual operation, a threshold value can be set based on experience (for example, a value within 85%-100%). When the semantic similarity between the media object keyword and the word corresponding to the media object tag exceeds the threshold, the media object is recalled; otherwise, it is not recalled.
[0134] Continue to refer Figure 2 In step S250, the target collection title, target collection introduction, target collection cover and target media object are assembled to obtain the target media object collection.
[0135] Among them, this step is to combine the target collection title, target collection introduction, target collection cover and target media object together, for example: combine them into a playlist, combine them into an album, etc.
[0136] In actual operation, the keywords input by the user may not be qualified. In this case, it is necessary to determine qualified keywords based on the keywords input by the user to make the recalled media objects more accurate. For details, refer to Figure 15 The method for generating a media object collection may further include the following steps S1510 and S1520:
[0137] Step S1510: Determine whether the media object keyword meets the preset requirements.
[0138] The preset requirements are used to measure whether the media object keywords input by the user are qualified. The preset requirements can be set based on experience or relevant literature, and are not limited here.
[0139] Generally, words that cannot directly describe the characteristics of a media object are considered keywords that do not meet the preset requirements, while words that can directly describe the characteristics of a media object are considered keywords that meet the preset requirements. For example, "rainy day" is a word that does not meet the preset requirements for describing a media object, while "popular," "light music," "sad," and "lyrical" are words that meet the preset requirements for describing a media object.
[0140] Step S1520: When the media object keywords do not meet the preset requirements, label mapping is performed on the media object keywords in the media object keywords to obtain media object keywords that meet the preset requirements.
[0141] Among them, label mapping can be understood as mapping the unqualified media object keywords input by the user to qualified media object keywords (words corresponding to the labels). Figure 16 The above step S1520 of "mapping the media object keywords in the media object keywords to obtain media object keywords that meet the preset requirements" may further include the following steps S1610 and S1620:
[0142] Step S1610: Determine corresponding tags from tag databases of different dimensions based on the media object keywords.
[0143] In practice, we create tag databases of different dimensions to map media object keywords that don't meet preset requirements. For example, the unqualified media object keyword "rainy day" entered by a user is mapped to "popular, light music" from the genre dimension, "morning, night, sleep aid" from the scene dimension, "sad, lyrical" from the emotion dimension, "theme A" from the theme dimension, and "artist B" from the artist dimension. The number of dimensions used can be determined based on actual circumstances and is not limited here.
[0144] Step S1620: The words corresponding to the determined tags are determined as keywords of the media object that meet preset requirements.
[0145] For example, the above-mentioned "popular, light music", "early morning, night, sleep aid", "sad, lyrical", "theme A", and "artist B" are determined as media object keywords that meet the preset requirements.
[0146] The present disclosure provides a method for generating a media object collection. Generative AI generates a target collection title based on a set of title keywords; generates a target collection introduction based on a set of introduction keywords; determines a target collection cover based on a set of cover keywords; retrieves media objects from a media object library based on a set of media object keywords to obtain a target media object; and assembles the target collection title, target collection introduction, target collection cover, and target media object to obtain a target media object collection. Generative AI can generate a media object collection based on keywords, reducing labor costs while meeting users' personalized needs and maintaining diversity.
[0147] Exemplary devices
[0148] After introducing the method for generating a media object collection according to an exemplary embodiment of the present disclosure, Figure 17 An apparatus for generating a media object collection according to an exemplary embodiment of the present disclosure is described.
[0149] refer to Figure 17 As shown, a device for generating a media object collection includes: a title generation module 1710, configured to generate a target collection title through generative artificial intelligence based on title keywords; a description generation module 1720, configured to generate a target collection description through generative artificial intelligence based on description keywords; a cover determination module 1730, configured to determine a target collection cover based on cover keywords; a media object determination module 1740, configured to recall media objects from a media object library based on media object keywords to obtain target media objects; an assembly module 1750, configured to assemble the target collection title, target collection description, target collection cover and target media objects to obtain a target media object collection.
[0150] In one embodiment, the title generation module 1710 is configured to: determine the target collection title prompt template selected by the user from the collection title prompt template library; obtain the title prompt and title keywords input by the user based on the target collection title prompt template; and generate the target collection title based on the title prompt and title keywords through a text-based large model.
[0151] In one embodiment, the introduction generation module 1720 is configured to: determine the target collection introduction prompt template selected by the user from the collection introduction prompt template library; obtain the introduction prompt and introduction keywords input by the user based on the target collection introduction prompt template; and generate the target collection introduction based on the introduction prompt and introduction keywords through a text-based large model.
[0152] In one embodiment, the cover determination module 1730 is configured to: determine the target collection cover prompt template selected by the user from the collection cover prompt template library; obtain the cover prompt and cover keywords input by the user based on the target collection cover prompt template; and generate the target collection cover through a large picture model based on the target collection title, target collection introduction, cover prompt and cover keywords.
[0153] In one embodiment, the cover determination module 1730 is configured to: obtain cover keywords; match the cover keywords with words corresponding to the labels of the collection covers in the collection cover database; and determine the collection cover to which the label corresponding to the word with the highest similarity to the cover keyword belongs as the target collection cover.
[0154] In one embodiment, the media object determination module 1740 is configured to: obtain a media object keyword; and perform intersection / subtraction selection on media objects in the media object library based on the media object keyword to obtain a target media object.
[0155] In one embodiment, the media object determination module 1740 is configured to: determine a selection condition according to a media object keyword; and perform intersection-and-subtraction selection on media objects in the media object library according to the selection condition to obtain a target media object.
[0156] In one embodiment, the media object determination module 1740 is configured to: obtain a media object keyword; and retrieve media objects from a media object library based on the media object keyword to obtain a target media object.
[0157] In one embodiment, the device also includes a label mapping module, which is configured to: determine whether the media object keywords meet the preset requirements; if the media object keywords do not meet the preset requirements, perform label mapping on the media object keywords in the media object keywords to obtain media object keywords that meet the preset requirements.
[0158] In one embodiment, the tag mapping module is configured to: determine corresponding tags from tag databases of different dimensions based on media object keywords; and determine the words corresponding to the determined tags as media object keywords that meet preset requirements.
[0159] Exemplary Storage Media
[0160] The storage medium according to the exemplary embodiment of the present disclosure will be described below.
[0161] In this exemplary embodiment, the above method can be implemented by a program product, such as a portable compact disc read-only memory (CD-ROM) that includes program code and can be executed on a device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0162] The program product can be implemented in any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0163] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0164] The program code contained on the readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RE, etc., or any suitable combination of the foregoing.
[0165] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0166] Exemplary electronic devices
[0167] refer to Figure 18 An electronic device according to an exemplary embodiment of the present disclosure will be described.
[0168] Figure 18 The electronic device 1800 shown is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0169] like Figure 18 As shown, electronic device 1800 is implemented as a general-purpose computing device. Components of electronic device 1800 may include, but are not limited to, at least one processing unit 1810, at least one storage unit 1820, a bus 1830 connecting various system components (including storage unit 1820 and processing unit 1810), and a display unit 1840.
[0170] The storage unit stores program codes, which can be executed by the processing unit 1810, so that the processing unit 1810 performs the steps described in the "Exemplary Method" section of the present disclosure according to various exemplary embodiments. For example, the processing unit 1810 can perform the following steps: Figure 1 The method steps shown, etc.
[0171] The storage unit 1820 may include a volatile storage unit, such as a random access memory unit (RAM) 1821 and / or a cache memory unit 1822 , and may further include a read-only memory unit (ROM) 1823 .
[0172] The storage unit 1820 may also include a program / utility 1824 having a set (at least one) of program modules 1825, such program modules 1825 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0173] The bus 1830 may include a data bus, an address bus, and a control bus.
[0174] The electronic device 1800 can also communicate with one or more external devices 2000 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.). This communication can be performed via an input / output (I / O) interface 1850. The electronic device 1800 also includes a display unit 1840, which is connected to the input / output (I / O) interface 1850 for display. Furthermore, the electronic device 1800 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 1860. As shown, the network adapter 1860 communicates with other modules of the electronic device 1800 via a bus 1830. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 1800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0175] It should be noted that although several modules or submodules of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in a single unit / module. Conversely, the features and functions of a single unit / module described above can be further divided and embodied by multiple units / modules.
[0176] Furthermore, although the operations of the disclosed method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0177] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division into various aspects does not mean that the features in these aspects cannot be combined to benefit. Such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the appended claims.
Claims
1. A method for generating a media object collection, characterized in that: The method comprises: Generate target collection titles using generative AI based on title keywords; Generate a target collection profile using generative AI based on the profile keywords; Determine the target album cover based on cover keywords; According to the media object keywords, media objects are recalled from the media object library to obtain the target media object; The target collection title, the target collection introduction, the target collection cover and the target media object are assembled to obtain a target media object collection.
2. The method according to claim 1, characterized in that The target collection title is generated by generative artificial intelligence based on the title keywords, including: Determining a target collection title prompt template selected by the user from a collection title prompt template library; Obtaining the title prompt and title keywords input by the user according to the target collection title prompt template; The target collection title is generated according to the title prompt and the title keywords through a large text model.
3. The method according to claim 1, characterized in that The generating of the target collection introduction by generative artificial intelligence based on the introduction keywords includes: Determine a target collection introduction prompt template selected by the user from a collection introduction prompt template library; Obtaining the introduction prompt and introduction keywords input by the user according to the introduction prompt template of the target collection; The target collection introduction is generated according to the introduction prompts and the introduction keywords through a text-based large model.
4. The method according to claim 1, wherein Determining the target album cover according to the cover keyword includes: Determining a target album cover prompt template selected by the user from an album cover prompt template library; Obtaining the cover hint and cover keywords input by the user according to the target album cover hint template; The target album cover is generated through a large picture model according to the target album title, the target album introduction, the cover prompt and the cover keywords.
5. The method according to claim 1, characterized in that Determining the target album cover according to the cover keyword includes: Obtaining the cover keywords; Matching the cover keywords with words corresponding to labels of the album covers in the album cover database; The album cover to which the tag corresponding to the word with the highest similarity to the cover keyword belongs is determined as the target album cover.
6. The method according to claim 1, characterized in that The step of retrieving the media object from the media object library according to the media object keyword to obtain the target media object includes: Obtaining the media object keyword; According to the media object keywords, the media objects in the media object library are selected by intersection and difference to obtain the target media object.
7. The method according to claim 6, characterized in that The step of performing intersection, union, and difference selection on the media objects in the media object library according to the media object keywords to obtain the target media object includes: Determining a selection condition based on the media object keywords; The media objects in the media object library are subjected to intersection and difference selection according to the selection condition to obtain the target media object.
8. A device for generating a media object collection, characterized in that: The device comprises: A title generation module is configured to generate a target collection title through generative artificial intelligence based on the title keywords; An introduction generation module is configured to generate an introduction to a target collection through generative artificial intelligence based on introduction keywords; A cover determination module is configured to determine a target album cover based on cover keywords; The media object determination module is configured to recall media objects from the media object library according to the media object keywords to obtain a target media object; The assembling module is configured to assemble the target collection title, the target collection introduction, the target collection cover and the target media object to obtain a target media object collection.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the method according to any one of claims 1 to 7 by executing the executable instructions.
Citation Information
Patent Citations
Song menu display information generation method and device, electronic equipment and storage medium
CN114943006A
Video search result-free processing method, system and equipment and medium
CN117453950A
System and method for using a list of audio media to create a list of audiovisual media
US20120232681A1
Generating titles for content segments of media items using machine-learning
US20230402065A1