Memory album generation method and device, electronic equipment and readable storage medium
By combining cluster analysis and user portraits, the memory album is generated, and the problem of high manual intervention cost in the prior art is solved, and the effect of reducing generation costs and improving creativity is achieved.
Patent Information
- Application Number
- CN202510190300.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-16
AI Technical Summary
When generating recollection albums, developers need to pre-develop a large number of album themes, editing templates and matching rules in the prior art, resulting in high cost of manual intervention and corresponding increase in generation costs.
By determining image feature information, clustering analysis, obtaining image sets and themes, determining clip templates based on user portraits, and generating memory albums.
Eliminate users to manually generate memory albums, reducing generation costs and generating rich and diverse themes, music, special effects and copywriting through AI technology, improving creativity and imagination.
Smart Images

Figure CN120011580A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a memory album generation method, device, electronic device and readable storage medium. Background Art
[0002] A memory album is an automatically created personalized collection of images or videos. Currently, electronic devices can automatically identify important people, places, and events based on the images and videos in the local albums, and generate corresponding memory albums. These memory albums usually contain selected images, videos, and music to help users relive the good times.
[0003] In the related art, developers are required to prepare multiple optional album themes in advance, create multiple optional editing templates (including text, music, special effects, transitions, etc.) for different album themes, and create different matching rules to form a resource library. When generating a memory album, some candidate images are obtained from the local album of the electronic device, and album themes and editing templates that match these candidate images are selected from the resource library according to the matching rules. The memory album is generated based on these candidate images, matching album themes and editing templates.
[0004] Although the above method can generate a memory album, since the image content of the albums in the electronic devices of different users is very diverse, developers are required to prepare a large number of album themes, create a large number of editing templates and matching rules in advance to cover the image content of the albums in the electronic devices of different users. The cost of manual intervention is relatively high, resulting in a relatively high cost for generating a memory album. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide a memory album generation method, device, electronic device and readable storage medium, which can reduce the generation cost of the memory album.
[0006] In a first aspect, an embodiment of the present application provides a method for generating a memory album, the method comprising:
[0007] Determining image feature information of each first image;
[0008] performing cluster analysis on each of the first images according to image feature information of each of the first images to obtain a first image set and a theme corresponding to the first image set;
[0009] Determining a clipping template corresponding to the first image set according to the user portrait and the first image set;
[0010] A memory album is generated according to the first image set, a theme corresponding to the first image set, and a editing template corresponding to the first image set.
[0011] In a second aspect, an embodiment of the present application provides a memory album generation device, the device comprising:
[0012] A first determining module, used to determine image feature information of each first image;
[0013] A clustering module, configured to perform cluster analysis on each of the first images according to image feature information of each of the first images, to obtain a first image set and a theme corresponding to the first image set;
[0014] A second determination module, configured to determine a clipping template corresponding to the first image set according to the user portrait and the first image set;
[0015] A generating module is used to generate a memory album according to the first image set, a theme corresponding to the first image set and a editing template corresponding to the first image set.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the memory album generation method described in the first aspect are implemented.
[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the memory album generation method described in the first aspect are implemented.
[0018] In a fifth aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the memory album generation method as described in the first aspect.
[0019] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the memory album generation method as described in the first aspect.
[0020] In an embodiment of the present application, image feature information of each first image is determined; based on the image feature information of each first image, a cluster analysis is performed on each first image to obtain a first image set and a theme corresponding to the first image set; based on the user portrait and the first image set, a clipping template corresponding to the first image set is determined; based on the first image set, the theme corresponding to the first image set, and the clipping template corresponding to the first image set, a memory album is generated. Since there is no need for the user to manually generate the memory album, the generation cost of the memory album can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of a memory album generation method provided by some embodiments of the present application;
[0022] Figure 2 is a structural block diagram of a memory album generating device provided by some embodiments of the present application;
[0023] Figure 3 is a schematic diagram of the structure of an electronic device provided by some embodiments of the present application;
[0024] Figure 4 It is a schematic diagram of the hardware structure of an electronic device implementing various embodiments of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0026] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0027] To facilitate understanding, some relevant concepts and application scenarios involved in the embodiments of the present application are first introduced below.
[0028] 1. Related concepts
[0029] Local albums refer to the collection of images and videos stored locally on electronic devices (such as smartphones, computers, etc.). On smartphones, local albums are usually located in the phone's built-in photo album or gallery app, which contains all the images and videos taken, and can be browsed by time, location, person, etc. On computers, images are usually saved in a folder named "Pictures" or stored and managed using specific image management software.
[0030] A memory album is an automatically created personalized collection of images or videos that can automatically identify important people, places, and events based on images and videos in the local albums of the user's electronic device, and generate corresponding memory albums. These memory albums usually contain selected images, videos, and music to help users relive good times. For example, the cute pet-themed memory album aggregates various cute pet images in the local albums of the user's electronic device.
[0031] Artificial Intelligence (AI) is a technology that enables machines to "think" and "learn" like humans. It simulates human intelligence to enable machines to process various information such as language, audio, images, and videos, and to learn and infer intelligently from them. Artificial Intelligence is not a single technology, but a combination of multiple technologies and algorithms including deep learning, machine learning, computer vision, and natural language processing.
[0032] AI capabilities refer to the technology and science of machine simulation of human intelligence, enabling computer systems to perform tasks that normally require human intelligence, such as learning, reasoning, perception, and problem solving.
[0033] Artificial Intelligence Generated Content (AIGC) is a technology that uses artificial intelligence technology to automatically generate various types of content (such as text, images, audio, video, etc.) by learning and analyzing large amounts of data.
[0034] Large models refer to machine learning models with large-scale parameters and complex computational structures. These models are usually built from deep neural networks and have billions or even hundreds of billions of parameters. Due to their huge size, large models have very powerful representation and generalization capabilities and can perform well in tasks in various fields, such as speech recognition, natural language processing, computer vision, and other fields.
[0035] 2. Application Scenarios
[0036] In the related art, developers are required to prepare multiple optional album themes in advance, create multiple optional editing templates (including text, music, special effects, transitions, etc.) for different album themes, and create different matching rules to form a resource library. For example, for images of cats and dogs, developers prepare the album theme of "cute pets" and create 8 different texts, 10 pieces of music, 5 special effects, and 6 transitions for this album theme.
[0037] When generating a memory album, some candidate images are obtained from the local album of the electronic device, and album themes and editing templates matching these candidate images are selected from the resource library according to the matching rules, and the memory album is generated based on these candidate images, matching album themes and editing templates. For example, when there are many images of cats or dogs in the local album of the user's electronic device, the generation of the "Cute Pets" memory album is triggered.
[0038] Although the above method can generate a memory album, due to the wide variety of image content in the local albums of different users' electronic devices, developers need to prepare a large number of album themes in advance, such as "holiday memories", "weekend memories", "memories of a certain day", "monthly memories", "quarterly memories", "annual memories", "travel - within the city", "travel - outside the city", "someone's memories", "owner's annual memories", "group photos", "children's highlights", "Children's Day", "outing", "anniversary", "cute pets", etc. In addition, developers are required to create a large number of editing templates and matching rules for different album themes. For example, memory albums with different themes have different playback forms. The travel theme will be equipped with rotation entry, fast cut-in and cut-out, etc., and more cheerful music; while the anniversary theme will be equipped with fade-in and fade-out, slower music, to create an atmosphere of peace and tranquility, so as to cover the image content of the albums in the electronic devices of different users. The cost of manual intervention is relatively high, resulting in a relatively high cost for generating memory albums.
[0039] In order to solve the above technical problems, the embodiments of the present application provide a memory album generation method, device, electronic device and readable storage medium.
[0040] A memory album generation method provided in an embodiment of the present application is described in detail below in conjunction with the accompanying drawings.
[0041] It should be noted that the memory album generation method provided in the embodiment of the present application is applicable to electronic devices. In practical applications, the electronic devices include but are not limited to: mobile terminals such as smart phones and tablet computers, and computer devices such as desktop computers and servers. The embodiment of the present application does not limit this.
[0042] Figure 1 is a flowchart of a memory album generation method provided by some embodiments of the present application, such as Figure 1 As shown, the method may include the following steps: step 101, step 102, step 103 and step 104.
[0043] In step 101, image feature information of each first image is determined.
[0044] In the embodiment of the present application, the first image is usually an image in a local photo album of the electronic device.
[0045] In the embodiments of the present application, it is taken into account that the sources of images in the local album of the electronic device are relatively wide, for example, images taken by the electronic device itself, images transmitted between other electronic devices via the network or Bluetooth, images downloaded from the network, images from social media applications, etc., and the memory album is usually strongly related to the owner of the electronic device. For example, the memory album contains images of the user, images of the user's family, images of the user's pets, images of places visited by the user, or images of the user's friends, etc. Therefore, when generating a memory album, usually not all images in the local album of the electronic device are used, but images are selectively used, that is, the first image.
[0046] In the embodiment of the present application, the first image includes but is not limited to: an image taken by the electronic device itself, an image transmitted between other electronic devices via a network or Bluetooth, an image from a social media application, and the image contains a portrait and the portrait matches the top five portrait images taken by the device itself.
[0047] In the embodiment of the present application, the image feature information refers to a set of a series of features that can characterize the characteristics or content of the first image.
[0048] In the embodiment of the present application, the image feature information may include at least one of the following: content description information, category label, aesthetic score value, user attention score value, etc.
[0049] In an embodiment of the present application, the content description information is used to describe the picture content of the first image, for example, a little girl running on the beach with a bright smile under the sunset, or camping on the top of a mountain.
[0050] In the embodiment of the present application, the category label is used to characterize the category of the picture element in the first image, for example, landscape, person, cat, or dog.
[0051] In the embodiment of the present application, the aesthetic score value is used to characterize the overall aesthetic feeling (ie, beauty) of the first image, wherein the higher the aesthetic score value of the first image, the better the overall aesthetic feeling of the first image.
[0052] In the embodiment of the present application, the user attention score value is used to characterize the user's preference for the first image, wherein the higher the user attention score value of the first image, the higher the user's preference for the first image.
[0053] It can be understood that the richer the image feature information of the first image, the more information can be referenced when generating a memory album, the more targeted the generated memory album is, the more it can fit the user's emotions and aesthetics, and the higher the user satisfaction. Therefore, preferably, the image feature information includes content description information, category labels, aesthetic score values and user attention score values.
[0054] In some embodiments, when the image feature information includes content description information, category labels, aesthetic score values, and user attention score values, the above step 101 may include the following steps: step 1011, step 1012, step 1013, and step 1014.
[0055] In step 1011, content description information of each first image is determined through an image understanding network.
[0056] In the embodiment of the present application, the image understanding network has AI image understanding capabilities. Through the AI image understanding capabilities of the image understanding network, the picture elements in each first image are identified and marked to form semantic content description information.
[0057] In step 1012, the category label of each first image is determined through an image recognition network.
[0058] In an embodiment of the present application, the image recognition network has AI image recognition capabilities. Through the AI image recognition capabilities of the image recognition network, the picture elements in each first image are classified, for example, people, cats, dogs, natural scenery, parents and children, beaches, etc., to facilitate subsequent clustering of first images of the same category.
[0059] In step 1013, the aesthetic score value of each first image is determined by an aesthetic scoring network.
[0060] In an embodiment of the present application, the aesthetic scoring network has AI aesthetic scoring capabilities. Through the AI aesthetic scoring capabilities of the aesthetic scoring network, scoring is performed based on at least one of the clarity, color, distortion, character expression, action and composition of each first image to obtain an aesthetic score value for each first image.
[0061] For example, the higher the clarity of the first image, the higher the aesthetic score of the first image; the richer and brighter the colors of the first image, the higher the aesthetic score of the first image; the smaller the distortion of the first image, the higher the aesthetic score of the first image; the richer the expressions and movements of the characters in the first image, and the more they meet the set requirements, the higher the aesthetic score of the first image; the better the composition of the first image, the higher the aesthetic score of the first image.
[0062] In the embodiment of the present application, the overall aesthetic feeling of the first image can be scored by AI intelligent scoring from multiple dimensions. On the one hand, no human intervention is required and the cost is relatively low; on the other hand, scoring in multiple dimensions can ensure the accuracy of the aesthetic score value.
[0063] In step 1014, the user attention score value of each first image is determined through a statistical network.
[0064] In an embodiment of the present application, the statistical network has AI statistical capabilities. Through the AI statistical capabilities of the statistical network, a score is performed based on at least one of the historical editing times, historical saving times, historical sharing times, historical collection times, historical browsing times, the number of images of the same category, and the historical update frequency of images of the same category for each first image to obtain a user attention score value for each first image.
[0065] In an embodiment of the present application, if the screen element in the first image is a cat, then images of the same category refer to other images whose screen elements are cats, the number of images of the same category refers to the number of first images whose screen elements are cats, and the historical update frequency of images of the same category refers to the frequency of increase of first images whose screen elements are cats.
[0066] For example, the more times the first image is historically edited, the higher the user attention score of the first image; the more times the first image is historically saved, the higher the user attention score of the first image; the more times the first image is historically shared, the higher the user attention score of the first image; the more times the first image is historically collected, the higher the user attention score of the first image; the more times the first image is historically browsed, the higher the user attention score of the first image; the more images of the same category as the first image, the higher the user attention score of the first image; the higher the historical update frequency of images of the same category as the first image, the higher the user attention score of the first image.
[0067] In the embodiment of the present application, the user attention of the first image can be scored by AI intelligent scoring from multiple dimensions. On the one hand, no human intervention is required and the cost is relatively low; on the other hand, scoring in multiple dimensions can ensure the accuracy of the user attention score value.
[0068] In an embodiment of the present application, each of the above-mentioned image understanding network, image recognition network, aesthetic scoring network and statistical network can be an independent AI model or a sub-network of a large AI model, which is not limited in the embodiment of the present application.
[0069] It can be seen that in the embodiment of the present application, the image feature information of each first image can be generated with the help of the AI capability of the network. Since the AI capability can learn the prior knowledge of the image feature information of a large number of images, on the one hand, no manual intervention is required and the cost is relatively low; on the other hand, accurate image feature information of each first image can be generated.
[0070] In step 102, cluster analysis is performed on each first image according to image feature information of each first image to obtain a first image set and a theme corresponding to the first image set.
[0071] In the embodiment of the present application, since the image feature information is a collection of a series of features that can characterize the characteristics or content of the first image, the commonalities and differences between different first images can be analyzed based on the image feature information of each first image, thereby achieving accurate clustering of each first image.
[0072] In the embodiment of the present application, when clustering analysis is performed on each first image according to the image feature information of each first image, one image set or multiple image sets can be obtained by clustering. When one image set is obtained by clustering, the image set is the first image set. When multiple image sets are obtained by clustering, the first image set can be one of the multiple image sets. For example, three image sets with similar image contents are obtained by clustering, and the image set with the highest image quality is selected as the first image set.
[0073] In some embodiments, AI capabilities may be used to perform image clustering. Accordingly, the above step 102 may include the following steps: step 1021 .
[0074] In step 1021, cluster analysis is performed on each first image based on image feature information of each first image through an image clustering network to obtain a first image set and a theme corresponding to the first image set.
[0075] In an embodiment of the present application, the image clustering network has AI image clustering capabilities. Through the AI image clustering capabilities of the image clustering network, each first image can be clustered and analyzed based on the image feature information of each first image to obtain a first image set. Thereafter, the user's personalized interest preferences are identified for the first image set. For example, it is identified that the user likes online games, camping and mountaineering, etc., and the theme of the image set is formulated based on the user's interest preferences.
[0076] For example, through the image clustering network, the images of the beach in Bali taken three times in the past are clustered, and the same portrait is found, and it is recognized that this is a child in the family. Based on this, the theme of child growth is inferred, and the memory album of this theme will realize the arrangement of images according to the growth path of parent-child travel.
[0077] It can be seen that in the embodiments of the present application, since the image clustering network is trained based on a large number of sample images, it can learn the prior knowledge of the category information and user preference information of the large number of sample images, and the image set is clustered based on the category information, and the theme of the image set is formulated based on the user preference information. Therefore, the image clustering network can cluster and obtain an accurate first image set and the theme of the first image set.
[0078] In step 103, a clipping template corresponding to the first image set is determined according to the user portrait and the first image set.
[0079] In the embodiment of the present application, the user portrait refers to a general description of the user formed by analyzing the user's basic information, behavioral data, consumption habits, etc.
[0080] In the embodiment of the present application, the user portrait may include at least one of the following: basic information, behavioral data, consumption habits, hobbies and interests, and psychological characteristics. Among them, the basic information may include age, gender, region, occupation, education level, etc. Behavioral data may include the user's activity, browsing history, and interactive behavior on the new media platform. Consumption habits may include the type of products purchased by the user, the frequency of consumption, the amount of consumption, etc. Hobbies and interests may include the user's favorite entertainment, sports, reading, etc. Psychological characteristics may include the user's motivations, needs, expectations, etc.
[0081] In the embodiment of the present application, the editing template may include at least one of the following: text, music, special effects and transitions.
[0082] In the embodiment of the present application, the commonality analysis of the user portrait and the first image set can be performed to obtain the editing template corresponding to the first image set. For example, for the same "travel" theme and the same landscape set, the editing templates corresponding to the elderly and the young are different. In the elderly's memory album, the editing template presents the magnificent mountains and rivers of the motherland, and the music also sets off the atmosphere of the beautiful mountains and rivers. In the young people's memory album, the editing template presents the unexpected encounters during the trip, and the music is full of joy and relaxation.
[0083] In some embodiments, in order to improve the editing effect, the first image set may be optimized, and the editing template may be determined based on the optimized first image set. Accordingly, the above step 103 may include the following steps: step 1031 and step 1032.
[0084] In step 1031, filtering is performed on the second image in the first image set according to the filtering strategy, wherein the second image is an image in the first image set whose image quality is lower than an image quality threshold.
[0085] In an embodiment of the present application, the filtering strategy may include at least one of the following: the aspect ratio of the image exceeds a preset aspect ratio range, the source of the image exceeds the image source whitelist, the existence duration of the image is greater than a first threshold, the category label of the image is different from the category labels of other images, the aesthetic score value is lower than a second threshold, the user attention score value is lower than a third threshold, and the similarity between the image and other images is greater than a fourth threshold.
[0086] In the embodiment of the present application, considering that the first image set may contain a large number of images, such as dozens or hundreds of images, while the number of images in the memory album is usually limited, such as a dozen or twenty images containing essential content, the images can be optimized within the first image set and some low-quality images can be filtered out to reduce the number of images in the image set and improve the overall content quality of the images in the image set.
[0087] In the embodiment of the present application, since the display size of the screen of the electronic device is fixed, the display effect of some images of special sizes, such as long screenshots, on the screen is not good, so such images that exceed the preset aspect ratio range are filtered out.
[0088] In the embodiment of the present application, the user can pre-set a whitelist of image sources, and filter out images whose sources exceed the whitelist of image sources to meet the user's personalized requirements.
[0089] In the embodiment of the present application, since the memory album has a certain degree of expiration, such older images in the image collection are filtered out.
[0090] In an embodiment of the present application, since some image sets are image sets of specific image categories, for example, an image set with the theme of "cat", if some images of dogs or other animals are mixed in the image set, it will be inconsistent with the theme. Therefore, images in the image set whose category labels are different from the category labels of other images are filtered out.
[0091] In the embodiment of the present application, in order to improve the overall quality of the images in the image set, the images with lower aesthetic score values in the image set are filtered out.
[0092] In the embodiment of the present application, in order to better meet the user's interest preferences, images in the image set with low user attention are filtered out.
[0093] In the embodiment of the present application, in order to avoid excessive homogeneity of images in the image set, a portion of images with high similarity in the image set are filtered out.
[0094] In step 1032, commonality analysis is performed on the user portrait and the first image set after filtering to obtain a clipping template corresponding to the first image set.
[0095] In the embodiment of the present application, AIGC technology can be used to perform commonality analysis on the user portrait and the first image set to obtain a editing template corresponding to the first image set, that is, automatically generate an adaptive editing template, which can expand the resource library infinitely. Whether it is themes, text descriptions, transitions, special effects or music, they are no longer limited to a few fixed types manually formulated, enriching the types and quantities of editing templates and improving the creativity and imagination of the finally generated memory album.
[0096] It can be seen that in the embodiment of the present application, since the overall image quality of the first image set after filtering is higher, determining the clipping template based on the user portrait and the first image set after filtering can improve the adaptability of the determined clipping template to the first image set.
[0097] In some embodiments, when the editing template includes text, music, special effects and transitions, the above step 103 may include the following steps: step 1033, step 1034 and step 1035.
[0098] In step 1033 , a cover image of the first image set is determined.
[0099] In the embodiment of the present application, the cover image may include at least one of the following: a selected image in the first image set or an image that has received the most attention from users.
[0100] In an embodiment of the present application, when determining the cover image of the first image set, the following rules can be followed: first, select the batch of images in the first image set that is closest to the current time; then select a selected image from the most recent batch (the electronic device automatically determines it according to the strategy); if there are multiple selected images, select the selected image with the highest user attention as the cover image; if there is no selected image in the most recent batch, select the image with the highest user attention as the cover image.
[0101] In step 1034 , the copy corresponding to the first image set is determined based on the user portrait, the first image set, and the cover image of the first image set.
[0102] In the embodiment of the present application, the length of the text can be limited to 2-8 words to ensure that the text is concise and intuitive.
[0103] In the embodiment of the present application, the copywriting network has AI sentiment analysis capabilities and AI copywriting capabilities. Through the AI sentiment analysis capabilities of the copywriting network, a comprehensive sentiment analysis is performed on the user portrait, the first image set, and the cover image of the first image set. Through the AI copywriting capabilities of the copywriting network, the copywriting corresponding to the first image set is accurately written based on the sentiment information obtained through the analysis. For example, for the theme of "summer", if it is for young people, the copywriting is "summer light years"; if it is for middle-aged people, the copywriting is "summer memories".
[0104] In step 1035, music, special effects and transitions corresponding to the first image set are determined based on the user portrait, the first image set, the cover image of the first image set and the text corresponding to the first image set.
[0105] In an embodiment of the present application, AIGC technology can be used to more accurately generate music, special effects and transitions corresponding to the first image set based on the user portrait, the first image set, the cover image of the first image set and the text corresponding to the first image set, thereby expanding the types and number of editing templates.
[0106] It can be seen that in the embodiment of the present application, since the cover image of the memory album is usually the most representative image, the cover image is determined first, and then the case is determined based on the cover image. Finally, the transition, special effects, music, etc. are determined based on the image set, the theme of the image set, the cover image and the text. This enables the determined editing template to be accurately adapted to the image set, thereby improving the overall audio-visual experience of the generated memory album.
[0107] In step 104, a memory album is generated according to the first image set, the theme corresponding to the first image set, and the editing template corresponding to the first image set.
[0108] In the embodiment of the present application, if the first image set is filtered, a memory album is generated based on the filtered first image set, the theme corresponding to the first image set, and the editing template corresponding to the first image set. If the first image set is not filtered, a memory album is directly generated based on the original first image set, the theme corresponding to the first image set, and the editing template corresponding to the first image set.
[0109] In the embodiment of the present application, all images in the image set are edited and laid out according to the theme and editing template corresponding to the image set to obtain the memory album corresponding to the image set. In addition, the memory album can be further optimized and edited, for example, the size of all images in the memory album is unified to the same size.
[0110] In the embodiment of the present application, after the memory album is generated, the memory album can be pushed to the user based on time information, such as this day in previous years. Alternatively, the memory album can be pushed to the user based on behavior information, such as when the user browses a certain image. Alternatively, the memory album can be pushed to the user based on location information, such as when the permanent residence changes. Alternatively, the memory album can be pushed to the user based on artificial settings, such as when a new memory album is generated.
[0111] It can be seen that in the embodiment of the present application, the AI capability of the electronic device is used to intelligently generate a memory album. On the one hand, the labor cost of preparing album themes and editing templates is reduced. On the other hand, the limitations of manually imported content are broken. The types of themes, music, special effects, transitions, and text descriptions generated by AI are very rich, and the number can basically be understood as unlimited. They are no longer limited to a few fixed types drawn up by humans, which improves creativity and imagination.
[0112] It can be seen from the above embodiments that in this embodiment, the image feature information of each first image is determined; based on the image feature information of each first image, a cluster analysis is performed on each first image to obtain a first image set and a theme corresponding to the first image set; based on the user portrait and the first image set, a editing template corresponding to the first image set is determined; based on the first image set, the theme corresponding to the first image set and the editing template corresponding to the first image set, a memory album is generated. Since the user does not need to manually generate the memory album, the generation cost of the memory album can be reduced.
[0113] The memory album generation method provided in the embodiment of the present application can be executed by a memory album generation device. In the embodiment of the present application, the memory album generation method executed by the memory album generation device is taken as an example to illustrate the memory album generation device provided in the embodiment of the present application.
[0114] Figure 2 is a structural block diagram of a memory album generation device provided by some embodiments of the present application, such as Figure 2 As shown, the memory album generating device 200 may include: a first determining module 201, a clustering module 202, a second determining module 203 and a generating module 204;
[0115] The first determining module 201 is used to determine image feature information of each first image;
[0116] The clustering module 202 is used to perform cluster analysis on each of the first images according to image feature information of each of the first images to obtain a first image set and a theme corresponding to the first image set;
[0117] The second determination module 203 is used to determine a clipping template corresponding to the first image set according to the user portrait and the first image set;
[0118] The generating module 204 is used to generate a memory album according to the first image set, the theme corresponding to the first image set and the editing template corresponding to the first image set.
[0119] It can be seen from the above embodiments that in this embodiment, the image feature information of each first image is determined; based on the image feature information of each first image, a cluster analysis is performed on each first image to obtain a first image set and a theme corresponding to the first image set; based on the user portrait and the first image set, a editing template corresponding to the first image set is determined; based on the first image set, the theme corresponding to the first image set and the editing template corresponding to the first image set, a memory album is generated. Since the user does not need to manually generate the memory album, the generation cost of the memory album can be reduced.
[0120] Optionally, as an embodiment, the image feature information may include: content description information, category label, aesthetic score value and user attention score value;
[0121] The first determining module 201 may include:
[0122] A first determination submodule, configured to determine content description information of each first image through an image understanding network;
[0123] A second determination submodule, used to determine the category label of each first image through an image recognition network;
[0124] A third determination submodule is used to determine the aesthetic score value of each first image through an aesthetic scoring network;
[0125] The fourth determination submodule is used to determine the user attention score value of each first image through a statistical network.
[0126] Optionally, as an embodiment, the second determining module 203 may include:
[0127] a filtering submodule, configured to filter the second image in the first image set according to a filtering strategy, wherein the second image is an image in the first image set whose image quality is lower than an image quality threshold;
[0128] The analysis submodule is used to perform commonality analysis on the user portrait and the first image set after filtering to obtain a clipping template corresponding to the first image set.
[0129] The generating module 204 may include:
[0130] A generating submodule is used to generate a memory album according to the first image set after filtering, the theme corresponding to the first image set and the editing template corresponding to the first image set.
[0131] Optionally, as an embodiment, the third determining submodule may include:
[0132] The first determination unit is used to score each first image based on at least one of clarity, color, distortion, character expression, action and composition through an aesthetic scoring network to obtain an aesthetic score value of each first image.
[0133] Optionally, as an embodiment, the fourth determining submodule may include:
[0134] The second determination unit is used to score each first image through a statistical network based on at least one of the historical editing times, historical saving times, historical sharing times, historical collection times, historical browsing times, the number of images of the same category, and the historical update frequency of images of the same category, so as to obtain a user attention score value for each first image.
[0135] Optionally, as an embodiment, the clustering module 202 may include:
[0136] The clustering submodule is used to perform cluster analysis on each of the first images based on image feature information of each of the first images through an image clustering network to obtain a first image set and a theme corresponding to the first image set.
[0137] Optionally, as an embodiment, the editing template includes: text, music, special effects and transitions;
[0138] The second determining module 203 may include:
[0139] a fifth determining submodule, configured to determine a cover image of the first image set;
[0140] a sixth determination submodule, configured to determine a text corresponding to the first image set based on the user portrait, the first image set, and the cover image of the first image set;
[0141] The seventh determination submodule is used to determine the music, special effects and transitions corresponding to the first image set based on the user portrait, the first image set, the cover image of the first image set and the text corresponding to the first image set.
[0142] The memory album generation device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or it can be other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-Mobile Personal Computer, UMPC), a netbook or a personal digital assistant (Personal Digital Assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.
[0143] The memory album generation device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0144] The memory album generation device provided in the embodiment of the present application can achieve the above Figure 1 To avoid repetition, the various processes implemented in the illustrated method embodiment will not be described again here.
[0145] Alternatively, if Figure 3 As shown, an embodiment of the present application further provides an electronic device 300, including a processor 301 and a memory 302, wherein the memory 302 stores programs or instructions that can be executed on the processor 301, and when the program or instructions are executed by the processor 301, the various steps of the above-mentioned memory album generation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, they are not described here.
[0146] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0147] Figure 4 It is a schematic diagram of the hardware structure of an electronic device implementing various embodiments of the present application.
[0148] The electronic device 400 includes but is not limited to components such as a radio frequency unit 401 , a network module 402 , an audio output unit 403 , an input unit 404 , a sensor 405 , a display unit 406 , a user input unit 407 , an interface unit 408 , a memory 409 , and a processor 410 .
[0149] Those skilled in the art will appreciate that the electronic device 400 may also include a power source (such as a battery) for supplying power to each component, and the power source may be logically connected to the processor 410 through a power management system, thereby implementing functions such as managing charging, discharging, and power consumption management through the power management system. Figure 4 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be described in detail here.
[0150] In some embodiments, the processor 410 is used to determine image feature information of each first image; perform cluster analysis on each first image based on the image feature information to obtain a first image set and a theme corresponding to the first image set; determine a clipping template corresponding to the first image set based on the user portrait and the first image set; and generate a memory album based on the first image set, the theme corresponding to the first image set, and the clipping template corresponding to the first image set.
[0151] It can be seen that in the embodiment of the present application, since the user does not need to manually generate a memory album, the generation cost of the memory album can be reduced.
[0152] Optionally, as an embodiment, the image feature information includes: content description information, category label, aesthetic score value and user attention score value;
[0153] Processor 410 is specifically used to determine the content description information of each first image through an image understanding network; determine the category label of each first image through an image recognition network; determine the aesthetic score value of each first image through an aesthetic scoring network; and determine the user attention score value of each first image through a statistical network.
[0154] Optionally, as an embodiment, the processor 410 is specifically used to filter the second image in the first image set according to a filtering strategy, wherein the second image is an image in the first image set whose image quality is lower than an image quality threshold; perform commonality analysis on the user portrait and the first image set after filtering to obtain a clipping template corresponding to the first image set; and generate a memory album according to the first image set after filtering, the theme corresponding to the first image set, and the clipping template corresponding to the first image set.
[0155] Optionally, as an embodiment, the processor 410 is specifically used to score each first image based on at least one of clarity, color, distortion, character expression, action and composition through an aesthetic scoring network to obtain an aesthetic score value of each first image.
[0156] Optionally, as an embodiment, the processor 410 is specifically used to score each first image through a statistical network based on at least one of its historical editing times, historical saving times, historical sharing times, historical collection times, historical browsing times, the number of images of the same category, and historical update frequency of images of the same category, so as to obtain a user attention score value for each first image.
[0157] Optionally, as an embodiment, the processor 410 is specifically configured to perform cluster analysis on each of the first images based on image feature information of each of the first images through an image clustering network to obtain a first image set and a theme corresponding to the first image set.
[0158] Optionally, as an embodiment, the editing template includes: text, music, special effects and transitions; the processor 410 is specifically used to determine the cover image of the first image set; based on the user portrait, the first image set and the cover image of the first image set, determine the text corresponding to the first image set; based on the user portrait, the first image set, the cover image of the first image set and the text corresponding to the first image set, determine the music, special effects and transitions corresponding to the first image set.
[0159] It should be understood that in the embodiment of the present application, the input unit 404 may include a graphics processor (Graphics Processing Unit, GPU) 4041 and a microphone 4042, and the graphics processor 4041 processes the image data of the static picture or video obtained by the image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 406 may include a display panel 4061, and the display panel 4061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 407 includes a touch panel 4071 and at least one of other input devices 4072. The touch panel 4071 is also called a touch screen. The touch panel 4071 may include two parts: a touch detection device and a touch controller. Other input devices 4072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0160] The memory 409 can be used to store software programs and various data. The memory 409 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 409 may include a volatile memory or a non-volatile memory, or the memory 409 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 409 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0161] The processor 410 may include one or more processing units; optionally, the processor 410 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 410.
[0162] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned memory album generation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0163] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0164] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned memory album generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0165] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0166] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned memory album generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0167] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0168] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0169] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A method for generating a memory album, characterized in that: The method comprises: Determining image feature information of each first image; performing cluster analysis on each of the first images according to image feature information of each of the first images to obtain a first image set and a theme corresponding to the first image set; Determining a clipping template corresponding to the first image set according to the user portrait and the first image set; A memory album is generated according to the first image set, a theme corresponding to the first image set, and a editing template corresponding to the first image set.
2. The method according to claim 1, characterized in that The image feature information includes: content description information, category label, aesthetic score value and user attention score value; The determining of the image feature information of each first image includes: Determining content description information of each first image through an image understanding network; Determine the category label of each first image through an image recognition network; Determining an aesthetic score value of each first image through an aesthetic scoring network; The user attention score value of each first image is determined through a statistical network.
3. The method according to claim 2, characterized in that The determining, according to the user portrait and the first image set, a clipping template corresponding to the first image set includes: Performing filtering processing on a second image in the first image set according to a filtering strategy, wherein the second image is an image in the first image set whose image quality is lower than an image quality threshold; Performing commonality analysis on the user portrait and the first image set after filtering to obtain a clipping template corresponding to the first image set; The step of generating a memory album according to the first image set, a theme corresponding to the first image set, and a clipping template corresponding to the first image set includes: A memory album is generated according to the first image set after filtering, the theme corresponding to the first image set, and the editing template corresponding to the first image set.
4. The method according to claim 2, characterized in that: Determining the aesthetic score value of each first image through the aesthetic scoring network includes: The aesthetic scoring network is used to score each first image based on at least one of its clarity, color, distortion, character expression, action, and composition, to obtain an aesthetic score value for each first image.
5. The method according to claim 2, characterized in that: Determining the user attention score value of each first image through the statistical network includes: Through the statistical network, scores are performed based on at least one of the historical editing times, historical saving times, historical sharing times, historical collection times, historical browsing times, the number of images of the same category and the historical update frequency of images of the same category for each first image, so as to obtain the user attention score value of each first image.
6. The method according to claim 1, characterized in that The step of performing cluster analysis on each of the first images according to the image feature information of each of the first images to obtain a first image set and a theme corresponding to the first image set includes: By using an image clustering network, cluster analysis is performed on each of the first images based on image feature information of each of the first images to obtain a first image set and a theme corresponding to the first image set.
7. The method according to claim 1, characterized in that The editing template includes: copywriting, music, special effects and transitions; The determining, according to the user portrait and the first image set, a clipping template corresponding to the first image set includes: determining a cover image for the first image set; Determining a copy corresponding to the first image set based on the user portrait, the first image set, and the cover image of the first image set; Based on the user portrait, the first image set, the cover image of the first image set and the text corresponding to the first image set, music, special effects and transitions corresponding to the first image set are determined.
8. A memory album generation device, characterized in that: The device comprises: A first determining module, used to determine image feature information of each first image; A clustering module, configured to perform cluster analysis on each of the first images according to image feature information of each of the first images, to obtain a first image set and a theme corresponding to the first image set; A second determination module, configured to determine a clipping template corresponding to the first image set according to the user portrait and the first image set; A generating module is used to generate a memory album according to the first image set, a theme corresponding to the first image set and a editing template corresponding to the first image set.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the memory album generation method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps of the memory album generation method as described in any one of claims 1 to 7 are implemented.