Atlas acquisition method, electronic device, and computer-readable storage medium

CN120561329BActive Publication Date: 2026-08-11HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]然而,传统的图片展示方式比较单一,用户体验不高

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561329B_ABST
    Figure CN120561329B_ABST
Patent Text Reader

Abstract

This application relates to the field of terminal technology, providing a method for acquiring image sets, an electronic device, and a computer-readable storage medium. The method includes acquiring at least one component of a target recommendation template, wherein the at least one component includes at least a time element and / or a people element; converting initial keywords of the time element and / or people element to obtain time search keywords and / or people search keywords; and searching a gallery based on the search keywords to obtain a target image set matching the search keywords, wherein the search keywords include at least time search keywords and / or people search keywords. This method can obtain corresponding image sets based on the image elements of an image, which can be used to recommend images to users, thereby personalizing the display of images according to the recommendation template and improving the user's browsing experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, specifically to a method for acquiring image atlases, an electronic device, and a computer-readable storage medium. Background Technology

[0002] With the development of internet and terminal technologies, terminal devices are being used more and more widely and deeply into people's production and lives, becoming an inseparable part of them.

[0003] Modern mobile devices are no longer limited to traditional communication functions; they also have the ability to take photos and store them. When a mobile device stores a large number of photos, they can be categorized for easier viewing. For example, a user can open the Gallery app, click on the "Camera" folder to view photos taken with the camera app, or click on the "Screenshots" folder to view screenshots.

[0004] However, traditional image display methods are rather simplistic and offer a poor user experience. Summary of the Invention

[0005] This application provides a method, apparatus, chip, electronic device, computer-readable storage medium, and computer program product for acquiring atlases, which can improve the user experience.

[0006] Firstly, a method for obtaining an image set is provided, comprising: obtaining at least one component of a target recommendation template, wherein the at least one component includes at least a time element and / or a person element; converting the initial keywords of the time element and / or person element to obtain time search keywords and / or person search keywords; searching in an image library according to the search keywords to obtain a target image set that matches the search keywords, wherein the search keywords include at least a time search keyword and / or a person search keyword; wherein the first image is any image in the target image set, the image elements of the first image match at least one component, and the image elements of the first image include at least the generation time of the first image and / or a person identifier.

[0007] The constituent elements can be any one or more of time, location, people, and event elements. These constituent elements can be described using configuration information in a configuration file. The initial keywords are descriptive terms that conform to user's colloquial language. Elements of the time element (initial keywords) can be descriptive terms that conform to the user's verbal expression, such as "today," "tomorrow," or "last week." Elements of the people element (initial keywords) can be descriptive terms that conform to the user's verbal expression, such as "Dabao" or "lover." After converting the initial keywords, the terminal device obtains the corresponding time search keywords and / or people search keywords. The time search keyword is a specific date, and the people search keyword is the corresponding person's TagID or ReID. The terminal device can search the image library based on some or all of the time search keywords, people search keywords, location search keywords, and event search keywords to obtain the target image set matching the search keywords.

[0008] Terminal devices can derive corresponding image sets based on image elements and recommend them to users, thus personalizing image display according to recommended templates and enhancing the user's browsing experience. Furthermore, by converting colloquial initial keywords into corresponding search keywords that the terminal device can recognize, image library searches can be achieved. Additionally, the colloquial initial keywords result in recommendation statements that align with users' conversational expression habits, leading to a better reading experience.

[0009] In some possible implementations, at least one component includes: a first component and other components. The first component is any one of a time element, a person element, a location element, and an event element. The other components include one or more of the time element, person element, location element, and event element, and each of the first component and other components is different. When the first component includes at least two initial keywords, the method further includes: combining each of the at least two initial keywords with the initial keywords of the other components to form at least two initial keyword combinations. The first initial keyword combination is any one of the at least two initial keyword combinations. The first initial keyword combination includes multiple initial keywords, including a first initial keyword and a second initial keyword. The component corresponding to the first initial keyword and the component corresponding to the second initial keyword are different components among at least one component. The multiple initial keywords include at least two initial keywords. Any one of the following: Transform the initial keywords of the time element and / or the person element to obtain time search keywords and / or person search keywords, including: when multiple initial keywords in the first initial keyword combination include time keywords, transform the time keywords to obtain time search keywords corresponding to the time keywords; when multiple initial keywords in the first initial keyword combination include person keywords, transform the person keywords to obtain person search keywords corresponding to the person keywords; search the image library according to the search keywords to obtain target image sets matching the search keywords, including: replacing the corresponding time keywords in the first initial keyword combination with time search keywords, and / or replacing the corresponding person keywords in the first initial keyword combination with person search keywords to obtain a first search keyword combination corresponding to the first initial keyword combination, wherein the first search keywords include at least one search keyword; search the image library according to at least one search keyword to obtain target image sets matching at least one search keyword.

[0010] The first component mentioned above can be any one of at least one component, and the other components are different from the first component. Each component can contain empty or one or more elements. Not all elements in a recommendation template will be empty. The terminal device can first combine different elements from different components to form all possible combinations of different elements from multiple components, thereby obtaining at least two initial keyword combinations. Taking the first initial keyword combination as an example, the initial keywords are converted to obtain the corresponding search keywords.

[0011] It should be noted that the search keywords for people mentioned above can be either the names of the people or the feature vectors of the people used to search for them; the search keywords for events mentioned above can be either the names of the events or the semantic feature vectors of the events used to search for them. See the preceding text for details.

[0012] The terminal device searches for the target image library based on the converted search keywords, along with other search keywords that do not require conversion from the same initial keyword group. For example, TagID-1 + February 10, 2024 + City C1 + ClipID-1. Optionally, one or more search keywords in the search keyword combination may be empty; the specific form of the search keyword combination is not limited here.

[0013] Based on this, terminal devices can generate various image sets corresponding to different elements, making the presented image sets more diverse and enriching the user experience.

[0014] In some possible implementations, at least one component also includes: a location element, at least one search keyword also includes a location search keyword, and the image element of the first image also includes the location where the first image was generated.

[0015] In some possible implementations, the location search keyword is one of a number of candidate locations, which are obtained by clustering the locations generated from multiple images in the gallery.

[0016] In some possible implementations, when the location element includes a first referential address, the method further includes: converting the first referential address into a first map address, and including the location name corresponding to the first map address in the location search keywords.

[0017] In some possible implementations, at least one component also includes: a person element, at least one search keyword also includes a person search keyword, and the image element of the first image also includes a person feature vector of the first image, which is used to characterize the person appearing in the first image.

[0018] In some possible implementations, the feature vector of the person in the first image is one or more of a plurality of feature vectors, which are obtained by clustering the faces and / or bodies of people appearing in multiple images in the image library.

[0019] Elements derived from the clustering of images in the image library, specifically person feature vectors and / or location elements, can effectively narrow the search scope, saving search time and resources. Simultaneously, it enables personalized recommendations, meeting user expectations and enhancing the user experience.

[0020] In some possible implementations, at least one component further includes: an event element; at least one search keyword further includes an event search keyword; the image element of the first image further includes a semantic feature vector of the first image; the semantic feature vector of the first image is used to characterize the image content of the first image; and the image content of the first image includes the event corresponding to the event search keyword.

[0021] For detailed descriptions of the constituent elements and image elements, please refer to other relevant descriptions in the text, which will not be repeated here.

[0022] In some possible implementations, the method further includes: obtaining a first recommendation statement template of the target recommendation template, the first recommendation statement template including at least one element slot, the at least one element slot corresponding one-to-one with at least one component element; filling the first element slot with the first initial keyword of the first component element corresponding to the first element slot, generating a first recommendation statement corresponding to the first recommendation statement template, wherein the first element slot is any one of the at least one element slots, the first component element is one of the at least one component elements corresponding to the first element slot, and the first initial keyword is one of the elements of the first component element.

[0023] The terminal device fills the corresponding element slots in the recommendation statement template with the aforementioned initial keywords, thereby generating the corresponding recommendation statement. A detailed description of this can be found in the preceding text and will not be repeated here. The initial keywords in this recommendation statement conform to users' colloquial expression habits, making it easier to read. Furthermore, filling the statement based on element slots offers high flexibility and improves the user experience.

[0024] In some possible implementations, the method further includes: receiving a first operation input by a user, the first operation being used to display an interface of a gallery application; responding to the first operation, displaying a first interface of the gallery application, the first interface including a first control; receiving a second operation performed by the user on the first control; responding to the second operation, displaying a second interface, the second interface displaying a first recommendation statement; receiving a third operation performed by the user on a second control of the first recommendation statement; responding to the third operation, displaying a third interface, the third interface displaying a target gallery.

[0025] In some possible implementations, the third interface also includes a third control, and the method further includes: receiving a fourth operation input by the user on the third control; in response to the fourth operation, generating a video file containing a target atlas according to the target recommended template; and playing and / or storing the video file.

[0026] The terminal device can respond to user operations and generate short videos from the images in the target image set based on the target recommendation template, thereby achieving personalized display, enriching the image display methods, and improving the user experience.

[0027] In some possible implementations, before searching the image library based on search keywords to obtain the target image set matching the search keywords, the process also includes: analyzing the images in the image library to obtain the image elements of each image, where the image elements include at least the generation time.

[0028] In some possible implementations, the image element also includes: the location where the image was generated, which can be one or more of a city, a POI region, and a custom location.

[0029] In some possible implementations, the location where the image was generated is obtained by converting the longitude and latitude of the location where the image was generated.

[0030] In some possible implementations, image elements also include: feature vectors of people and / or semantic features in the image.

[0031] The methods for image analysis and clustering are described above and will not be repeated here. The image elements mentioned above describe the image's attributes, facilitating classification.

[0032] In a second aspect, an atlas acquisition device is provided, comprising a unit consisting of software and / or hardware, the unit being used to perform any one of the methods described in the first aspect.

[0033] Thirdly, embodiments of this application provide a chip including a processor; the processor is used to read and execute a computer program stored in a memory to perform any one of the methods described in the first aspect.

[0034] Optionally, the chip further includes a memory, which is connected to the processor via a circuit or wire.

[0035] Alternatively, the chip may further include a communication interface.

[0036] Fourthly, an electronic device is provided, comprising: a processor, a memory, and an interface; the processor, memory, and interface cooperate with each other to enable the electronic device to perform any one of the methods described in the first aspect.

[0037] Fifthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, the processor performs any one of the methods described in the first aspect.

[0038] In a sixth aspect, a computer program product is provided, the computer program product comprising: computer program code, which, when executed on an electronic device, causes the electronic device to perform any one of the methods described in the first aspect. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the structure of a terminal device 100 provided in an embodiment of this application;

[0040] Figure 2 This is a software structure block diagram of the terminal device 100 provided in the embodiments of this application;

[0041] Figure 3 This is an interactive diagram of an example of an image set acquisition method provided in an embodiment of this application;

[0042] Figure 4 This is a schematic diagram of an example gallery interface provided in an embodiment of this application;

[0043] Figure 5 This is another example of a smart image processing interface diagram provided in the embodiments of this application;

[0044] Figure 6 This is another example of an interface diagram related to smart film processing provided in the embodiments of this application;

[0045] Figure 7 This is another example of an interface diagram related to smart film processing provided in the embodiments of this application;

[0046] Figure 8 This is a flowchart illustrating an example of an image set acquisition method provided in an embodiment of this application;

[0047] Figure 9 This is another example of a schematic diagram of a picture set acquisition device provided in the embodiments of this application. Detailed Implementation

[0048] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0049] Hereinafter, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include one or more of that feature.

[0050] The image acquisition method provided in this application can be applied to terminal devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of terminal device.

[0051] For example, Figure 1 This is a schematic diagram of the structure of a terminal device 100 provided in an embodiment of this application. The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0052] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0053] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.

[0054] The software system of terminal device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of terminal device 100.

[0055] Figure 2 This is a software structure block diagram of the terminal device 100 according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer may include a series of application packages.

[0056] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0057] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0058] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0059] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0060] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0061] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0062] The phone manager is used to provide communication functions for terminal device 100. For example, it manages call status (including connection, hang-up, etc.).

[0063] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0064] The notification manager allows applications to display notification information in the status bar. It can be used to convey informational messages and can disappear automatically after a short time without user interaction.

[0065] In this embodiment, the application framework layer also includes a media recommendation service, cloudification parameters, and a recommendation database.

[0066] The cloud-based parameters include a configuration file for a recommended template. This configuration file can be downloaded from the cloud or pre-installed on the terminal device and updated periodically from the cloud.

[0067] The media recommendation service (MrsgService) includes a media recommendation generation engine and a recommendation template parser. Both the engine and parser can adapt different graph atlases to different recommendation templates based on configuration files and generate corresponding recommendation statements.

[0068] A recommendation database is used to store adapted image sets and recommendation statements for easy presentation later.

[0069] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.

[0070] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0071] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0072] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0073] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0074] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0075] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0076] A 2D graphics engine is a graphics engine for 2D drawing.

[0077] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0078] For ease of understanding, the following embodiments of this application will be described using the following methods: Figure 1 and Figure 2 Taking the terminal device with the structure shown as an example, and in conjunction with the accompanying drawings and application scenarios, the atlas acquisition method provided in this application embodiment will be specifically described.

[0079] Typically, users can take photos using the camera application (hereinafter referred to as the camera app) installed on their terminal device. Photos taken can include pictures of people, landscapes, other objects, or different objects together. Photos taken by the camera app can be stored in the gallery application (hereinafter referred to as the gallery). Terminal devices can store not only photos taken by the camera app, but also images from other sources, such as screenshots, images downloaded from web pages, and images transmitted through other communication tools (such as chat applications).

[0080] When users want to browse images, they can open the gallery. For example, a user can open the gallery, click the "Camera" folder to view photos taken with the camera app, or click the "Screenshots" folder to view screenshots. Due to the limited screen size of terminal devices, the number of images displayed at once is limited. Users can also navigate to folders and scroll up and down to view other images within those folders. Traditional image display methods are relatively simple, making image browsing monotonous and resulting in a poor user experience.

[0081] The image atlas acquisition method provided in this application can cluster images in an image library according to one or more elements, and then select images from the image library whose image elements match the elements of the recommended template to form an image atlas. This allows the resulting image atlas to vary based on the image elements of different images. It should be noted that the image atlas obtained by this method based on the image elements is used to recommend images to users, thereby personalizing the image display according to the recommended template. This image display method is richer and more diverse, enhancing the user's browsing experience.

[0082] In this embodiment, when new images are stored in the image library, the terminal device can perform computer vision (CV) analysis on each image stored in the library in the background. CV analysis can include the following three aspects:

[0083] Aspect 1: Location information conversion.

[0084] Specifically, the terminal device can convert longitude and latitude location information into specific generated locations, such as cities and / or points of interest (POIs).

[0085] When generating an image, the terminal device can record the image's location information and then convert that information into a specific location. Taking a camera application taking a photo as an example, if the terminal device's positioning function (such as GPS) is enabled when taking the photo, it can obtain and record its current longitude and latitude as the location information of the photo. For instance, if the terminal device's longitude and latitude at the time of taking image P1 are (X+Y) / 2 E, (Z+W) / 2 N, and city C1's longitude range is XX to YY E, and its latitude range is ZZ to WW N, then the terminal device, through CV analysis, can determine that image P1 was generated in city C1, thus converting the location information of image P1 into the location of city C1. Optionally, the terminal device can also determine the POI area corresponding to the longitude and latitude of image P1 as the location of image P1 based on the longitude and latitude ranges of some preset POI areas. Optionally, the POI area can be a tourist attraction, an office building, a residential area, etc.

[0086] If the location function of the terminal device is not enabled when taking a photo, the terminal device cannot obtain the longitude and latitude at that time, so the location information of the captured image can be empty, and the generated location is also empty.

[0087] Optionally, if the terminal device's location function is not enabled when taking a picture, the user can manually input the image's location information. For example, the user can directly input their current longitude and latitude for the image. In this case, the terminal device can store the user-input longitude and latitude as the location information of the captured image. The user can also directly input the name of their current city, the name of a scenic spot, or a user-defined location, such as "Old Place." The terminal device can then store the user-input location name as the image's generated location. Optionally, the terminal device can also determine its current location based on the coverage of the currently connected local area network and use that area as the image's generated location.

[0088] After the terminal device performs location information conversion on multiple images, it can obtain the generation locations of these multiple images. The terminal device can also perform cluster analysis on the generation locations of multiple images, for example, by taking the union of the generation locations of multiple images to obtain multiple locations. Optionally, these generation locations can be cities and / or POI areas, which is not limited in this embodiment. Some or all of these multiple locations can be used in the configuration information of the recommendation template to form location elements.

[0089] Aspect 2: Extracting character feature vectors.

[0090] Specifically, the terminal device can use a facial recognition model to identify faces appearing in an image, thereby obtaining facial feature vectors corresponding to different faces. It should be noted that if an image contains multiple faces, a corresponding facial feature vector can be extracted from each face, resulting in multiple facial feature vectors. After the terminal device extracts multiple facial feature vectors from multiple images, it can also perform cluster analysis on these extracted facial feature vectors to obtain a person feature vector representing each individual. If the same person's face appears in multiple images, then the faces in these multiple images correspond to the same person feature vector (TagID). It should also be noted that the terminal device can set up a list of people, where each person's name corresponds to a TagID. That is, each TagID represents a natural person. Therefore, if an image contains multiple faces, multiple TagIDs corresponding one-to-one with each face can be extracted from that image.

[0091] Optionally, the terminal device can also employ a human body recognition model to identify the shape of human bodies appearing in an image, thereby obtaining human body feature vectors corresponding to different human bodies. It should be noted that if an image contains multiple human bodies, a corresponding human body feature vector can be extracted for each human body, resulting in multiple human body feature vectors. After the terminal device extracts multiple human body feature vectors from multiple images, it can also perform cluster analysis on the extracted multiple human body feature vectors to obtain a person feature vector representing each person. If the same person's body appears in multiple images, then the body of that person in all these images corresponds to the same person feature vector (ReID). It should also be noted that the terminal device can set up a list of people, where each person's name corresponds to a ReID. That is, each ReID represents a natural person. Therefore, if an image contains multiple human bodies, multiple ReIDs corresponding one-to-one with each human body can be extracted from the image.

[0092] Optionally, the terminal device can also use a combination of face recognition model and human body recognition model to recognize both faces and bodies in the image, as long as it can obtain the TagID, ReID, or combined ID that represents each person.

[0093] The following explanation uses TagID as an example. The names of the people included in the above list can be entered by the user. For example, after the terminal device identifies and clusters people in multiple images, it obtains two TagIDs: TagID-1 and TagID-2. For example, taking the union of facial and / or body feature vectors of multiple people yields multiple person feature vectors. The names of the people corresponding to these multiple feature vectors can be used as multiple elements in the person element. When a user browses the album, if the user selects to browse by portrait, the terminal device can display two folders. Folder 1 stores images of the person corresponding to TagID-1, and folder 2 stores images of the person corresponding to TagID-2. Folder 1 can use an image corresponding to TagID-1 as its cover, and folder 2 can use an image corresponding to TagID-2 as its cover. The user can name these two folders separately. For example, if the user names folder 1 (corresponding to TagID-1) as "Little Daughter," it means that the person corresponding to TagID-1 is "Little Daughter." If a user names the folder 2 corresponding to TagID-2 as "lover", it means that the person corresponding to TagID-2 is "lover".

[0094] Users can also tag individual people in an image. For example, tagging an image with the tag "young daughter" means that the image's TagID corresponds to the young daughter. If other images have the same TagID, it means that those other images also correspond to the young daughter.

[0095] Users can use any of the above methods to label different people in the album one by one. Some or all of the names of these people can be used in the configuration information of the recommendation template to form the person element.

[0096] Additionally, user-defined names, such as "youngest daughter" and "lover," can be added to the character list. In other words, some or all of the names of characters in the character list can serve as character elements that make up the character element.

[0097] If the user does not enter a name for a person for a certain TagID, the terminal device can set the name of the person corresponding to that TagID to empty, for example, displaying: Unnamed (or Add Name).

[0098] It should be noted that a person's face and body can correspond to one or more TagIDs. If an image contains multiple people, it can correspond to multiple TagIDs.

[0099] Optionally, the same face or body may have both TagID-1 and ReID-1, or the same face may be identified with different TagIDs due to recognition errors. Users can then operate their terminal devices to change the names of the people corresponding to TagID-1 and ReID-1 to the same name, thus associating the person's feature vectors and defining TagID-1 and ReID-1 as the same person. During image search, if searching for images of this person, images with the person's feature vector set to either TagID-1 or ReID-1 can be filtered out.

[0100] When new images are added to the photo library, the terminal device can also perform facial and / or human body recognition on the new images to obtain the corresponding TagID. Optionally, if the new image contains a person already in the album, the image can be associated with an existing TagID. Optionally, if the new image contains a person not present in the album, a new TagID can also be obtained to correspond to the newly added person.

[0101] Optionally, the aforementioned face recognition model and body recognition model can be trained neural network models.

[0102] Aspect 3: Extracting semantic feature vectors.

[0103] Specifically, the terminal device can employ a semantic recognition model to perform semantic recognition on the images, thereby obtaining a semantic feature vector (ClipID) corresponding to each image. A semantic feature vector represents the image content contained in the corresponding image. Optionally, the image content can include objects different from people, such as pets, cats, dogs, food, etc.; it can also include scenes, such as painting, traveling, sports, dancing, weddings, parties, etc. This application embodiment does not limit the image content that can be recognized. The objects and scenes contained in the image content of an image can be one or more. For example, if the image is a scene of many people having a picnic in the wild, and there is also a pet dog in the picture, then the recognized semantic feature vector can represent that the image includes image content such as: traveling, parties, pets, and pet dogs.

[0104] Optionally, the semantic recognition model described above can be a trained neural network model.

[0105] The aforementioned CV analysis process can be performed when the terminal device is in an idle period. This idle period can be any time when the user does not use the terminal device, such as when the device is charging and the screen is off, or between 2:00 AM and 4:00 AM, or when the screen is off. This embodiment of the application does not limit this. Performing CV analysis on the terminal device during idle periods does not affect user operation and is therefore highly reasonable.

[0106] The objects of the aforementioned CV analysis can include not only individual images stored in the image library, but also frame images from video files within the library. CV analysis of video files can be performed by analyzing each frame image individually, or by extracting frames from the video to obtain frame images, and then performing the aforementioned CV analysis on these frame images to obtain the CV analysis results. Optionally, during frame extraction, the terminal device can choose to extract one frame at fixed intervals, such as one frame every ten frames, or it can extract one frame at fixed time intervals, such as one frame per second. This application embodiment does not limit the specific method of frame extraction. Including video frame images as objects of CV analysis expands the selection range of the image set, making the selected image set more comprehensive and the subsequent presentation richer, thus improving the user experience.

[0107] As mentioned above, the result of CV analysis is the identification of image elements, which can include one or more of the following: location, person feature vectors, and semantic feature vectors. It should be noted that when images in the image library are generated, the terminal device also records the image's generation time. Specifically, the image generation time can include the year, month, and day, such as February 15, 2023; it can also include the hour and minute, such as 14:00. The generation time of each image, along with one or more of its location, person feature vectors, and semantic feature vectors, can be associated with that image and stored together in the recommendation database for later use.

[0108] In addition, the terminal device can also parse the recommended template according to the configuration file of the recommended template during idle periods or other triggered events. Optionally, the terminal device can pre-install the configuration file of the recommended template locally, or it can download the configuration file of the recommended template from the cloud; this embodiment of the application does not limit this. Optionally, the terminal device can also periodically query whether there is a new configuration file in the cloud. When a new configuration file exists in the cloud, it can download the new configuration file to replace the original configuration file, thereby ensuring that the configuration file is updated in a timely manner.

[0109] Optionally, when downloading a new configuration file, the terminal device can also determine whether it currently supports the execution of the new configuration file. If it does, the terminal device can download the new configuration file to replace the original configuration file. If the terminal device does not currently support the execution of the new configuration file, for example, if the system version installed on the terminal device does not match the configuration file, the terminal device can use the original configuration file for template analysis.

[0110] The configuration information in the above configuration file can be called cloudification parameters, which can be used to describe the recommendation templates. Optionally, the configuration file can include multiple topics. Each topic includes multiple different recommendation templates. The configuration information of each recommendation template includes four components. Specifically, the four components of a recommendation template are: time, location, human, and event. To clearly illustrate the meaning of the configuration information of the recommendation template, an example of a recommendation template is used here:

[0111] Component 1: Time element.

[0112] The configuration information for recommended templates can include time elements. Specifically, time elements can include one or more time descriptive words, each descriptive word being one of the elements of a time element, called a time element. Time descriptive words do not need to use dates; instead, they use colloquial descriptive words, ensuring that the descriptions in the later presentation conform to users' expression habits and improving user experience. For example, time descriptive words (time elements) can include, but are not limited to, one or more of the following: today, yesterday, last weekend, this week, last week, this month, this year, last year, the year before last, New Year's Eve, Spring Festival, Mid-Autumn Festival, Father's Day, Mother's Day, Children's Day, Valentine's Day, Qixi Festival, Christmas, Thanksgiving, wedding anniversary, graduation anniversary, and birthday. The number and types of time elements in the configuration information of different recommended templates can be the same or different.

[0113] Component 2: Location.

[0114] In this embodiment, the location element in the configuration information can be empty. In the actual template parsing process, since locations cannot be exhaustively listed, the location element can be obtained by clustering the generation locations of multiple images obtained from CV analysis. Therefore, compared to listing a large number of existing locations to cover user needs, it is unnecessary to search for invalid locations unrelated to the images in the image library; only the generation locations of the images obtained from image clustering need to be searched, which can effectively narrow the search scope. For example, in the above CV analysis process, if the results of CV analysis of multiple images are clustered, and some images have location C1, some images have location B2, and other images have empty locations, then if the recommended template includes the location element during template parsing, the terminal device can search only by city C1 and location B2, without needing to search for images by searching other locations. The specific template parsing process can be found in the description below, which will not be repeated here.

[0115] Component 3: Character elements.

[0116] The recommended template configuration information also includes a "person element." Specifically, a person element can include descriptive terms for one or more people, each descriptive term being an element of the person element, referred to as a person element. Descriptive terms for people can use colloquial terms, nicknames, names, and descriptions of relationships, including but not limited to: spouse, eldest daughter, youngest son, boyfriend, Zhang San, grandfather, grandmother, maternal grandfather, maternal grandmother, etc. Optionally, the person element in the recommended template configuration information can be empty. In the actual template parsing process, it can be obtained by exhaustively listing the names of people or relationships, or by clustering images. Using image clustering analysis to obtain the elements of the person element is more effective in narrowing the search scope compared to listing a large number of people or relationships. For example, in the above CV analysis process, after clustering the results of CV analysis of multiple images, if all the images contain only the youngest daughter and spouse, then in the template parsing process, only these two people need to be searched, without searching for other people.

[0117] Component 4: Event Element.

[0118] The recommended template configuration information also includes event elements. Specifically, event elements can include one or more event descriptors, each descriptor being an element of the event element, called an event element. The descriptors for event element types can also use colloquial terms, ensuring the descriptions conform to user expression habits and improving user experience. Event element descriptors include, but are not limited to: travel, daily exercise, hobbies, daily gatherings, group photos, food, food preparation process, cats, dogs, pets, travel experiences, travel, happy parties, team building meals, playing music, drawing, singing, dancing, etc., one or more of these. Optionally, since the types of events cannot be exhaustively listed, the event fields in the configuration information can be treated as open fields, passed in by developers during the development phase. In the actual template parsing process, the terminal device can search for images based on the event fields passed in by the developers, a process known as value-based search matching.

[0119] The configuration information for the aforementioned recommendation template also includes a recommendation statement template. This template sets one or more of several element slots. These slots can include: a time slot for time elements, a location slot for location elements, a person slot for person elements, and an event slot for event elements. Once the elements of these four components are filled into their corresponding slots, a complete recommendation statement is generated. It should be noted that the recommendation statement templates differ across different recommendation templates. Even using the same template, different elements will result in different generated recommendation statements. Furthermore, the number, type, and order of element slots vary between different recommendation statement templates.

[0120] Optionally, elements within the same component can also be multiple similar synonyms. For example, elements of an event element can include: travel, play, and outing. When generating recommendation statements, travel, play, and outing can be interchanged or alternated, making the recommendation statements richer and improving the user experience.

[0121] Optionally, multiple element slots may also include modifier slots for filling in modifiers. Based on this, modifier slots in the same recommendation statement template, besides element slots, can also use synonyms or related words interchangeably to conform to everyday language expression habits and ensure differentiated descriptions of the recommendation statements. For example, the recommendation statement template is "Generate %_time% videos of outings". Here, "generate" is the modifier, and the corresponding slot is the modifier slot. Taking "time" as today as an example, the recommendation statement generated by the terminal device would be any one of: Generate videos of today's outings, Create videos of today's playtime, or Generate videos of today's picnic. The position of the modifier slots and the corresponding modifiers in the recommendation statements differ in different recommendation statement templates.

[0122] Optionally, the configuration file may include configuration information for multiple recommendation templates, such as the types of the four components of each recommendation template, the types and number of elements for each component, and whether the recommendation statement template is partially or completely different.

[0123] Based on the meaning of the above configuration information, the following section will detail the template parsing process and the process of matching the atlas based on the results of template parsing, using the element slots in the recommendation statement template as an example.

[0124] This section uses the example of a terminal device searching for images one by one based on the four components of a recommended template's configuration information to obtain an image set that matches the four components. The process includes the following four steps:

[0125] Step 1: Generate an initial keyword combination based on the four components of the configuration information.

[0126] For specific examples of configuration information, please refer to the following:

[0127]

[0128]

[0129] In this configuration, "topic" represents the name of the topic, and each topic can correspond to one or more recommended templates. "_time" represents the time element; "_location" represents the location element; "_human" represents the human element; and "_event" represents the event element. It should be noted that one or more of these components may be empty, but not all of them will be empty, and among the non-empty components, one or more elements can be present. In the example configuration information above, the topic "topic" represents daily Vlogs, "_time" includes the five elements: today, yesterday, last weekend, this week, last week, and this month, "_human" is empty (no element, represented by []), and the element "_event" is travel. In other words, each component can be understood as a defined range. For example, if the time element includes the elements "today, yesterday, and last weekend," then the range of this time element is: today, yesterday, and last weekend.

[0130] Taking the time element as an example, the time element and its corresponding element can be understood as a key-value pair relationship. For example, if the time element includes the elements: today, yesterday, and last weekend, then the resulting key-value pairs can include: time-today, time-yesterday, and time-last weekend.

[0131] Optionally, in the configuration information, if a component is empty, it can be represented by "[]", that is, the recommendation template corresponding to the configuration information does not include the element of the component, and the recommendation statement corresponding to the recommendation template does not have the element slot corresponding to the component (e.g., "_location":[] and "_human":[]" in the example above); if the element of a component is written by passing values ​​from outside, such as the result of image clustering or the result obtained by learning user profiles, it can be represented by "null" (e.g., "title":null" in the example above); if the element of a component uses the descriptive words in the configuration information, it can be represented by "%***%" (where *** represents the category of the component, such as ["%_time%"] and ["%_event%"] in the example above).

[0132] Optionally, the configuration information can also include other components, such as "_argument," which can be used as an extended component. This component can act as an extended variable, enriching the recommendation statement template. The specific types of extended variables can be any one or more of the four types of components mentioned above. Optionally, the number of extended variables can also be one or more. For example: "_argument": [certain little daughter], "sentence": "Generate videos of %_human% and %_argument%%_time% going on an outing". Another example: "_argument": [whole little daughter], where "whole" indicates that the people are the intersection, meaning that the search should retrieve images of the person "human" and his little daughter. Correspondingly, the recommendation statement template includes the corresponding element slots, and the generated recommendation statement contains elements of the extended component.

[0133] Optionally, the configuration information may also include a field for the search strategy, such as "how_", which characterizes the search strategy for the images. If "how_" is "TIME_AXIS", it means that the image set corresponding to the recommended template can be searched from far to near based on the timeline during the search process. This application embodiment does not limit the specific form of the search strategy.

[0134] Optionally, in the configuration information, "sentence" is the recommended statement template. The space between the two % symbols is filled with the element corresponding to the feature slot. "priority" indicates the recommendation priority of the corresponding statement; the smaller the number, the higher the priority.

[0135] Optionally, the same recommendation template may correspond to different recommendation statement templates. For example, a recommendation template might have two different phrases: "sentence": "Generate videos of %_human% and %_argument%%_time% traveling" and "sentence": "Generate videos of %_human% and %_argument% traveling". Based on this, the terminal device can generate a set of recommendation statements for the same recommendation template. This set of recommendation statements can be one or more, with the number of recommendation statements corresponding to the number of recommendation statement templates. When generating recommendation statement templates, the terminal device can prioritize recommendation statements based on their "priority," recommending those with the highest priority first.

[0136] Optionally, the configuration information may also include a field limiting the number of recommended statements, such as "topk". When "topk":1, it means one set of recommended statements is generated. When "topk":3, it means three sets of recommended statements are generated.

[0137] Optionally, the configuration information may also include a scheduling level field, indicating the priority level for parsing the corresponding recommendation template. For example, "level". When "level": 0, it means that the priority level for parsing the corresponding recommendation template is based on other recommendation timings, and will not be triggered based on the background polling strategy of the terminal device, i.e., it will not be triggered during idle periods. When "level" is not 0, it means that the corresponding recommendation template is parsed based on the background polling strategy, such as being triggered during idle periods.

[0138] Optionally, the configuration information may also include a field for the scan interval, such as "scan_interval_hour", which indicates the minimum interval for parsing the recommended template, and can be in hours. For example, "scan_interval_hour":4 means that the recommended template will not repeatedly search the image library within four hours to update the image atlas corresponding to the recommended template.

[0139] Optionally, the configuration information may also include a name field, such as "title", which represents the name of the recommended card when the template is displayed on the desktop. When "title":null, it means that the card name corresponding to the recommended template can be entered as a value.

[0140] The following are the configuration details for several recommended templates for everyday Vlog themes:

[0141]

[0142]

[0143]

[0144]

[0145] The above configuration information is just an example. It should be noted that the configuration file can also include more configuration information for recommended templates, which will not be elaborated here.

[0146] Optionally, the elements in the time element above can also be specific dates. When the elements in the time element are specific dates, no conversion is required.

[0147] During template parsing, elements from the four components mentioned above can be used as initial keywords. Then, the image library is queried according to different combinations of these initial keywords to obtain the corresponding image sets. Specifically, the initial keyword combinations can be obtained by combining each element in "_time" with each element in "_location", each element in "_human", and each element in "_event". For ease of description, taking the elements in "_time" (today and yesterday), "_location" (city C1 and attraction B2), "_human" (youngest daughter and lover), and "_event" (painting and dancing) as examples, the generated initial keyword combinations include the following 16 types:

[0148] Today + City C1 + Little Daughter + Dancing;

[0149] Today + City C1 + Little Daughter + Painting;

[0150] Today + City C1 + Lover + Dancing;

[0151] Today + City C1 + Lover + Painting;

[0152] Today + Attraction B1 + Little Daughter + Dancing;

[0153] Today + Scenic Spot B1 + Little Daughter + Painting;

[0154] Today + Attraction B1 + Lover + Dancing;

[0155] Today + Attraction B1 + Lover + Painting;

[0156] Yesterday + City C1 + Little Daughter + Dancing;

[0157] Yesterday + City C1 + Little Daughter + Drawing;

[0158] Yesterday + City C1 + Lover + Dancing;

[0159] Yesterday + City C1 + Lover + Painting;

[0160] Yesterday + B1 of the attraction + my youngest daughter + dancing;

[0161] Yesterday + B1 of the tourist attraction + my youngest daughter + drawing;

[0162] Yesterday + B1 of the attraction + my lover + dancing;

[0163] Yesterday + Scenic Spot B1 + My Lover + Painting.

[0164] When any of the four components contains an empty element, it indicates that the recommendation template does not consider empty elements, and therefore, the generated initial keyword combinations will not contain initial keywords corresponding to empty elements. For example, if "_time" includes elements such as "today" and "yesterday," "_location" includes elements such as "city C1," "_human" contains an empty element, and "_event" contains an empty element, then the generated initial keyword combinations will include the following two types:

[0165] Today + City C1;

[0166] Yesterday + City C1.

[0167] Optionally, the number of elements in the above four components can also be other numbers, and the specific composition and number of the resulting initial keyword combinations will vary accordingly, making them impossible to exhaustively list. As long as the same initial keyword combination does not contain two or more elements of the same type of component, and the same initial keyword combination contains one element of a component whose elements are not empty, it is acceptable; further details will not be elaborated here.

[0168] Step 2: Keyword conversion.

[0169] After obtaining the initial keyword combination, the terminal device converts the time-related and person-related keywords within the initial keyword combination. For example, consider the initial keyword combination: Today + City C1 + Little Daughter + Dancing.

[0170] Specifically, the terminal device converts the initial keyword "today" into a specific date. If today's date is February 14, 2024, then the initial keyword "today" is converted to the time period from 0:00:00 to 23:59:59 on February 14, 2024, as the time search keyword. Optionally, if the time keyword in the initial keyword combination is "yesterday," and today is February 14, 2024, then the initial keyword "yesterday" is converted to the time period from 0:00:00 to 23:59:59 on February 13, 2024, as the time search keyword.

[0171] Optionally, when converting time elements in the time element, the time element is described as start time and end time, and optionally, a validity period may also be included.

[0172] For example, the initial keyword "today" is converted to a start time of 0:00:00 on February 14, 2024, and an end time of 23:59:59 on February 14, 2024. Optionally, the validity period corresponding to the initial keyword "today" is 23:59:59 on February 14, 2024.

[0173] The terminal device can also convert the initial keyword "youngest daughter" into the TagID corresponding to the person "youngest daughter". For example, in the photo album, if the TagID corresponding to the youngest daughter's face or body is TagID-1, then the initial keyword "youngest daughter" can be converted into the corresponding TagID-1.

[0174] If the location keywords are cities, attractions, or other locations whose corresponding longitude and latitude can be obtained, then the location keywords do not need to be converted. If the location keywords are manually set addresses, then the location keywords can be converted to the corresponding map address, and the location keywords corresponding to this map address can be obtained. For example, the location keyword is: "My Company". The terminal device, after learning user habits, can obtain the user's workplace as "XX Office Building". The terminal device can then convert "My Company" to the map address "XX Office Building", and then use the keyword "XX Office Building" corresponding to the map address "XX Office Building" as the converted location keyword. This "XX Office Building" is an already identified POI area.

[0175] After keyword conversion, the initial keyword combination described above can be replaced with the converted keywords by the terminal device, thus transforming the initial keyword combination into a corresponding search keyword combination. For example:

[0176] The initial keyword combination: Today + My Company can be converted into the search keyword combination: 0:00 to 23:59 on February 13, 2024 + XX Office Building.

[0177] It should be noted that event keywords do not need to be converted.

[0178] When the initial keywords do not need to be converted, such as when the location keywords are cities or attractions, or when the initial keywords are event keywords, the initial keywords can be used as search keywords to form a combination of search keywords to search the image library.

[0179] Step 3: Search the image library based on the converted search keyword combinations to obtain the image sets corresponding to the search keyword combinations.

[0180] Specifically, terminal devices can search for and match images with search keywords in the image library, which will then be used as the image set corresponding to that search keyword combination.

[0181] Let's continue with an example using an initial keyword combination: Today + City C1 + Little Daughter + Dancing. After keyword conversion, the search keyword combination can be obtained as: 0:00 to 23:59 on February 14, 2024 + City C1 + TagID-2 + Dancing.

[0182] The terminal device first searches the image library to obtain all images generated between 0:00 and 23:59 on February 14, 2024, as the first image set. Then, it searches the first image set for all images located in city C1, as the second image set. Next, it searches the second image set for all images with TagID corresponding to TagID-2, as the third image set. Finally, it searches the third image set for all images whose semantic feature vectors match the drawing event, resulting in the fourth image set. Based on this, the terminal device selects the image set corresponding to: Today + City C1 + Little Daughter + Dancing. The images in this set are all photos of the little daughter dancing in city C1 taken between 0:00 and 23:59 on February 14, 2024. Optionally, the order of the four components searched during the process of obtaining the fourth image set can be adjusted. For example, the terminal device first searches for all images whose semantic feature vectors match the event of drawing. Then, it searches for images whose creation time was between 0:00 and 23:59 on February 14, 2024. Next, it searches for images featuring a young daughter. Finally, it searches for images with the location city C1. The order in which the terminal device searches for the four elements is not limited in this embodiment; the only requirement is to select image sets that match all four elements.

[0183] Optionally, when searching the image library, the terminal device can also determine whether an image matches the search keywords in the aforementioned search keyword combination. If so, the image is added to the image set; otherwise, it is not added. The terminal device checks each image individually to determine whether it matches the search keywords in the search keyword combination, thereby generating the required image set.

[0184] It should be noted that the specific process of searching for images that match semantic feature vectors and events can be as follows: The terminal device can traverse the semantic feature vectors of all images to be searched according to the semantic feature vectors describing event elements, such as the semantic feature vector of drawing. If the similarity between the semantic feature vector ClipID-1 corresponding to image 1 and the semantic feature vector of drawing is higher than or equal to a preset similarity threshold, for example, higher than 90%, it means that the image content of image 1 likely contains the event of drawing, and the terminal device determines that the semantic feature vectors of image 1 and drawing match. If the similarity between the semantic feature vector ClipID-2 corresponding to image 2 and the semantic feature vector of drawing is lower than the preset similarity threshold, it means that the image content of image 2 likely does not contain the event of drawing, and the semantic feature vectors of image 2 and drawing do not match. Optionally, the preset similarity threshold can also be other values ​​such as 85% or 80%, and this embodiment does not limit this.

[0185] It should be noted that one or more of the four components of a recommendation template may be empty, such as the location component being empty. When a component is empty, there is no need to search for that component. Additionally, a component of a recommendation template may have multiple elements; for example, the "person" component may include "lover" and "young daughter." When a component has multiple elements, all elements of that component need to be searched. Optionally, AND / OR relationships between multiple elements can be distinguished using symbols. Taking the "person" component as an example, if two elements are written within the same square bracket in the configuration file, such as "human":[["I", "lover"]], it represents "I" and "lover". Therefore, the retrieved images should include both "I" and "lover". If the elements are written within two separate square brackets in the configuration file, such as "human":[["I"], ["lover"]], it represents "I" or "lover". Therefore, the retrieved images may include "I", or any one or both of "lover". Alternatively, other components, such as time, location, and event elements, can also be distinguished by symbols to indicate the AND or OR relationship between multiple elements, which will not be elaborated here.

[0186] Optionally, after the terminal device finds the image set corresponding to the recommended template, it can determine whether the number of images in the image set meets a quantity threshold. The quantity threshold can be a natural number, such as six. If the number of images in the image set is less than the quantity threshold, for example, less than six, it means that the number of images in the searched image set is too small, and the effect of generating a video is not rich enough. Therefore, the image set will no longer be recommended to the user, and the corresponding recommendation statement will not be displayed. The terminal device can then execute a search for the next combination of search keywords. If the number of images in the image set is greater than or equal to the quantity threshold, for example, greater than or equal to six, then the image set can be recommended to the user, and the corresponding recommendation statement can be displayed.

[0187] Optionally, if the image set searched by the terminal device includes multiple frames from a video, these frames can be used as images in the image set to generate a short video. Alternatively, the video segment containing these frames can be extracted, and all frames from this extracted segment can be combined with other images from the image set to generate a short video. It should be noted that the playback order of all frames from the extracted video segment remains unchanged from the original video's playback order in the generated short video to ensure optimal playback quality.

[0188] After keyword conversion, the terminal device can generate a recommendation statement. The specific process for generating the recommendation statement is as follows: the terminal device can fill the initial keywords from the initial keyword combination into the corresponding recommendation statement template, thereby generating the corresponding recommendation statement. This process can be called recommendation statement instantiation. In the configuration information, the recommendation statement template "sentence" is: "Generate a video of %_time%%%_human% at %_location%%%_event%". This recommendation statement template includes a time slot: time, a human slot: human, a location slot: location, and an event slot: event. Each element slot to be filled is distinguished by the symbol %; two consecutive % represent one slot. For example, if the terminal device fills the time "today" into the time slot of the recommendation statement template, the human slot "little daughter" into the human slot, the location slot "attraction B" into the location slot, and the event slot "dancing" into the event slot, the recommendation statement generated is: Generate a video of the little daughter dancing at attraction B today. Optionally, the initial keyword in the above recommended statement template is: %_time%%_human%%_location%%_event%.

[0189] For example, if the recommended statement template "sentence" is: "Generate a video of %_time%%%_event%", this template includes a time slot: time and an event slot: event. Optionally, if the event slot in this template is already filled with "travel", then the recommended statement template "sentence" would be: "Generate a video of %_time% travel". During template parsing, the terminal device can search for the keywords %_time% and travel to obtain the corresponding image set. Simultaneously, it fills the time slot with time information, thus generating the corresponding recommended statement: "Generate a video of today's travel". The initial keyword in this recommended statement template is: %_time%%%_event%.

[0190] Alternatively, the combination of the above search keywords can be implemented using the following five-level for loop:

[0191]

[0192]

[0193] Based on this, the terminal device can obtain one or more sets of search keyword combinations corresponding to a recommendation template, as well as the image set corresponding to each set of search keyword combinations and the recommendation statement corresponding to each set of search keywords, thus obtaining the template parsing result. The terminal device then stores the template parsing result in the recommendation database. The image set is stored by storing the identifiers (IDs) of the images within the image set.

[0194] When there are configuration information for multiple recommendation templates in the configuration file, the above steps one to three can be executed for the configuration information of each recommendation template to complete the parsing of all recommendation templates and form one or more recommendation statements for each recommendation template and the graph set corresponding to each recommendation statement.

[0195] The method described above, which generates recommendation statements by filling element slots with templates containing element slots, is more flexible than pre-set fixed recommendation statements. The generated recommendation statements conform to users' conversational expression habits, resulting in a better reading experience. Furthermore, the elements within the location and people elements can be clustered from images in the gallery, allowing the generated recommendation statements to vary from user to user, achieving personalized recommendation statements and enhancing the user experience.

[0196] To clearly describe the template parsing process of this application embodiment, combined with Figure 3 The interaction diagram shown is described in detail. For example... Figure 3 As shown, it includes:

[0197] S301. The user uses the camera application to take pictures or videos.

[0198] S302. The camera application responds to the user's shooting action, generates images or videos, and records the generation time and location information of the images or videos.

[0199] S303. The camera application sends images and / or videos, along with the time and location information of their creation, to the gallery for storage.

[0200] S304. The image library performs CV analysis on frame images of pictures and videos during idle periods to obtain the image elements of the pictures.

[0201] The image library can also cluster the locations where images were generated to obtain location elements.

[0202] S305. After analyzing all the images and frames, the gallery sends a start command to the media recommendation generation engine.

[0203] Specifically, the startup command is used to launch the media recommendation generation engine.

[0204] S306. In response to the startup command, the media recommendation generation engine sends an initialization command to the recommendation template parser.

[0205] S307. In response to the initialization directive, it is recommended that the template parser be initialized.

[0206] Specifically, the process of initializing the recommended template parser involves obtaining a configuration file. Optionally, the configuration file can be a JSON file. The configuration file includes the configuration information for the recommended templates. For a detailed description of the configuration information, please refer to the preceding text; it will not be repeated here.

[0207] S308. After the template parser initialization is complete, it returns an initialization completion message and a configuration file (or configuration information in the configuration file) to the media recommendation generation engine.

[0208] Optionally, if the recommendation template parser has not completed initialization, it is not necessary to send a message returning initialization complete, or to return a message returning initialization incomplete to the media recommendation generation engine.

[0209] S309. The media recommendation generation engine responds to the initialization completion message and executes the template parsing process according to the configuration file (or configuration information).

[0210] The template parsing process includes: repeatedly executing steps S310 to S317 for the configuration information of each recommended template.

[0211] S310. The media recommendation generation engine splits the recommendation statement template in the configuration information into at least one initial keyword combination according to the four element slots and the elements of at least one of the four component elements.

[0212] It should be noted that the split recommendation statement template is based on the element slots corresponding to the four constituent elements. Each initial keyword combination resulting from the split includes at least one initial keyword. Taking one set of initial keyword combinations as an example, each initial keyword in a set of initial keyword combinations corresponds to an element in one of the four constituent elements.

[0213] The terminal device executes steps S311 to S317 for each initial keyword group. Taking an initial keyword group as an example, the initial keyword group includes at least one initial keyword.

[0214] S311. The media recommendation generation engine performs keyword transformation on at least one initial keyword to obtain at least one search keyword.

[0215] It should be noted that initial keywords that do not require conversion can be used as corresponding search keywords.

[0216] S312. The media recommendation generation engine sends a search command to the image library, the search command carrying at least one search keyword.

[0217] S313. The image library searches for images in the image library using at least one search keyword to obtain the corresponding image set.

[0218] The image elements in the image set match at least one search keyword. For example, the image in the image set was generated at a time within the time period indicated by the time search keyword in at least one search keyword; or the image in the image set was generated at a location indicated by the location search keyword in at least one search keyword.

[0219] S314. The image library uses the identifiers of images in the image collection as search results and returns them to the media recommendation generation engine.

[0220] S315. The media recommendation generation engine determines whether the number of images in the image set is greater than six based on the search results.

[0221] If the number of images in the image set is greater than or equal to six, then execute S316; if the number of images in the image set is less than six, then continue to execute the template parsing process for the next recommended template.

[0222] S316. The media recommendation generation engine instantiates recommendation statements.

[0223] Specifically, the media recommendation generation engine fills at least one of the above initial keywords into the corresponding element slots of the recommendation statement template, thus completing the assembly of the recommendation statement and realizing the recommendation statement instance.

[0224] S317. The media recommendation generation engine sends the recommendation statement generated by at least one initial keyword and the search-obtained graph to the recommendation database for storage.

[0225] Optionally, the image set sent by the media recommendation generation engine can be a set of identifiers for images in the image set; alternatively, the media recommendation generation engine can also send the identifiers for images in the image set used as covers; alternatively, the media recommendation generation engine can also send assembled recommendation statements.

[0226] The set of image identifiers in the above image collection, the identifier of the cover image, and the recommendation statements can all be persistently stored in the recommendation database.

[0227] After all the recommendation templates have been parsed, the media recommendation generation engine can execute S318 and subsequent processes.

[0228] S318. Media Recommendation Generation Engine End Process.

[0229] Optionally, if an anomaly occurs during the CV analysis process, such as a sudden change in the charging / screen-off state, indicating that the user may need to use the terminal device, the image library can send a termination command to the media recommendation generation engine. The media recommendation generation engine responds to the termination command by ending the template parsing process, thereby releasing resources to ensure the user can use the terminal device normally.

[0230] Optionally, after completing the CV analysis, the image library can also send a termination command to the media recommendation generation engine, instructing the media recommendation generation engine to release resources.

[0231] The above Figure 3 The implementation principles and technical effects of each step in the interactive diagram shown can be found in the previous text, and will not be repeated here.

[0232] The previous section detailed the image set corresponding to the recommendation template and the process of obtaining the recommendation statement. Next, we will introduce the recommendation process of the recommendation statement in conjunction with user operations.

[0233] When a user opens the gallery, the terminal device displays the gallery's homepage interface, for example... Figure 4 As shown in Figure a. In Figure 4 The gallery homepage interface, shown in image a, includes a "Smart Photo" card. The homepage also includes other cards such as "One-Click Masterpiece," "Editing," and portrait thumbnails. At the bottom of the homepage interface are icons for several tabs, including: Photos, Albums, Moments, and Creations. Users can access different gallery interfaces by clicking on different tab icons.

[0234] When a user clicks the "Smart Photo Generation" card on the gallery homepage, the Smart Photo Generation function is activated. The terminal device can display images such as... Figure 4 The first-level page of the "Smart Projection" section, shown in image b, displays multiple recommendation quotes. Figure 4 Figure b in the example shows three recommendation statements. Each recommendation statement entry also displays a thumbnail of the corresponding cover image. Optionally, if the cover image is not defined in the image set corresponding to the recommendation statement, the recommendation statement entry may not display any image thumbnail; only the text of the corresponding recommendation statement needs to be displayed, for example... Figure 5 As shown in Figure a. Alternatively, as... Figure 4 Figure b in the middle and Figure 5 As shown in Figure a, on the first-level page of Smart Integration, a prompt can also be displayed above the recommended quotes.

[0235] by Figure 4Taking image b as an example, the first-level page of the "Smart Integrated Display" also includes theme switching controls, such as "More Skills," and recommendation statement switching controls, such as "Change." When a user clicks the recommendation statement switching control, the terminal device can switch the recommendation statement. For example, when a user clicks the recommendation statement switching control, the terminal device can switch the recommendation statement. Figure 4 Clicking the "Change" control in image b will display the following on the terminal device: Figure 6 The recommended statements shown.

[0236] Users can click on a recommendation statement on the Smart Film page to enter a secondary page of Smart Film, for example... Figure 7 As shown. The secondary page of the smart image gallery can be displayed in a dialog format, including recommended statements and the response statements generated by the terminal device in response to the recommended statements, as well as cards for the image galleries corresponding to the recommended statements. Optionally, the image gallerie cards may include thumbnails of multiple images in the gallery; they may also include controls for displaying all images, such as "View All"; and they may include video generation controls, such as "Generate Video". When the user clicks the control for displaying all images, the terminal device displays all or part of the images in the gallery. Optionally, the control for displaying all images may also display the number of images in the gallery. When the user clicks the video generation control, the terminal device generates a short video from the images in the gallery according to the recommended template corresponding to the recommended statement. The text, display effects, watermark, and background music in the short video are defined by the corresponding recommended template.

[0237] When a user clicks the theme switching control on the main page of Smart Image Generation, the terminal device displays the theme selection interface, for example... Figure 5 As shown in Figure b, the theme selection interface includes cards for multiple themes, such as daily vlogs, personal photos, and family photos. Clicking on any theme card displays a recommended quote for that theme. For example, clicking... Figure 5 When the card for the personal portrait theme in image b is displayed, the terminal device will show multiple recommended statements corresponding to the personal portrait theme, such as... Figure 7 As shown. Users can... Figure 7 On the secondary page shown, users can select the desired recommendation phrase to generate a corresponding short video. The terminal device can play, store, and forward this short video. This enriches the ways to browse and share images, enhancing the user experience.

[0238] Optionally, the recommended statements can be displayed on the secondary page according to their priority. For example, when the secondary page is first displayed, the three highest-priority recommended statements are shown. When the user clicks "Change," the three recommended statements with lower priority than the third highest-priority statement are then displayed. Optionally, the priority of the recommended statements can be based on the number of images in the corresponding image set. The more images in the image set, the higher the priority of the corresponding recommended statement; the fewer images in the image set, the lower the priority of the corresponding recommended statement. Optionally, the priority of the recommended statements can also be determined based on the update time of the recommendation template corresponding to the recommended statement. The more recent the update time of the recommendation template, the higher the priority of the corresponding recommended statement; the more distant the update time of the recommendation template, the lower the priority of the corresponding recommended statement. Optionally, if a recommended statement exists in the history of recommended statements, it means that the recommended statement has already been recommended to the user, and the priority of that recommended statement is reduced, while the priority of other unrecommended recommended statements is increased. This can improve the differentiation of the recommended statements, making the generated short videos more differentiated and enriching the user experience.

[0239] In this embodiment, when a user activates the smart video creation function, the terminal device can automatically display recommended statements, guiding the user to generate a short video corresponding to the recommended statement. Optionally, the user can also input similar statements via voice or text based on the guidance of the recommended statements. The terminal device can then perform semantic recognition on the user's input statements to obtain the initial keyword combination corresponding to the statements. Then, it performs keyword conversion on the initial keywords in the initial keyword combination to obtain the corresponding search keyword combination. The terminal device can search the image library based on this search keyword combination to obtain the image set corresponding to the user's input statement and display it. The display method can be seen in... Figure 6 The dialogue format shown here uses user-inputted statements as the recommended phrases. For details regarding the acquisition of initial keyword combinations, keyword conversion, and searching the image library based on search keyword combinations, please refer to the preceding descriptions; they will not be repeated here.

[0240] Optionally, the above-described CV analysis process can also be executed after the user clicks on "Smart Image Generation". This application does not limit this approach.

[0241] Figure 8 A method for acquiring an atlas provided in this application includes:

[0242] S801. Obtain at least one component of the target recommendation template, wherein the at least one component includes at least a time element and / or a person element.

[0243] The constituent elements can be any one or more of the following: time, location, people, and events. These constituent elements can be described using configuration information in a configuration file. For details on configuration files and configuration information, please refer to the previous descriptions; they will not be repeated here.

[0244] S802. Transform the initial keywords of time elements and / or people elements to obtain time search keywords and / or people search keywords.

[0245] The initial keywords mentioned above are descriptive terms that conform to users' colloquial speech. The time element (initial keywords) can be descriptive terms such as "today," "tomorrow," "last week," etc., that conform to users' verbal expressions. The person element (initial keywords) can be descriptive terms such as "darling," "lover," etc., that conform to users' verbal expressions. After the terminal device converts the initial keywords, it obtains the corresponding time search keywords and / or person search keywords. The time search keywords are specific dates, and the person search keywords are the corresponding person's TagID or ReID.

[0246] S803. Search the image library according to the search keywords to obtain the target image set that matches the search keywords. The search keywords include at least time search keywords and / or people search keywords. The first image is any image in the target image set. The image elements of the first image match at least one component element. The image elements of the first image include at least the generation time of the first image and / or the person identifier.

[0247] Terminal devices can search the image library based on some or all of the search keywords, including time, people, location, and events, to obtain the target image set that matches the search keywords. A detailed description of the search image library can be found above and will not be repeated here.

[0248] Figure 8 In the illustrated embodiment, the terminal device can obtain corresponding image sets based on the image elements of an image, which can then be used to recommend images to the user. This allows for personalized image display according to a recommendation template, enhancing the user's browsing experience. Furthermore, by converting colloquial initial keywords into corresponding search keywords that the terminal device can recognize, image library searches can be achieved. Simultaneously, the colloquial initial keywords result in recommendation statements that align with users' conversational expression habits, leading to a better reading experience.

[0249] In some embodiments, at least one component element includes: a first component element and other components, wherein the first component element is any one of a time element, a person element, a location element, and an event element, and the other components element includes one or more of a time element, a person element, a location element, and an event element, and the first component element and any one of the other components element are different; when the first component element includes at least two initial keywords, the method further includes: combining each of the at least two initial keywords with the initial keywords of the other components to form at least two initial keyword combinations, wherein the first initial keyword combination is any one of the at least two initial keyword combinations, the first initial keyword combination includes multiple initial keywords, the multiple initial keywords include the first initial keyword and the second initial keyword, the component element corresponding to the first initial keyword and the component element corresponding to the second initial keyword are different components among at least one component element, and the multiple initial keywords include at least two initial keywords. Any one of the following: Transform the initial keywords of time elements and / or people elements to obtain time search keywords and / or people search keywords, including: when multiple initial keywords in the first initial keyword combination include time keywords, transform the time keywords to obtain the time search keywords corresponding to the time keywords; when multiple initial keywords in the first initial keyword combination include people keywords, transform the people keywords to obtain the people search keywords corresponding to the people keywords; search the image library based on the search keywords to obtain target image sets matching the search keywords, including: replacing the corresponding time keywords in the first initial keyword combination with time search keywords, and / or replacing the corresponding people keywords in the first initial keyword combination with people search keywords to obtain a first search keyword combination corresponding to the first initial keyword combination, wherein the first search keywords include at least one search keyword; search the image library based on at least one search keyword to obtain target image sets matching at least one search keyword.

[0250] The first component mentioned above can be any one of at least one component, and the other components are different from the first component. Each component can contain empty or one or more elements. Not all elements in a recommendation template will be empty. The terminal device can first combine different elements from different components to form all possible combinations of different elements from multiple components, thereby obtaining at least two initial keyword combinations. Taking the first initial keyword combination as an example, the initial keywords are converted to obtain the corresponding search keywords.

[0251] It should be noted that the aforementioned keywords for searching for people can be feature vectors for searching for people; the aforementioned keywords for searching for events can be semantic feature vectors for searching for events. See the preceding text for details.

[0252] The terminal device searches for the target image library based on the converted search keywords, along with other search keywords that do not require conversion from the same initial keyword group. For example, TagID-1 + February 10, 2024 + City C1 + ClipID-1. Optionally, one or more search keywords in the search keyword combination may be empty; the specific form of the search keyword combination is not limited here.

[0253] Based on this, terminal devices can generate various image sets corresponding to different elements, making the presented image sets more diverse and enriching the user experience.

[0254] In some embodiments, at least one component further includes: a location element; at least one search keyword further includes a location search keyword; and the image element of the first image further includes the location where the first image was generated.

[0255] In some embodiments, the location search keyword is one of a plurality of candidate locations, which are obtained by clustering the locations where multiple images in the gallery were generated.

[0256] In some embodiments, when the location element includes a first referential address, the method further includes: converting the first referential address into a first map address, and the location search keywords include the location name corresponding to the first map address.

[0257] In some embodiments, at least one component further includes: a person element, at least one search keyword further includes a person search keyword, and the image element of the first image further includes a person feature vector of the first image, the person feature vector of the first image being used to characterize the person appearing in the first image.

[0258] In some embodiments, the feature vector of the person in the first image is one or more of a plurality of feature vectors, which are obtained by clustering the faces and / or bodies of people appearing in multiple images in the image library.

[0259] Elements derived from the clustering of images in the image library, specifically person feature vectors and / or location elements, can effectively narrow the search scope, saving search time and resources. Simultaneously, it enables personalized recommendations, meeting user expectations and enhancing the user experience.

[0260] In some embodiments, at least one component further includes: an event element; at least one search keyword further includes an event search keyword; the image element of the first image further includes a semantic feature vector of the first image; the semantic feature vector of the first image is used to characterize the image content of the first image; and the image content of the first image includes the event corresponding to the event search keyword.

[0261] For a detailed description of the constituent elements and image elements, please refer to the previous text, which will not be repeated here.

[0262] In some embodiments, the method further includes: obtaining a first recommendation statement template of the target recommendation template, the first recommendation statement template including at least one element slot, the at least one element slot corresponding one-to-one with at least one component element; filling the first element slot with a first initial keyword of the first component element corresponding to the first element slot, generating a first recommendation statement corresponding to the first recommendation statement template, wherein the first element slot is any one of the at least one element slot, the first component element is one of the at least one component element corresponding to the first element slot, and the first initial keyword is one of the elements of the first component element.

[0263] The terminal device fills the corresponding element slots in the recommendation statement template with the aforementioned initial keywords, thereby generating the corresponding recommendation statement. A detailed description of this can be found in the relevant descriptions within the text, and will not be repeated here. The initial keywords in this recommendation statement conform to users' colloquial expression habits, making it easier to read. Furthermore, filling the statement based on element slots offers high flexibility and improves the user experience.

[0264] In some embodiments, the method further includes: receiving a first operation input by a user, the first operation being used to display an interface of a gallery application; responding to the first operation, displaying a first interface of the gallery application, the first interface including a first control; receiving a second operation performed by the user on the first control; responding to the second operation, displaying a second interface, the second interface displaying a first recommendation statement; receiving a third operation performed by the user on a second control of the first recommendation statement; responding to the third operation, displaying a third interface, the third interface displaying a target gallery.

[0265] The first action could be for the user to open the gallery application. The terminal device responds to the user's first action by displaying the gallery's initial interface, for example... Figure 4 As shown in Figure a, this first interface includes first controls, such as... Figure 4 The image in Figure a shows a smart photo card. Users can perform a second operation on the first control, such as clicking on a smart photo card. The terminal device then activates the smart photo function, displaying a second interface including one or more recommended statements.

[0266] In the second interface, the user performs a third operation on a control related to one of the recommended statements. For example, clicking the control for the first recommended statement (generating a video of Nanjing tourism) will display a third interface on the terminal device. This third interface includes a thumbnail of the target image set corresponding to the first recommended statement. Figure 6 As shown in the figure.

[0267] In some embodiments, the third interface further includes a third control, and the method further includes: receiving a fourth operation input by the user to the third control; in response to the fourth operation, generating a video file containing a target atlas according to a target recommendation template; and playing and / or storing the video file.

[0268] The third interface can also include third controls, such as... Figure 6 The "Generate Video" option is shown in the image. When a user performs a fourth operation on the third control, such as clicking "Generate Video," the terminal device generates a short video from the images in the target image set based on the recommended template. This achieves personalized display, enriches the image display methods, and enhances the user experience.

[0269] In some embodiments, before searching the image library based on search keywords to obtain a target image set matching the search keywords, the method further includes: analyzing the images in the image library to obtain image elements for each image, wherein the image elements include at least the generation time.

[0270] In some embodiments, the image element further includes: the location where the image was generated, which can be any one or more of a city, a POI area, and a custom location.

[0271] In some embodiments, the location where the image was generated is obtained by converting the longitude and latitude of the location where the image was generated.

[0272] In some embodiments, the image elements further include: a person feature vector and / or a semantic feature vector of the image.

[0273] The methods for image analysis and clustering are described above and will not be repeated here. The image elements mentioned above describe the image's attributes, facilitating classification.

[0274] The foregoing has detailed examples of the methods provided in this application. It is understood that the corresponding apparatus, in order to achieve the above functions, includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0275] This application can divide the atlas acquisition device into functional modules based on the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0276] Figure 9 A schematic diagram of a picture atlas acquisition device provided in this application is shown. The device 900 includes:

[0277] The acquisition module 901 is used to acquire at least one component of the target recommendation template, wherein the at least one component includes at least a time element and / or a person element.

[0278] The conversion module 902 is used to convert the initial keywords of time elements and / or people elements to obtain time search keywords and / or people search keywords.

[0279] Search module 903 is used to search in the image library according to search keywords to obtain a target image set that matches the search keywords. The search keywords include at least time search keywords and / or people search keywords. The first image is any image in the target image set. The image elements of the first image match at least one component element. The image elements of the first image include at least the generation time of the first image and / or a person identifier.

[0280] In some embodiments, at least one component includes: a first component and other components, wherein the first component is any one of a time element, a person element, a location element, and an event element, and the other components include one or more of a time element, a person element, a location element, and an event element, and each of the first component and other components is different;

[0281] When the first component element includes at least two initial keywords, the conversion module 902 is specifically used to combine each of the at least two initial keywords with the initial keywords of other component elements to form at least two initial keyword combinations. The first initial keyword combination is any one of the at least two initial keyword combinations. The first initial keyword combination includes multiple initial keywords, including a first initial keyword and a second initial keyword. The component element corresponding to the first initial keyword and the component element corresponding to the second initial keyword are different components from at least one component element. The multiple initial keywords include any one of the at least two initial keywords. Furthermore, when the multiple initial keywords in the first initial keyword combination include a time keyword, the time keyword is converted to obtain a time search keyword corresponding to the time keyword. When the multiple initial keywords in the first initial keyword combination include a person keyword, the person keyword is converted to obtain a person search keyword corresponding to the person keyword.

[0282] Search module 903 is specifically used to replace the corresponding time keyword in the first initial keyword combination with a time search keyword, and / or replace the corresponding person keyword in the first initial keyword combination with a person search keyword, to obtain a first search keyword combination corresponding to the first initial keyword combination, wherein the first search keyword includes at least one search keyword; and to search in the image library based on at least one search keyword to obtain a target image set that matches at least one search keyword.

[0283] In some embodiments, at least one component further includes: a location element; at least one search keyword further includes a location search keyword; and the image element of the first image further includes the location where the first image was generated.

[0284] In some embodiments, the location search keyword is one of a plurality of candidate locations, which are obtained by clustering the locations where multiple images in the gallery were generated.

[0285] In some embodiments, when the location element includes a first referential address, the conversion module 902 is further configured to convert the first referential address into a first map address, and the location search keywords include the location name corresponding to the first map address.

[0286] In some embodiments, at least one component further includes: a person element, at least one search keyword further includes a person search keyword, and the image element of the first image further includes a person feature vector of the first image, the person feature vector of the first image being used to characterize the person appearing in the first image.

[0287] In some embodiments, the feature vector of the person in the first image is one or more of a plurality of feature vectors, which are obtained by clustering the faces and / or bodies of people appearing in multiple images in the image library.

[0288] In some embodiments, at least one component further includes: an event element; at least one search keyword further includes an event search keyword; the image element of the first image further includes a semantic feature vector of the first image; the semantic feature vector of the first image is used to characterize the image content of the first image; and the image content of the first image includes the event corresponding to the event search keyword.

[0289] In some embodiments, the device 900 further includes: a recommendation module, configured to obtain a first recommendation statement template of a target recommendation template, the first recommendation statement template including at least one element slot, the at least one element slot corresponding one-to-one with at least one component element; fill the first element slot with a first initial keyword of a first component element corresponding to the first element slot, and generate a first recommendation statement corresponding to the first recommendation statement template, wherein the first element slot is any one of the at least one element slots, the first component element is one of the at least one component element corresponding to the first element slot, and the first initial keyword is one of the elements of the first component element.

[0290] In some embodiments, the device 900 further includes a receiving module for receiving a first operation input by a user, the first operation being used to display the interface of a gallery application.

[0291] The display module is used to respond to the first operation and display the first interface of the gallery application, which includes the first control.

[0292] The receiving module is also used to receive a second operation performed by the user on the first control.

[0293] The display module is also used to respond to the second operation by displaying a second interface, in which the first recommended statement is displayed.

[0294] The receiving module is also used to receive a third operation performed by the user on the second control of the first recommendation statement.

[0295] The display module is also used to respond to a third operation by displaying a third interface, in which the target atlas is displayed.

[0296] In some embodiments, the third interface further includes a third control and a receiving module, which is also used to receive a fourth operation input by the user on the third control.

[0297] The display module is also used to respond to the fourth operation by generating a video file containing the target atlas according to the target recommended template, and playing and / or storing the video file.

[0298] In some embodiments, the apparatus 900 further includes an image analysis module for analyzing images in the image library to obtain image elements for each image, wherein the image elements include at least the generation time.

[0299] In some embodiments, the image element further includes: the location where the image was generated, which can be any one or more of a city, a POI area, and a custom location.

[0300] In some embodiments, the location where the image was generated is obtained by converting the longitude and latitude of the location where the image was generated.

[0301] In some embodiments, the image elements further include: a person feature vector and / or a semantic feature vector of the image.

[0302] The specific method by which device 900 performs the atlas acquisition method and the beneficial effects thereof can be found in the relevant descriptions in the method embodiments, and will not be repeated here.

[0303] This application also provides an electronic device, including the processor described above. The electronic device provided in this embodiment may be... Figure 1 The terminal device 100 shown is used to execute the above-described atlas acquisition method. When using integrated units, the terminal device may include a processing module, a storage module, and a communication module. The processing module can be used to control and manage the actions of the terminal device; for example, it can support the terminal device in executing the steps performed by the display unit, detection unit, and processing unit. The storage module can support the terminal device in executing stored program code and data. The communication module can support communication between the terminal device and other devices.

[0304] The processing module can be a processor or a controller. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination of functions that implement computing capabilities, such as a combination of one or more microprocessors, a digital signal processor (DSP), and a microprocessor, etc. The storage module can be a memory. The communication module can specifically be a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, or other devices that interact with other terminal devices.

[0305] In one embodiment, when the processing module is a processor and the storage module is a memory, the terminal device involved in this embodiment can be a device having... Figure 1 The device with the structure shown.

[0306] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the atlas acquisition method described in any of the above embodiments.

[0307] This application also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the atlas acquisition method described in the above embodiments.

[0308] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0309] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units. The replaced units may or may not be physically separate. The component shown as a unit may be one physical unit or multiple physical units, that is, it may be located in one place or distributed in multiple different places. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0310] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0311] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0312] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for acquiring image atlases, characterized in that, include: Obtain at least one component of the target recommendation template, wherein the at least one component includes at least a time element and / or a person element; The initial keywords of the time element and / or person element are transformed to obtain time search keywords and / or person search keywords; The image library is searched according to the search keywords to obtain the target image set that matches the search keywords. The search keywords include at least the time search keywords and / or the people search keywords. Wherein, the first image is any image in the target image set, the image elements of the first image match the at least one component element, and the image elements of the first image include at least the generation time of the first image and / or the person identifier. The method further includes: Obtain the first recommendation statement template of the target recommendation template. The first recommendation statement template includes at least one element slot, and the at least one element slot corresponds one-to-one with the at least one component element. Fill the first element slot with the first initial keyword of the first component element corresponding to the first element slot, and generate the first recommended statement corresponding to the first recommended statement template. The first element slot is any one of the at least one element slots, the first component element is one of the at least one component elements corresponding to the first element slot, and the first initial keyword is one of the elements of the first component element. The user input is received as a first operation, which is used to display the interface of the gallery application. In response to the first operation, a first interface of the gallery application is displayed, the first interface including a first control; Receive the second operation performed by the user on the first control; In response to the second operation, a second interface is displayed, in which the first recommended statement is displayed; Receive a third operation performed by the user on the second control of the first recommendation statement; In response to the third operation, a third interface is displayed, in which the target atlas is displayed.

2. The method according to claim 1, characterized in that, The at least one component includes: a first component and other components, wherein the first component is any one of the time element, the person element, the location element, and the event element, and the other components include one or more of the time element, the person element, the location element, and the event element, and the first component and any one of the other components are different; When the first component includes at least two initial keywords, the method further includes: Each of the at least two initial keywords is combined with the initial keywords of the other constituent elements to form at least two initial keyword combinations. The first initial keyword combination is any one of the at least two initial keyword combinations. The first initial keyword combination includes multiple initial keywords, including a first initial keyword and a second initial keyword. The constituent element corresponding to the first initial keyword and the constituent element corresponding to the second initial keyword are different constituent elements among the at least one constituent element. The multiple initial keywords include any one of the at least two initial keywords. The process of converting the initial keywords of the time element and / or person element to obtain time search keywords and / or person search keywords includes: When the initial keywords in the first initial keyword combination include time keywords, the time keywords are converted to obtain the time search keywords corresponding to the time keywords; When the first initial keyword combination includes a person keyword among the multiple initial keywords, the person keyword is converted to obtain the person search keyword corresponding to the person keyword; The step of searching the image library based on search keywords to obtain a target image set matching the search keywords includes: The time-related search keywords are replaced with the time-related search keywords, and / or the person-related search keywords are replaced with the person-related search keywords, to obtain a first search keyword combination corresponding to the first initial keyword combination, wherein the first search keywords include at least one search keyword; A search is performed in the image library based on the at least one search keyword to obtain the target image set that matches the at least one search keyword.

3. The method according to claim 2, characterized in that, The at least one component further includes: a location element; the at least one search keyword further includes a location search keyword; and the image element of the first image further includes the location where the first image was generated.

4. The method according to claim 3, characterized in that, The location search keyword is one of a plurality of candidate locations, which are obtained by clustering multiple images in the image library based on their origin.

5. The method according to claim 3 or 4, characterized in that, When the location element includes a first referential address, the method further includes: The first referential address is converted into a first map address, and the location search keywords include the location name corresponding to the first map address.

6. The method according to any one of claims 2 to 5, characterized in that, The at least one component further includes: a person element; the at least one search keyword further includes a person search keyword; the image element of the first image further includes a person feature vector of the first image; the person feature vector of the first image is used to characterize the person appearing in the first image.

7. The method according to claim 6, characterized in that, The feature vector of the person in the first image is one or more of a plurality of feature vectors, which are obtained by clustering the faces and / or bodies of people appearing in multiple images in the image library.

8. The method according to any one of claims 2 to 7, characterized in that, The at least one component further includes: an event element; the at least one search keyword further includes an event search keyword; the image element of the first image further includes a semantic feature vector of the first image; the semantic feature vector of the first image is used to characterize the image content of the first image; and the image content of the first image includes the event corresponding to the event search keyword.

9. The method according to claim 1, characterized in that, The third interface also includes a third control, and the method further includes: Receive a fourth operation input by the user for the third control; In response to the fourth operation, a video file containing the target atlas is generated according to the target recommendation template; Play and / or store the video file.

10. The method according to claim 1, characterized in that, Before searching the image library based on search keywords to obtain a target image set matching the search keywords, the method further includes: The images in the image library are analyzed to obtain the image elements of each image, and the image elements include at least the generation time.

11. The method according to claim 10, characterized in that, The image elements also include: the location where the image was generated, which can be any one or more of a city, a POI area, and a custom location.

12. The method according to claim 11, characterized in that, The location where the image was generated is obtained by converting the longitude and latitude at the time of image generation.

13. The method according to claim 12, characterized in that, The image elements also include: the image's human feature vector and / or semantic feature vector.

14. An electronic device, characterized in that, include: Processor, memory, and interface; The processor, the memory, and the interface cooperate with each other to enable the electronic device to perform the method as described in any one of claims 1 to 13.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Method and system for searching files based on natural language and computing equipment

    CN117235014A

  • System and method for creating and sharing photo stories

    US20110280497A1