Image set acquisition method, electronic equipment and computer readable storage medium

By obtaining and converting image elements in the gallery, and using the target recommendation template to generate a personalized picture collection, it solves the problem of a single traditional image display method and improves the user experience.

CN120561329AActive Publication Date: 2025-08-29HONOR DEVICE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202410193106.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2025-08-29
Estimated Expiration
2044-02-20

AI Technical Summary

Technical Problem

The traditional image display method is single, the user experience is not high, and it is difficult to meet the diverse browsing needs of users.

Method used

By obtaining the components of the target recommendation template, including time, place, person and event elements, using colloquial keywords to convert them into search keywords that can be recognized by the terminal device, perform gallery searches, generate personalized picture albums and display them in person.

Benefits of technology

It realizes diversified display of pictures, improves users' browsing experience, conforms to users' spoken expression habits, and enriches the presentation method of pictures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561329A_ABST
    Figure CN120561329A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of terminals, and provides an image set acquisition method, electronic equipment and a computer readable storage medium, the method comprises the following steps: acquiring at least one component of a target recommendation template, the at least one component at least comprising a time element and / or a character element; converting the initial keyword of the time element and / or the character element to obtain a time search keyword and / or a character search keyword; according to the search keywords, searching is carried out in a picture library, a target picture set matched with the search keywords is obtained, and the search keywords at least comprise time search keywords and / or character search keywords. According to the method, the corresponding picture set can be obtained according to the picture elements of the picture and is used for recommending to the user, so that the picture is displayed in a personalized manner according to the recommendation template, and the browsing experience of the user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a method for acquiring an atlas, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the development of Internet technology and terminal technology, the use of terminal equipment has become more and more extensive and penetrated into people's production and life, becoming an inseparable part.

[0003] Current terminal devices are no longer limited to traditional communication functions; they also have the ability to take photos and store captured images. When a terminal device stores a large number of images, it can also categorize and store them for easier viewing. For example, a user can open the Gallery app and tap the "Camera" folder to view photos taken with the camera app; they can also tap the "Screenshots" folder to view screenshots.

[0004] However, the traditional way of displaying pictures is relatively simple and the user experience is not good. Summary of the Invention

[0005] The present application provides a method, device, chip, electronic device, computer-readable storage medium and computer program product for obtaining an atlas, which can enhance the user experience.

[0006] In a first aspect, a method for acquiring an atlas is provided, comprising: acquiring at least one component element of a target recommendation template, wherein the at least one component element includes at least a time element and / or a person element; converting initial keywords of the time element and / or the person element to obtain a time search keyword and / or a person search keyword; searching in an atlas according to the search keyword to obtain a target atlas matching the search keyword, wherein the search keyword includes at least a time search keyword and / or a person search keyword; wherein the first picture is any picture in the target atlas, the picture element of the first picture matches the at least one component element, and the picture element of the first picture includes at least the generation time of the first picture and / or a person identification.

[0007] The constituent elements may be any one or more of the time elements, location elements, person elements and event elements. The constituent elements may be described using the configuration information in the configuration file. The above-mentioned initial keywords are descriptive words that conform to the user's spoken language. The elements of the time element (initial keywords) may be descriptive words such as today, tomorrow, last week, etc. that conform to the user's verbal expression. The elements of the person element (initial keywords) may be descriptive words such as Dabao and lover that conform to the user's verbal expression. After the terminal device converts the initial keywords, it obtains the corresponding time search keywords and / or person search keywords. The time search keyword is a specific date, and the person search keyword is the TagID or ReID of the corresponding person. The terminal device can search the gallery based on part or all of the time search keywords, person search keywords, location search keywords and event search keywords to obtain the target atlas that matches the search keywords.

[0008] The terminal device can use the corresponding atlas obtained from the image elements of the image to recommend to the user, thereby personalizing the display of the image according to the recommendation template, which can enhance the user's browsing experience. In addition, by converting the colloquial initial keywords into corresponding search keywords that the terminal device can recognize for searching, it is possible to search the gallery. At the same time, the recommendation sentences generated later by the colloquial initial keywords conform to the user's colloquial expression habits, which provides a better reading experience.

[0009] In some possible implementations, at least one component element includes: a first component element and other components, the first component element is any one of a time element, a person element, a place element and an event element, the other components include one or more of a time element, a person element, a place element and an event element, and any one of the first component element and the other components are different; when the first component element includes at least two initial keywords, the method further includes: combining each of the at least two initial keywords with the initial keywords of the other components to form at least two initial keyword combinations, wherein the first initial keyword combination is any one of the at least two initial keyword combinations, the first initial keyword combination includes multiple initial keywords, the multiple initial keywords include the first initial keyword and the second initial keyword, the component element corresponding to the first initial keyword and the component element corresponding to the second initial keyword are different component elements in at least one component, and the multiple initial keywords include at least two initial keywords any one of them; converting the initial keywords of the time element and / or the person element to obtain the time search keyword and / or the person search keyword, including: when the multiple initial keywords in the first initial keyword combination include the time keyword, converting the time keyword to obtain the time search keyword corresponding to the time keyword; when the multiple initial keywords in the first initial keyword combination include the person keyword, converting the person keyword to obtain the person search keyword corresponding to the person keyword; searching in the gallery according to the search keyword to obtain the target atlas matching the search keyword, including: replacing the corresponding time keyword in the first initial keyword combination with the time search keyword, and / or replacing the corresponding person keyword in the first initial keyword combination with the person search keyword, to obtain the first search keyword combination corresponding to the first initial keyword combination, the first search keyword including at least one search keyword; searching in the gallery according to at least one search keyword to obtain the target atlas matching at least one search keyword.

[0010] The above-mentioned first component is any one of at least one component, and the other components are different from the first component. If the elements in each component can be empty, it can also be one or more. The elements of all components in a recommendation template will not all be empty. The terminal device can first combine different elements in different components to form all combinations of different elements in multiple components, thereby obtaining at least two groups of initial keyword combinations. Taking the first initial keyword combination as an example, the initial keywords therein are converted into keywords to obtain corresponding search keywords.

[0011] It should be noted that the above-mentioned person search keyword can be a person's name or a person feature vector used to search for a person; the above-mentioned event search keyword can be an event name or a semantic feature vector used to search for an event. For details, please refer to the above description.

[0012] The terminal device searches the image library based on the converted search keywords, along with other search keywords corresponding to the same initial keyword combination that do not require conversion, to obtain the corresponding target atlas. Optionally, one or more search keywords in the search keyword combination may be empty. The specific form of the search keyword combination is not limited here.

[0013] Based on this, the terminal device can generate atlases corresponding to a variety of different elements, making the presented atlases diversified and enriching the user experience.

[0014] In some possible implementations, at least one component element further includes: a location element, at least one search keyword further includes a location search keyword, and the image element of the first image further includes the location where the first image was generated.

[0015] In some possible implementations, the location search keyword is one of a plurality of candidate locations, and the plurality of candidate locations are obtained by clustering locations generated according to a plurality of pictures in a gallery.

[0016] In some possible implementations, when the place element includes a first reference address, the method further includes: converting the first reference address into a first map address, and the place search keyword includes a place name corresponding to the first map address.

[0017] In some possible implementations, at least one component element also includes: a character element, at least one search keyword also includes a character search keyword, and the image element of the first image also includes a character feature vector of the first image, and the character feature vector of the first image is used to represent the character appearing in the first image.

[0018] In some possible implementations, the character feature vector of the first image is one or more of a plurality of character feature vectors, and the plurality of character feature vectors are obtained by clustering faces and / or bodies of characters appearing in a plurality of images in a gallery.

[0019] By clustering the images in the gallery, we can obtain the character feature vectors and / or location elements, effectively narrowing the search scope and saving search time and resources. At the same time, we can also achieve personalized recommendations, meet user expectations, and improve the user experience.

[0020] In some possible implementations, at least one component element also includes: an event element, at least one search keyword also includes an event search keyword, the picture element of the first picture also includes a semantic feature vector of the first picture, the semantic feature vector of the first picture is used to represent the picture content of the first picture, and the picture content of the first picture includes the event corresponding to the event search keyword.

[0021] For detailed descriptions of the components and image elements, please refer to other relevant descriptions in the article and will not be repeated here.

[0022] In some possible implementations, the method also includes: obtaining a first recommendation statement template of the target recommendation template, the first recommendation statement template including at least one element slot, and the at least one element slot corresponds one-to-one to at least one component element; filling the first initial keyword of the first component element corresponding to the first element slot in the first element slot, generating a first recommendation statement corresponding to the first recommendation statement template, the first element slot is any one of the at least one element slot, the first component element is one of the at least one component element corresponding to the first element slot, and the first initial keyword is one of the elements of the first component element.

[0023] The terminal device populates the initial keywords into the corresponding element slots in the recommendation sentence template, generating the corresponding recommendation sentence. A detailed description of this can also be found in the previous section and will not be repeated here. The initial keywords in this recommendation sentence align with the user's spoken language, making it easier to read. Furthermore, by populating the sentence with element slots, the sentence is highly flexible and improves the user experience.

[0024] In some possible implementations, the method also includes: receiving a first operation input by the user, the first operation being used to display the interface of the gallery application; in response to the first operation, displaying the first interface of the gallery application, the first interface including a first control; receiving a second operation performed by the user on the first control; in response to the second operation, displaying the second interface, the first recommendation statement being displayed in the second interface; receiving a third operation performed by the user on the second control of the first recommendation statement; in response to the third operation, displaying the third interface, the target gallery being displayed in the third interface.

[0025] In some possible implementations, the third interface also includes a third control, and the method also includes: receiving a fourth operation input by the user for the third control; in response to the fourth operation, generating a video file containing a target atlas according to the target recommendation template; and playing and / or storing the video file.

[0026] The terminal device can respond to user operations and generate short videos from the pictures in the above target collection according to the target recommendation template to achieve personalized display, enrich the picture display method, and improve the user experience.

[0027] In some possible implementations, before searching the gallery based on the search keywords and obtaining the target gallery that matches the search keywords, the method further includes: analyzing the images in the gallery to obtain the image elements of each image, where the image elements include at least the generation time.

[0028] In some possible implementations, the picture element further includes: a location where the picture is generated, where the generation location is any one or more of a city, a POI area, and a custom location.

[0029] In some possible implementations, the location where the image is generated is converted based on the longitude and latitude at which the image is generated.

[0030] In some possible implementations, the picture elements further include: a character feature vector and / or a semantic feature vector of the picture.

[0031] The image analysis and clustering methods can be found in the previous article and will not be repeated here. The above image elements can describe the attributes of the image and facilitate classification.

[0032] In a second aspect, a device for acquiring an atlas is provided, comprising a unit composed of software and / or hardware, which is used to execute any one of the methods in the technical solution described in the first aspect.

[0033] In a third aspect, an embodiment of the present application provides a chip comprising a processor; the processor is used to read and execute a computer program stored in a memory to execute any one of the methods in the technical solution described in the first aspect.

[0034] Optionally, the chip further includes a memory, and the memory is connected to the processor via a circuit or wire.

[0035] Further optionally, the chip also includes a communication interface.

[0036] In a fourth aspect, an electronic device is provided, comprising: a processor, a memory, and an interface; the processor, the memory, and the interface cooperate with each other so that the electronic device executes any one of the methods in the technical solution described in the first aspect.

[0037] In a fifth aspect, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the processor executes any one of the methods in the technical solution described in the first aspect.

[0038] In a sixth aspect, a computer program product is provided, comprising: a computer program code, which, when executed on an electronic device, enables the electronic device to execute any one of the methods in the technical solution described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 1 is a schematic structural diagram of a terminal device 100 provided in an embodiment of the present application;

[0040] Figure 2 is a software structure block diagram of the terminal device 100 provided in an embodiment of the present application;

[0041] Figure 3 This is an interactive diagram of an atlas acquisition method provided in an embodiment of the present application;

[0042] Figure 4 This is a schematic diagram of an example of a gallery interface provided in an embodiment of the present application;

[0043] Figure 5 This is another example of an interface diagram related to smart photo creation in a gallery interface provided in an embodiment of the present application;

[0044] Figure 6 This is another example of an interface diagram related to smart filming provided in an embodiment of the present application;

[0045] Figure 7 This is another example of an interface diagram related to smart filming provided in an embodiment of the present application;

[0046] Figure 8 This is a flow chart of a method for obtaining an atlas provided in an embodiment of the present application;

[0047] Figure 9 This is another structural diagram of an atlas acquisition device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0049] In the following, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Therefore, a feature specified as "first," "second," or "third" may explicitly or implicitly include one or more of the features.

[0050] The atlas acquisition method provided in the embodiments of the present application can be applied to terminal devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific types of terminal devices.

[0051] For example, Figure 1 1 is a schematic diagram of the structure of an example terminal device 100 provided in an embodiment of the present application. The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0052] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0053] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may also adopt a different interface connection method from the above embodiments, or a combination of multiple interface connection methods.

[0054] The software system of the terminal device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software structure of the terminal device 100.

[0055] Figure 2 This is a software structure diagram of the terminal device 100 in an embodiment of the present application. The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer can include a series of application packages.

[0056] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0057] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0058] like Figure 2 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0059] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0060] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0061] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.

[0062] The phone manager is used to provide communication functions of the terminal device 100, such as management of call status (including answering, hanging up, etc.).

[0063] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0064] The notification manager enables applications to display notification information in the status bar, which can be used to convey informational messages and disappear automatically after a short stay without user interaction.

[0065] In an embodiment of the present application, the application framework layer also includes a media recommendation service, cloud parameters, and a recommendation database.

[0066] Cloud-based parameters include configuration files for recommended templates. This configuration file can be downloaded from the cloud or pre-installed on the terminal device and regularly updated from the cloud.

[0067] The media recommendation service (MrsgService) includes a media recommendation generation engine and a recommendation template parser. The media recommendation generation engine and recommendation template parser can adapt the corresponding image collections for different recommendation templates according to the configuration file and generate corresponding recommendation statements.

[0068] The recommendation database is used to store adapted atlases and recommendation statements for later presentation.

[0069] The Android runtime includes the core library and the virtual machine. The Android runtime is responsible for scheduling and management of the Android system.

[0070] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.

[0071] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0072] The system library can include multiple functional modules, such as a surface manager, media libraries, a 3D graphics processing library (such as OpenGL ES), and a 2D graphics engine (such as SGL).

[0073] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0074] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0075] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0076] A 2D graphics engine is a drawing engine for 2D drawings.

[0077] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0078] For ease of understanding, the following examples of this application will be described with Figure 1 and Figure 2 Taking the terminal device with the shown structure as an example, the atlas acquisition method provided in the embodiment of the present application is specifically explained in combination with the accompanying drawings and application scenarios.

[0079] Typically, a user can use a camera application (hereinafter referred to as the camera application) installed on a terminal device to take photos. The photos taken may include photos of people, photos of scenery, photos of other objects, or photos of different objects together. Photos taken by the camera application can be stored in a gallery application (hereinafter referred to as the gallery). The terminal device can store not only photos taken by the camera application, but also pictures from other sources, such as screenshots obtained by taking screenshots, pictures downloaded from web pages, and pictures transmitted through other communication tools (such as chat applications).

[0080] When users want to browse images, they can open the gallery. For example, they can open the gallery and tap the "Camera" folder to view photos taken with the camera app; they can also tap the "Screenshots" folder to view screenshots. Due to the limited screen size of terminal devices, the number of images that can be displayed at one time is limited. Users can also enter a folder and flip through the pages to view other images within it. Traditional image display methods are relatively simple, and the way users browse images is relatively monotonous, resulting in a poor user experience.

[0081] The atlas acquisition method provided in the embodiment of the present application can cluster and analyze the pictures in the gallery according to the picture elements based on one or more elements, and then screen out pictures from the gallery whose picture elements match the components of the recommendation template according to the components corresponding to the recommendation template to form an atlas, so that the formed atlas can be different according to the picture elements of different pictures. It should be noted that the corresponding atlas obtained by this method based on the picture elements of the pictures is used to recommend to the user, so that the pictures are personalized according to the recommendation template. This way of displaying pictures is richer and more diversified, and can enhance the user's browsing experience.

[0082] In the embodiment of the present application, when a new image is stored in the gallery, the terminal device can perform computer vision (CV) analysis on each image stored in the gallery in the background. The CV analysis can include the following three aspects:

[0083] Aspect 1: Location information conversion.

[0084] Specifically, the terminal device may convert the longitude and latitude location information into a specific generation location, such as a city and / or a point of interest (POI).

[0085] When a terminal device generates an image, it can record the image's location information and then convert that location information into a specific location. For example, if the terminal device's positioning function (e.g., GPS) is enabled when taking a photo using a camera app, the device can obtain the latitude and longitude at the time of capture and record them as the location information for the captured photo. For example, if the latitude and longitude of the terminal device at the time of capturing image P1 are (X+Y) / 2 east longitude and (Z+W) / 2 north latitude, and the longitude range of city C1 on Earth is from XX to YY degrees east longitude and from ZZ to WW degrees north latitude, the terminal device can determine, through CV analysis, that the location of image P1 is city C1, thereby converting the location information of image P1 into the location of city C1. Alternatively, the terminal device can also determine the POI area corresponding to the longitude and latitude ranges of pre-set POI areas as the location of image P1. POI areas can optionally be tourist attractions, office buildings, residential communities, and other areas.

[0086] If the positioning function of the terminal device is not turned on when the photo is taken, and the terminal device cannot obtain the longitude and latitude at the time, the location information of the taken picture may be empty and the generated location may also be empty.

[0087] Optionally, if the positioning function of the terminal device is not enabled when taking a picture, the user can also manually enter the location information of the picture. For example, the user can directly enter the current longitude and latitude for the picture. In this case, the terminal device can store the longitude and latitude entered by the user as the location information of the taken picture. The user can also directly enter the name of the current city, the name of a scenic spot, or the name of a user-defined place, such as "old place". The terminal device can then store the name of the place entered by the user as the location where the taken picture was generated. Optionally, the terminal device can also determine the current area based on the coverage of the local area network to which it is currently connected, and use the current area as the location where the picture was generated.

[0088] After the terminal device performs location information conversion on multiple images, the locations where these multiple images were generated can be obtained. The terminal device can also perform cluster analysis on the locations where multiple images were generated, for example, taking the union of the locations where multiple images were generated to obtain multiple locations. Optionally, this generation location can be a city and / or a POI area, which is not limited in this embodiment of the present application. Some or all of these multiple locations can be used in the configuration information of the recommendation template to constitute the location elements of the location factor.

[0089] Aspect 2: Extracting character feature vectors.

[0090] Specifically, the terminal device can use a face recognition model to perform face recognition on the faces appearing in the image, thereby obtaining face feature vectors corresponding to different faces. It should be noted that if there are multiple faces in a picture, a corresponding face feature vector can be extracted for each face in the picture, thereby obtaining multiple face feature vectors. After the terminal device extracts multiple face feature vectors from multiple pictures, it can also perform cluster analysis on the extracted multiple face feature vectors to obtain a character feature vector representing each character. If the face of the same character appears in multiple pictures, the faces of the character in these multiple pictures correspond to the same character feature vector (TagID). It should be noted that a character list can also be set in the terminal device, and the name of each character in the character list corresponds to a TagID. In other words, each TagID represents a natural person. Then, if a picture contains multiple faces, the picture can extract multiple TagIDs that correspond one to one with the faces.

[0091] Optionally, the terminal device can also use a human body recognition model to identify the morphology of the human body appearing in the picture, thereby obtaining human body feature vectors corresponding to different human bodies. It should be noted that if there are multiple human bodies in a picture, a corresponding human body feature vector can be extracted for each human body in the picture, thereby obtaining multiple human body feature vectors. After the terminal device extracts multiple human body feature vectors from multiple pictures, it can also perform cluster analysis on the extracted multiple human body feature vectors to obtain a character feature vector representing each character. If the human body of the same character appears in multiple pictures, the human body of the character in these multiple pictures corresponds to the same character feature vector (ReID). It should be noted that a character list can also be set in the terminal device, and the name of each character in the character list corresponds to a ReID. In other words, each ReID represents a natural person. Then, if a picture contains multiple human bodies, the picture can extract multiple ReIDs corresponding one to one to the human bodies.

[0092] Optionally, the terminal device can also use a combination of face recognition model and body recognition model to identify both faces and bodies in the picture, as long as it can obtain the TagID, ReID, or comprehensive ID that represents each person.

[0093] The following description uses TagID as an example. The names of the people included in the above-mentioned person list can be entered by the user. For example, after the terminal device identifies and clusters the people in multiple pictures, it obtains two TagIDs for the people: TagID-1 and TagID-2. For example, the union of the facial and / or body feature vectors of multiple people is taken to obtain multiple person feature vectors. The names of the people corresponding to these multiple person feature vectors can be used as multiple elements in the person element. When the user browses the photo album, if the user chooses to browse by portrait, the terminal device can display two folders. Folder 1 stores pictures of the person corresponding to TagID-1, and folder 2 stores pictures of the person corresponding to TagID-2. Folder 1 can use a picture corresponding to TagID-1 as its cover, and folder 2 can use a picture corresponding to TagID-2 as its cover. The user can name each of these two folders separately. For example, if the user names folder 1 corresponding to TagID-1 "Little Daughter", it means that the person corresponding to TagID-1 is the Little Daughter. If the user names the folder 2 corresponding to TagID-2 as Lover, it means that the person corresponding to TagID-2 is Lover.

[0094] Users can also label a person's name on a picture individually. For example, if a picture is labeled "little daughter," it means that the people corresponding to the TagID of the picture include the little daughter. If the TagID of other pictures contains the same TagID as the one in the picture, it means that the people corresponding to the other pictures also include the little daughter.

[0095] The user can use any of the above methods to mark different characters in the album one by one, and part or all of the names of these characters can be used in the configuration information of the recommendation template to constitute the character elements of the character factors.

[0096] At the same time, the names of the characters marked by the user, such as the little daughter and the lover, can be added into the character list. In other words, part or all of the names of the characters in the character list can be used as character elements that constitute the character elements.

[0097] If the user does not input the name of a person for a certain TagID, the terminal device may set the name of the person corresponding to the TagID to be empty, for example, displaying: Unnamed (or adding a name).

[0098] It should be noted that the face and body of the same person can correspond to one or more TagIDs. If a picture includes multiple people, it can correspond to multiple TagIDs.

[0099] Alternatively, the same face or body may be identified as TagID-1 and ReID-1, or the same face may be identified as different TagIDs due to recognition errors. The user can then operate the terminal device to change the names of the people corresponding to TagID-1 and ReID-1 to the same name, thereby associating the person's feature vectors and defining TagID-1 and ReID-1 as the same person. During the image search process, if you want to search for pictures of the person, you can filter out pictures whose feature vectors are TagID-1 or ReID-1.

[0100] When a new picture is added to the gallery, the terminal device can also perform face and / or body recognition on the new picture to obtain the TagID corresponding to the picture. Optionally, if the new picture contains a person already in the album, the picture can be mapped to the existing TagID. Optionally, if the new picture contains a person not in the album, a new TagID can be identified to correspond to the newly added person.

[0101] Optionally, the above-mentioned face recognition model and body recognition model may be trained neural network models.

[0102] Aspect three: Extracting semantic feature vectors.

[0103] Specifically, the terminal device can use a semantic recognition model to perform semantic recognition on the image, thereby obtaining a semantic feature vector (ClipID) corresponding to each image. A semantic feature vector represents the image content contained in the corresponding image. Optionally, the image content may include objects different from people, such as pets, cats, puppies, food, etc.; it may also include some scenes, such as painting, traveling, sports, dancing, weddings, parties, etc. The embodiment of the present application does not limit the image content that can be recognized. The objects and scenes contained in the image content of a picture can be one or more. For example, if the picture is a scene of multiple people having a meal outdoors, and there is a pet dog in the picture, the semantic feature vector obtained by recognition can represent the image content in the picture including: traveling, party, pets, pet dogs, etc.

[0104] Optionally, the semantic recognition model may be a trained neural network model.

[0105] The above-mentioned CV analysis process can be performed when the terminal device is in an idle period. Among them, the idle period can be a period when the user does not use the terminal device, such as the period when the screen is off during charging, from 2:00 to 4:00 in the morning, or the period when the screen is off. This embodiment of the application does not limit this. The terminal device performs CV analysis during the idle period, which does not affect the user's use of the terminal device and is highly reasonable.

[0106] In addition to the individual pictures stored in the gallery, the objects of the above-mentioned CV analysis may also include frame images in the video files in the gallery. The CV analysis of the video file may be to perform CV analysis on the frame images in the video file one by one, or to extract frames from the video to obtain frame images, and then perform the above-mentioned CV analysis on the frame images to obtain the CV analysis results of the frame images. Optionally, in the frame extraction process, the terminal device may choose to extract a frame at a fixed number of frame intervals, such as extracting a frame every ten frames, or to extract a frame at a fixed time interval, such as extracting a frame every second. The embodiment of the present application does not limit the specific method of frame extraction. Taking the frame images of the video as the object of CV analysis can expand the selection range of the atlas, make the screened atlas more comprehensive, and make the later presentation richer, which can enhance the user experience.

[0107] As can be seen from the above, the result of CV analysis is the picture elements obtained by identifying the picture, which may include: one or more of the location, character feature vectors, and semantic feature vectors. It should be noted that when the pictures in the gallery are generated, the terminal device will also record the generation time of the picture. Specifically, the generation time of the picture may include year, month, and day, such as February 15, 2023; it may also include hours and minutes, such as 14:00. The generation time of each picture, together with one or more of the location, character feature vector, and semantic feature vector of the picture, can form a corresponding relationship with the picture, and the pictures are connected and stored together in the recommendation database for subsequent use.

[0108] In addition, the terminal device can also parse the recommended template according to the configuration file of the recommended template during idle time or other triggering time. Optionally, the terminal device can pre-set the configuration file of the recommended template locally or download the configuration file of the recommended template from the cloud, which is not limited in the embodiment of the present application. Optionally, the terminal device can also periodically query the cloud for new configuration files. When a new configuration file is available in the cloud, the new configuration file can be downloaded to replace the original configuration file, thereby ensuring that the configuration file is updated in a timely manner.

[0109] Optionally, when downloading a new configuration file, the terminal device may also determine whether the terminal device currently supports the new configuration file. If so, the terminal device may download the new configuration file to replace the original configuration file. If the terminal device currently does not support the new configuration file, for example, because the installed system version on the terminal device does not match the configuration file, the terminal device may use the original configuration file for template analysis.

[0110] The configuration information in the above configuration file can be called cloud parameters, which can be used to describe the recommendation template. Optionally, the configuration file can include multiple topics. Each topic includes multiple different recommendation templates. The configuration information of each recommendation template includes four components. Specifically, the four components of the recommendation template include: time element (time), location element (location), person element (human) and event element (event). In order to clearly describe the meaning of the configuration information of the recommendation template, a recommendation template is used as an example for explanation:

[0111] Component 1: Time element.

[0112] The configuration information of the recommendation template may include a time element. Specifically, the time element may include one or more time descriptors, each of which is one of the elements of the time element, called a time element. The time descriptors do not need to use dates, but instead use colloquial descriptors, so that the description in the later presentation conforms to the user's expression habits, thereby improving the user experience. For example, the time descriptors (time elements) may include, but are not limited to, one or more of: today, yesterday, last weekend, this week, last week, this month, this year, last year, the year before last, New Year's Eve, Spring Festival, Mid-Autumn Festival, Father's Day, Mother's Day, Children's Day, Valentine's Day, Chinese Valentine's Day, Christmas, Thanksgiving, wedding anniversary, graduation anniversary, and birthday. The number and type of time elements in the configuration information of different recommendation templates may be the same or different.

[0113] Component 2: Location element.

[0114] In an embodiment of the present application, the location element in the configuration information can be empty. In the actual template parsing process, since the locations cannot be exhaustively listed, the location element can be obtained by clustering the generation locations of multiple pictures obtained by CV analysis of the pictures. Therefore, compared with the method of listing a large number of existing locations to cover user needs, there is no need to retrieve invalid locations that are not related to the pictures in the gallery. It is only necessary to retrieve the generation locations of the pictures obtained by clustering the pictures, which can effectively narrow the search scope. For example, in the above-mentioned CV analysis process, if the results of the CV analysis of multiple pictures are clustered, the locations of some pictures are C1, the locations of some pictures are scenic spot B2, and the locations of other pictures are empty. Then, during the template parsing process, if the recommended template includes the location element, the terminal device can only search according to the two locations of city C1 and scenic spot B2, without searching for pictures according to other locations. The specific template parsing process can be found in the description below, which will not be repeated here.

[0115] Component three: character elements.

[0116] The configuration information of the recommended template also includes character elements. Specifically, character elements can include one or more character descriptors, with each descriptor being an element of the character element, referred to as a character element. Character descriptors can use colloquial terms such as nicknames, given names, and character relationship descriptors, including but not limited to: lover, eldest daughter, youngest son, boyfriend, Zhang San, grandfather, grandmother, grandfather, grandmother, etc. Optionally, the character element in the recommended template configuration information can be empty. During the actual template parsing process, character elements can be obtained by exhaustively enumerating the names of characters or character relationships, or through cluster analysis of images. Using image cluster analysis to obtain character element elements effectively narrows the search scope compared to enumerating the names of a large number of characters or character relationships. For example, in the aforementioned CV analysis process, if the CV analysis results of multiple images are clustered and the only characters in all images are the youngest daughter and the lover, then during the template parsing process, the search can be performed solely on these two characters, eliminating the need to search for other characters.

[0117] Component four: event element.

[0118] The configuration information of the recommended template also includes event elements. Specifically, the event element may include one or more descriptive words for the event, and each descriptive word is one of the elements of the event element, which is called an event element. The descriptive words for the types of event elements can also use colloquial descriptive words, so that the description in the later presentation conforms to the user's expression habits, thereby improving the user experience. The descriptive words for event elements include but are not limited to: one or more of: outings, daily sports, hobbies, daily gatherings, group photos, food, food preparation processes, cats, dogs, pets, travel experiences, travel, happy gatherings, team-building dinners, performances, painting, singing, dancing, etc. Optionally, since the types of events cannot be exhaustive, the fields about events in the configuration information can be used as open fields and written by R&D personnel during the R&D stage. In the actual template parsing process, the terminal device can search for images based on the fields of the events passed in by the R&D personnel, which is referred to as value-passing search matching.

[0119] The configuration information of the above-mentioned recommendation template also includes: a recommendation statement template. One or more of the multiple element slots are set in the recommendation statement template. These multiple element slots may include: a time slot for filling the time element, a location slot for filling the location element, a character slot for filling the character element, and an event slot for filling the event element. After the elements of the above-mentioned four components are filled into the corresponding element slots, a complete recommendation statement can be generated. It should be noted that the recommendation statement templates in the configuration information of different recommendation templates are different. The same recommendation template uses the same recommendation statement template, but based on the different elements of the four components, the generated recommendation statements are different. In addition, the number, type and order of the element slots contained in different recommendation statement templates are also different.

[0120] Optionally, the same component element can also contain multiple similar synonyms. For example, the elements of the event element can include: outing, play, and picnic. When generating a recommendation sentence, outing, play, and picnic can be interchanged or used alternately, making the recommendation sentence more expressive and improving the user experience.

[0121] Optionally, multiple element slots may also include modifier slots for filling in modifiers. Based on this, in addition to the element slots, the modifier slots in the same recommendation sentence template may also be interchanged with synonyms or related words to conform to daily language expression habits and to make the recommendation sentence description differentiated. For example: the recommendation sentence template is "Generate a video of a trip at %_time%". Among them, "generate" is a modifier, and the corresponding slot is a modifier slot. Taking today as an example, the recommendation sentence generated by the terminal device is: Generate a video of today's trip, make a video of today's play, or generate a video of today's outing. The position of the modifier slots in different recommendation sentence templates and the corresponding modifiers in the recommendation sentence are different.

[0122] Optionally, the configuration file may include configuration information of multiple recommendation templates, and the types of the four components of each recommendation template, the types and quantities of elements of each component, and the recommendation statement templates are partially or completely different.

[0123] Based on the meaning of the above configuration information, the following describes in detail the template parsing process and the process of matching the atlas based on the template parsing results, combined with the element slots in the recommendation statement template.

[0124] Here, we use an example where a terminal device searches for images one by one based on the four components in the configuration information of a recommended template to obtain an atlas that matches the four components. This includes the following four steps:

[0125] Step 1: Generate an initial keyword combination based on the four components of the configuration information.

[0126] For specific configuration information, see the following example:

[0127]

[0128]

[0129] Among them, the theme "topic" is the name of the theme, and each theme can correspond to one or more recommended templates. Time "_time" is the time element; location "_location" is the location element; person "_human" is the person element; event "_event" is the event element. It should be noted that one or more of these components may be empty, but these multiple components will not all be empty, and among the components that are not empty, the elements can be one or more. In the example of the above configuration information, the theme "topic" is daily Vlog, "_time" includes the five elements: today, yesterday, last weekend, this week, last week, and this month, "_human" is empty (no element, represented by []), and the element of "_event" is travel, as an example. In other words, each component can be understood as a defined range. For example, the time element includes the elements: "today, yesterday, last weekend, then it means that the range of the time element is: the range of today, yesterday and last weekend.

[0130] Taking the time element as an example, the time element and the corresponding element can be understood as a key-value pair correspondence. For example, when the time element includes the elements: today, yesterday, and last weekend, the key-value formed can include: time-today, time-yesterday, and time-last weekend.

[0131] Optionally, in the configuration information, if a component is empty, it can be represented by "[]", that is, the recommendation template corresponding to the configuration information does not include the element of the component, and the element slot corresponding to the component does not exist in the recommendation statement corresponding to the recommendation template (for example, "_location": [] and "_human": [] in the above example); if the element of a component is written by external value, for example, the result of image clustering or the result of learning user portraits needs to be used, it can be represented by "null" (for example, "title": null in the above example); if the element of a component uses the descriptive word in the configuration information, it can be represented by "%***%" (where *** represents the category of the component, such as ["%_time%"] and ["%_event%"] in the above example).

[0132] Optionally, the configuration information may also include other components, such as "_argument", which can be used as an extended component. This component can be used as an extended variable to enrich the recommendation statement template, and the specific type of the extended variable can be any one or more of the four components mentioned above. Optionally, the number of extended variables can also be one or more. For example: "_argument": [certain little daughter], "sentence": "Generate a video of %_human% and %_argument%%_time% traveling". Another example: "_argument": [whole little daughter], where whole means that the characters are an intersection, indicating that pictures of the character human and his little daughter need to be retrieved. Correspondingly, the recommendation statement template includes corresponding element slots, and the generated recommendation statement contains elements of the extended components.

[0133] Optionally, the configuration information may also include a search strategy field, such as "how_," which indicates the search strategy for images. If "how_" is "TIME_AXIS," this indicates that the image collection corresponding to the recommended template can be searched from far to near based on the time axis. The specific form of the search strategy is not limited in this embodiment of the application.

[0134] Optionally, in the configuration information, "sentence" is the recommended sentence template. The space between the two % characters corresponds to the element slot. "priority" indicates the priority of the recommended sentence; a smaller number indicates a higher priority.

[0135] Optionally, the recommendation statement templates corresponding to the same recommendation template may also be different, for example, a recommendation template: "sentence": "Generate a video of %_human% and %_argument%%_time% traveling" and "sentence": "Generate a video of %_human% and %_argument% traveling". Based on this, the terminal device can generate a group of recommendation statements for the same recommendation template. The group of recommendation statements can be one or more recommendation statements, and the number of recommendation statements corresponds to the number of recommendation statement templates. When generating the recommendation statement template, the terminal device can recommend the recommendation statement with the highest priority first according to the recommendation priority "priority" of the recommendation statement.

[0136] Optionally, the configuration information may also include a field for limiting the number of recommended statements, such as "topk". When "topk": 1, it indicates that one set of recommended statements is generated. When "topk": 3, it indicates that three sets of recommended statements are generated.

[0137] Optionally, the configuration information may also include a field for the scheduling level, which indicates the priority level of the corresponding recommendation template when parsing. For example, "level". When "level": 0, it indicates that the priority level of the corresponding recommendation template when parsing is triggered based on other recommendation opportunities, and will not be triggered based on the background rotation policy of the terminal device, that is, it will not be triggered during idle periods. When "level" is not 0, it indicates that the corresponding recommendation template is parsed based on the background rotation policy, for example, during idle periods.

[0138] Optionally, the configuration information may also include a scan interval field, such as "scan_interval_hour," which indicates the minimum interval for resolving the recommended template, in hours. For example, a value of "scan_interval_hour": 4 indicates that the recommended template will not search the gallery repeatedly within four hours, thereby updating the corresponding atlas.

[0139] Optionally, the configuration information can also include a name field, such as "title", which represents the card name of the recommended template when displayed on the desktop. If "title":null, the card name corresponding to the recommended template can be input by value.

[0140] The following is the configuration information of several recommended templates corresponding to daily Vlog themes:

[0141]

[0142]

[0143]

[0144]

[0145] The above configuration information is only an example. It should be noted that the configuration file can also include configuration information of more recommended templates, which will not be repeated here.

[0146] Optionally, the elements in the time element may also be specific dates. When the elements in the time element are specific dates, no conversion is required.

[0147] During the template parsing process, the elements of the above four components can be used as initial keywords, and then the gallery can be queried according to different initial keyword combinations to obtain the corresponding atlas. Specifically, the initial keyword combination can be obtained by combining the above four components: each element in "_time" with each element in "_location", each element in "_human" and each element in "_event". In order to facilitate the description of logic, taking the elements included in "_time": today and yesterday, the elements included in "_location": city C1 and scenic spot B2, the elements included in "_human": little daughter and lover, and the elements included in "_event": painting and dancing as an example, the generated initial keyword combinations include the following 16 types:

[0148] Today + city C1 + little daughter + dancing;

[0149] Today + city C1 + little daughter + painting;

[0150] Today + city C1 + lover + dancing;

[0151] Today + city C1 + lover + painting;

[0152] Today + Attraction B1 + Little Daughter + Dancing;

[0153] Today + Attraction B1 + Little Daughter + Drawing;

[0154] Today + Attraction B1 + Lover + Dancing;

[0155] Today + Attraction B1 + Lover + Painting;

[0156] Yesterday + City C1 + Little Daughter + Dancing;

[0157] Yesterday + City C1 + Little Daughter + Painting;

[0158] Yesterday + City C1 + Lover + Dancing;

[0159] Yesterday + City C1 + Lover + Painting;

[0160] Yesterday + Attraction B1 + Little Daughter + Dancing;

[0161] Yesterday + Attraction B1 + Little Daughter + Painting;

[0162] Yesterday + Attraction B1 + Lover + Dancing;

[0163] Yesterday + attraction B1 + lover + painting.

[0164] If any of the four elements is empty, it means that the recommendation template does not focus on the empty element. Therefore, the generated initial keyword combination does not contain the initial keyword corresponding to the empty element. For example, if the elements of "_time" include: today and yesterday, the elements of "_location" include: city C1, the elements of "_human" are empty, and the elements of "_event" are empty, the generated initial keyword combinations include the following two types:

[0165] Today + City C1;

[0166] Yesterday + City C1.

[0167] Optionally, the number of elements of the above four components can also be other numbers, and the specific composition and number of the initial keyword combination formed will vary accordingly, which cannot be exhaustively listed. As long as the same initial keyword combination does not contain two or more elements of the same component, and the same initial keyword combination contains an element of a component in which each element is not empty, no further details will be given here.

[0168] Step 2: Keyword conversion.

[0169] After obtaining the initial keyword combination, the terminal device converts the time keyword and the person keyword in the initial keyword combination. For example, a group of initial keyword combinations is: today + city C1 + little daughter + dance:

[0170] The terminal device converts the initial keyword "today" into a specific date. If today's date is February 14, 2024, the initial keyword "today" is converted to the time period from 0:00:00 to 23:59:59 on February 14, 2024, as the time search keyword. Alternatively, if the time keyword in the initial keyword combination is "yesterday" and today is February 14, 2024, the initial keyword "yesterday" is converted to the time period from 0:00:00 to 23:59:59 on February 13, 2024, as the time search keyword.

[0171] Optionally, when the time element in the time element is converted, the time element is described as a start time and an end time, and optionally, a validity period may also be included.

[0172] For example, the initial keyword "today" is converted to a start time of 0:00:00 on February 14, 2024, and an end time of 23:59:59 on February 14, 2024. Optionally, the validity period corresponding to the initial keyword "today" is 23:59:59 on February 14, 2024.

[0173] The terminal device can also convert the initial keyword "little daughter" into the TagID corresponding to the little daughter. For example, in the album, the TagID corresponding to the little daughter's face or body is TagID-1, then the initial keyword "little daughter" is converted into the corresponding TagID-1.

[0174] If the location keyword is a city, scenic spot, or other location that can obtain the corresponding longitude and latitude, the location keyword does not need to be converted. If the location keyword is a manually set reference address, the location keyword can be converted to the corresponding map address, and the location keyword corresponding to this map address can be obtained. For example, the location keyword is: My Company. After learning the user's habits, the terminal device can obtain that the user's workplace is XX Office Building. The terminal device can then convert "My Company" into the map address "XX Office Building", and then use the keyword "XX Office Building" corresponding to the map address "XX Office Building" as the converted location keyword. The XX Office Building is an identified POI area.

[0175] After the above initial keyword combination undergoes keyword conversion, the terminal device can replace the corresponding initial keyword with the converted keyword, thereby converting the initial keyword combination into the corresponding search keyword combination. For example:

[0176] The initial keyword combination: today + my company can be converted into a search keyword combination: 0:00 to 23:59 on February 13, 2024 + XX office building.

[0177] It should be noted that event keywords do not need to be converted.

[0178] When the initial keyword does not need to be converted, for example, when the location keyword is a city or a scenic spot, or when the initial keyword is an event keyword, the initial keyword can be used as a search keyword to form a search keyword combination to search the gallery.

[0179] Step 3: Search the gallery based on the converted search keyword combination to obtain the gallery corresponding to the search keyword combination.

[0180] Specifically, the terminal device can search for pictures that match the search keyword in the gallery as the picture album corresponding to the search keyword combination.

[0181] Let's continue with an example using an initial keyword combination of: today + city C1 + little daughter + dancing. After keyword conversion, the search keyword combination can be obtained as: 0:00 to 23:59 on February 14, 2024 + city C1 + TagID-2 + dancing.

[0182] The terminal device first searches the gallery to obtain all pictures whose generation time is between 0:00 and 23:59 on February 14, 2024, as the first atlas; then, searches the first atlas for all pictures with the location of city C1 as the second atlas; then, searches the second atlas for all pictures with TagID corresponding to TagID-2 as the third atlas; finally, searches the third atlas for all pictures whose semantic feature vectors match the event of drawing to obtain the fourth atlas. Based on this, the terminal device filters out the atlas corresponding to: today + city C1 + little daughter + dancing. The pictures in this atlas are all pictures of the little daughter dancing in city C1 taken between 0:00 and 23:59 on February 14, 2024. Optionally, in the process of obtaining the fourth atlas, the order of the four search components can also be adjusted. For example, the terminal device first searches for all images whose semantic feature vectors match the event "drawing", then searches for images whose generation time is between 0:00 and 23:59 on February 14, 2024, then searches for images with the little daughter, and finally searches for images with the location being city C1. The order in which the terminal device searches for the four elements is not limited in this embodiment of the application, as long as the image set that matches the four components is selected.

[0183] Optionally, when searching the gallery, the terminal device may determine whether the image satisfies the search keyword in the search keyword combination. If so, the image is added to the gallery; if not, the image does not need to be added to the gallery. The terminal device determines whether each image satisfies the search keyword in the search keyword combination, thereby generating the required gallery.

[0184] It should be noted that the specific process of searching for pictures whose semantic feature vectors match events can be: the terminal device can traverse the semantic feature vectors of all pictures that need to be searched according to the semantic feature vectors that describe the event elements, such as the semantic feature vectors of drawing. If the similarity between the semantic feature vector ClipID-1 corresponding to picture 1 and the semantic feature vector of drawing is higher than or equal to the preset similarity threshold, for example, higher than ninety percent, it means that the picture content of picture 1 most likely contains the event of drawing, and the terminal device determines that the semantic feature vectors of picture 1 and drawing match. If the similarity between the semantic feature vector ClipID-2 corresponding to picture 2 and the semantic feature vector of drawing is lower than the preset similarity threshold, it means that the picture content of picture 2 most likely does not contain the event of drawing, and it is determined that the semantic feature vectors of picture 2 and drawing do not match. Optionally, the value of the preset similarity threshold can also be other values ​​such as eighty-five percent, eighty percent, etc., which is not limited in the embodiments of the present application.

[0185] It should be noted that one or more of the four components of a recommendation template may be empty, such as the location component. If a component is empty, there is no need to search for that component. Furthermore, a component within a recommendation template may contain multiple elements, such as the person component, which includes the two elements: lover and younger daughter. When a component contains multiple elements, it is necessary to search for all of them. Alternatively, symbols can be used to distinguish between "and" or "or" relationships between multiple elements. For example, if two elements are written within the same brackets in the configuration file, for example, "human":[["I", "lover"]], this represents "I" and "lover". Therefore, the retrieved images must include both "I" and "lover". If the elements are written within two brackets in the configuration file, for example, "human":[["I"], ["lover"]], this represents "I" or "lover". Therefore, the retrieved images may include "I", "lover", or both. Optionally, other components, such as time elements, location elements, and event elements, may also be distinguished by symbols to form AND or OR relationships between multiple elements, which will not be elaborated here.

[0186] Optionally, after the terminal device searches for the atlas corresponding to the recommended template, it determines whether the number of pictures in the obtained atlas meets the quantity threshold. The quantity threshold can be a natural number, such as six. If the number of pictures in the atlas is less than the quantity threshold, for example, less than six, it means that the number of pictures in the searched atlas is small, and the effect of the generated video is not rich enough, then the atlas will no longer be recommended to the user, and the corresponding recommendation statement will not be displayed. The terminal device can then perform a search for the next search keyword combination. If the number of pictures in the atlas is greater than or equal to the quantity threshold, for example, greater than or equal to six, then the atlas can be recommended to the user, and the corresponding recommendation statement can be displayed.

[0187] Optionally, if the atlas searched by the terminal device includes multiple frames from a video, these frames can be used as images in the atlas to generate a short video. Alternatively, the video segment containing these frames can be captured, and then all the frames in this captured segment and other images in the atlas can be combined to generate a short video. It should be noted that in the generated short video, the playback order of all the frames in the captured segment remains unchanged from the original video, ensuring the playback quality of the short video.

[0188] After performing keyword conversion, the terminal device can generate a recommendation statement. The specific process for generating a recommendation statement is as follows: the terminal device can populate the initial keyword from the initial keyword combination into the recommendation statement template corresponding to the initial keyword combination, thereby generating the corresponding recommendation statement. This process can be referred to as instantiating a recommendation statement. In the configuration information, the recommendation statement template "sentence" is: "Generate a video of %_time%%_human% at %_location%%_event%." This recommendation statement template includes a time slot: time, a person slot: person, a location slot: location, and an event slot: event. Each element slot to be filled is distinguished by the symbol %; each slot is separated by two consecutive %s. For example, the terminal device populates the time slot of the recommendation statement template with the time "today," the person slot "little daughter," the location slot "attraction B," and the event slot "dancing," thereby generating the recommendation statement: "Generate a video of little daughter dancing at attraction B today." Optionally, the initial keywords in the above recommendation statement template are: %_time% %_human% %_location% %_event%.

[0189] For another example, when the recommendation sentence template "sentence" is: "Generate a video of %_time%%_event%. The recommendation sentence template includes a time slot: time and an event slot: event. Optionally, the event slot in the recommendation sentence template has been filled with "travel", then the recommendation sentence template "sentence" is: "Generate a video of %_time% travel". During the template parsing process, the terminal device can search for the search keywords %_time% and travel to obtain the corresponding atlas. At the same time, the time is filled into the time slot to generate the corresponding recommendation sentence: Generate a video of today's travel. The initial keywords in the recommendation sentence template are: %_time%%_event%.

[0190] Optionally, the specific implementation of the above search keyword combination can be implemented using the following five-layer for loop statement:

[0191]

[0192]

[0193] Based on this, the terminal device can obtain one or more search keyword combinations corresponding to a recommended template, as well as the corresponding atlas for each search keyword combination and the recommended sentence for each search keyword combination, thereby obtaining the template parsing result. The terminal device then stores the template parsing result in the recommendation database. The storage of the atlas refers to the storage of the IDs of the images in the atlas.

[0194] When there is configuration information of multiple recommended templates in the configuration file, the above steps one to three can be executed for the configuration information of each recommended template to complete the parsing of all recommended templates and form one or more recommendation statements for each recommended template and a graph set corresponding to each recommended statement.

[0195] This method of generating recommendation sentences by filling element slots based on a recommendation sentence template containing element slots is more flexible than pre-setting fixed recommendation sentences. The generated recommendation sentences conform to the user's spoken language habits, providing a better reading experience. Furthermore, the location and person elements of the aforementioned components can be clustered by images in a gallery, allowing the generated recommendation sentences to be customized for each user, achieving personalized presentation and enhancing the user experience.

[0196] In order to clearly describe the template parsing process of the embodiment of the present application, Figure 3 The interaction diagram shown is fully described. Figure 3 Shown, including:

[0197] S301. The user operates the camera application to take pictures or videos.

[0198] S302. The camera application generates a picture or video in response to the user's shooting operation, and records the generation time and location information of the picture or video.

[0199] S303. The camera application sends the picture and / or video, as well as the generation time and location information, to the gallery for storage.

[0200] S304. The image library performs CV analysis on the frame images of the images and videos during the idle period to obtain the image elements of the images.

[0201] The gallery can also cluster the generated locations in the image elements to obtain location elements.

[0202] S305. After analyzing all the pictures and frame images, the gallery sends a start instruction to the media recommendation generation engine.

[0203] Specifically, the startup instruction is used to launch the media recommendation generation engine.

[0204] S306. In response to the start instruction, the media recommendation generation engine sends an initialization instruction to the recommendation template parser.

[0205] S307. In response to the initialization instruction, the recommended template parser is initialized.

[0206] Specifically, the process of initializing the recommendation template parser involves obtaining a configuration file. Optionally, the configuration file can be in JSON format. The configuration file includes configuration information for the recommendation template. A detailed description of the configuration information can be found above and will not be repeated here.

[0207] S308. After the recommendation template parser is initialized, it returns an initialization completion message and a configuration file (or configuration information in the configuration file) to the media recommendation generation engine.

[0208] Optionally, if the recommendation template parser has not completed initialization, there is no need to send a return initialization completion message, or return an initialization incomplete message to the media recommendation generation engine.

[0209] S309. In response to the initialization completion message, the media recommendation generation engine executes a template parsing process according to the configuration file (or configuration information).

[0210] The template parsing process includes: looping through steps S310 to S317 for each recommended template's configuration information.

[0211] S310. The media recommendation generation engine splits the recommendation sentence template in the configuration information into at least one group of initial keyword combinations according to the four element slots and in combination with elements of at least one of the four components.

[0212] It should be noted that the recommendation statement template is split based on the element slots corresponding to the four components. Each initial keyword combination includes at least one initial keyword. Taking one initial keyword combination as an example, each initial keyword in the initial keyword combination corresponds to an element of one of the four components.

[0213] The terminal device performs steps S311 to S317 for each group of initial keyword groups. Taking a group of initial keyword groups as an example, the group of initial keyword groups includes at least one initial keyword.

[0214] S311. The media recommendation generation engine performs keyword conversion on at least one initial keyword to obtain at least one search keyword.

[0215] It should be noted that the initial keywords that do not need to be converted can be used as corresponding search keywords.

[0216] S312. The media recommendation generation engine sends a search instruction to the gallery, where the search instruction carries at least one search keyword.

[0217] S313. The gallery searches for pictures in the gallery according to at least one search keyword to obtain a corresponding gallery.

[0218] The image elements of the images in the collection match at least one search keyword. For example, the time when the images in the collection were generated is a time within the time period indicated by the time search keyword in at least one search keyword; for another example, the location where the images in the collection were generated is a location indicated by the location search keyword in at least one search keyword.

[0219] S314. The gallery returns the identifier of the image in the gallery as the search result to the media recommendation generation engine.

[0220] S315. The media recommendation generation engine determines whether the number of images in the collection is greater than six based on the search results.

[0221] If the number of pictures in the atlas is greater than or equal to six, execute S316; if the number of pictures in the atlas is less than six, continue to execute the template parsing process for the next recommended template.

[0222] S316. The media recommendation generation engine instantiates the recommendation statement.

[0223] Specifically, the media recommendation generation engine fills the at least one initial keyword into a corresponding element slot of the recommendation sentence template to complete the assembly of the recommendation sentence, that is, to realize the recommendation sentence instance.

[0224] S317. The media recommendation generation engine sends the recommendation sentence generated by the at least one initial keyword and the searched atlas to the recommendation database for storage.

[0225] Optionally, the atlas sent by the media recommendation generation engine may be a collection of identifiers of the images in the atlas; optionally, the media recommendation generation engine may also send identifiers of images in the atlas that serve as covers; optionally, the media recommendation generation engine may also send assembled recommendation statements.

[0226] The set of identifiers of the pictures in the above-mentioned collection, the identifier of the picture serving as the cover, and the recommendation statement may all be persistently stored in the recommendation database.

[0227] After completing the parsing of all recommendation templates, the media recommendation generation engine may execute S318 and subsequent processes.

[0228] S318. The media recommendation generation engine ends the process.

[0229] Optionally, if an abnormal situation occurs during the CV analysis, such as a sudden change from charging to off-screen status, indicating that the user may need to use the terminal device, the gallery can send an end instruction to the media recommendation generation engine. In response to the end instruction, the media recommendation generation engine ends the template parsing process, thereby freeing up resources to ensure normal user use of the terminal device.

[0230] Optionally, after completing the CV analysis, the gallery may also send an end instruction to the media recommendation generation engine to instruct the media recommendation generation engine to release resources.

[0231] above Figure 3 In the interaction diagram shown, the implementation principle and technical effect of each step can be found in the previous description and will not be repeated here.

[0232] The previous article detailed the image collection corresponding to the recommendation template and the process of obtaining the recommended sentence. Next, we will introduce the recommendation process of the recommended sentence in combination with user operations.

[0233] When the user opens the gallery, the terminal device displays the gallery homepage interface, such as Figure 4 As shown in Figure a. Figure 4 The gallery homepage, shown in Figure a, includes a card for Smart Photos. The gallery homepage also includes other cards, such as One-Click Blockbuster, Clips, and thumbnails of portraits. The bottom of the gallery homepage also displays icons for multiple tabs, including Photos, Albums, Moments, and Creations. Users can access different gallery interfaces by clicking on different tab icons.

[0234] When the user clicks on the Smart Photo function in the gallery homepage, the Smart Photo function is activated. The terminal device can display the following Figure 4 The first-level page of the smart film is shown in Figure b. The first-level page of the smart film displays multiple recommended sentences. Figure 4 In the diagram b in FIG, three recommendation statements are shown as an example. In each recommendation statement entry, a thumbnail of the corresponding cover is also displayed. Optionally, if the cover is not defined in the collection corresponding to the recommendation statement, the entry of the recommendation statement may not display any image thumbnail, and only the text of the corresponding recommendation statement may be displayed, for example Figure 5 As shown in Figure a. Alternatively, Figure 4 Figure b and Figure 5 As shown in Figure a, on the first-level page of Smart Film, a prompt can be displayed above the entry of the recommended statement.

[0235] by Figure 4Taking Figure b in the figure as an example, the first-level page of the smart film also includes a theme switch control, such as "More Skills"; it also includes a recommended sentence switch control, such as "Change". When the user clicks the recommended sentence switch control, the terminal device can switch the recommended sentence. For example, the user is in the following Figure 4 Click the "Change" control in Figure b, and the terminal device will display the following Figure 6 Recommended statement shown.

[0236] Users can click on a recommended sentence in the Smart Video page to enter the secondary page of Smart Video, such as Figure 7 As shown. The secondary page of the smart film can be displayed in the form of a dialogue, including a recommendation statement and an answer statement generated by the terminal device in response to the recommendation statement, as well as a card of the album corresponding to the recommendation statement. Optionally, the card of the album may include thumbnails of multiple pictures in the album; it may also include a control for displaying all, such as "View All"; it may also include a video generation control, such as "Generate Video". When the user clicks on the control for displaying all, the terminal device displays all or part of the pictures in the album. Optionally, the control for displaying all can also display the number of pictures in the album. When the user clicks on the video generation control, the terminal device generates a corresponding short video for the pictures in the album according to the recommendation template corresponding to the recommendation statement. The text, display effect, watermark and background music in the short video are defined by the corresponding recommendation template.

[0237] When the user clicks the theme switch control on the first-level page of Smart Film, the terminal device displays the theme selection interface, such as Figure 5 As shown in Figure b. The theme selection interface includes multiple theme cards, such as daily Vlog, personal portrait, family photo, etc. When a user clicks on any theme card, the recommended sentence corresponding to the theme will be displayed. For example, when a user clicks Figure 5 When a card with the theme of personal photo is displayed in Figure b, the terminal device displays multiple recommended sentences corresponding to the theme of personal photo, such as Figure 7 As shown. Users can Figure 7 Select the desired recommendation statement on the secondary page shown, and a corresponding short video will be generated. The terminal device can play the short video, store it, and forward it. This enriches the way to browse and share pictures, and improves the user experience.

[0238] Optionally, the way to display the recommended statements in the above-mentioned secondary page can also be to display them according to the priority of the recommended statements. For example, when the secondary page is displayed for the first time, the three recommended statements with the highest priority are displayed. When the user clicks "Change", three recommended statements with a lower priority than the third recommended statement are displayed. Optionally, the priority of the recommended statement can be displayed based on the number of pictures in the corresponding collection. The more pictures there are in the collection, the higher the priority of the corresponding recommended statement; the fewer pictures there are in the collection, the lower the priority of the corresponding recommended statement. Optionally, the priority of the recommended statement can also be determined based on the update time of the recommendation template corresponding to the recommended statement. The more recent the update time of the recommended template, the higher the priority of the corresponding recommended statement; the further the update time of the recommended template, the lower the priority of the corresponding recommended statement. Optionally, if a recommended statement exists in the history of the recommended statement, it means that the recommended statement has been recommended to the user, then the priority of the recommended statement is reduced, and the priority of other recommended statements that have not been recommended is increased. This can improve the differentiation of the presentation of the recommended statements, make the generated short videos differentiated, and make the user experience richer.

[0239] In an embodiment of the present application, when the user starts the smart video function, the terminal device can automatically display the recommended sentence and guide the user to generate a short video corresponding to the recommended sentence. Optionally, the user can also input a similar sentence by voice or text according to the guidance of the recommended sentence. The terminal device can then perform semantic recognition on the sentence input by the user, obtain the initial keyword combination corresponding to the sentence, and then perform keyword conversion on the initial keywords in the initial keyword combination to obtain the corresponding search keyword combination. The terminal device can search in the gallery based on the search keyword combination, obtain the atlas corresponding to the sentence input by the user, and display it. The display method can be referred to as Figure 6 The dialog form shown here shows the recommended sentence entered by the user. The specific description of the acquisition of the initial keyword combination, keyword conversion, and searching in the gallery based on the search keyword combination in this embodiment can be found in the relevant description above and will not be repeated here.

[0240] Optionally, the above CV analysis process can also be executed after the user clicks on the smart film. This embodiment of the present application does not limit this.

[0241] Figure 8 A method for obtaining an atlas provided in an embodiment of the present application includes:

[0242] S801: Acquire at least one component element of a target recommendation template, where the at least one component element includes at least a time element and / or a person element.

[0243] The constituent elements may be any one or more of a time element, a location element, a person element, and an event element. The constituent elements may be described using configuration information in a configuration file. For details about the configuration file and configuration information, please refer to the previous description and will not be repeated here.

[0244] S802: Convert the initial keywords of the time element and / or the person element to obtain a time search keyword and / or a person search keyword.

[0245] The above-mentioned initial keywords are descriptive words that conform to the user's spoken language. The elements of the time factor (initial keywords) can be descriptive words such as today, tomorrow, last week, etc. that conform to the user's spoken expression. The elements of the person factor (initial keywords) can be descriptive words such as Dabao, lover, etc. that conform to the user's spoken expression. After the terminal device converts the initial keywords, it obtains the corresponding time search keywords and / or person search keywords. The time search keyword is a specific date, and the person search keyword is the TagID or ReID of the corresponding person.

[0246] S803. Search the gallery according to the search keyword to obtain a target gallery that matches the search keyword, where the search keyword includes at least a time search keyword and / or a person search keyword; wherein the first picture is any picture in the target gallery, the picture element of the first picture matches at least one component element, and the picture element of the first picture includes at least the generation time of the first picture and / or the person identification.

[0247] The terminal device can search the gallery based on some or all of the time search keywords, person search keywords, location search keywords, and event search keywords to obtain the target gallery that matches the search keywords. The detailed description of the search gallery can be found in the previous section and will not be repeated here.

[0248] Figure 8 In the illustrated embodiment, the terminal device can use the corresponding atlas obtained from the image elements of the image to recommend to the user, thereby personalizing the display of the image according to the recommendation template, which can enhance the user's browsing experience. In addition, by converting the colloquial initial keywords into corresponding search keywords that the terminal device can recognize for searching, a gallery search can be implemented. At the same time, the recommendation sentences generated later by the colloquial initial keywords conform to the user's colloquial expression habits, providing a better reading experience.

[0249] In some embodiments, at least one component element includes: a first component element and other components, the first component element is any one of a time element, a person element, a place element and an event element, the other component elements include one or more of a time element, a person element, a place element and an event element, and any one of the first component element and the other component elements are different; when the first component element includes at least two initial keywords, the method further includes: combining each of the at least two initial keywords with the initial keywords of the other components to form at least two initial keyword combinations, wherein the first initial keyword combination is any one of the at least two initial keyword combinations, the first initial keyword combination includes multiple initial keywords, the multiple initial keywords include the first initial keyword and the second initial keyword, the component element corresponding to the first initial keyword and the component element corresponding to the second initial keyword are different component elements in at least one component, and the multiple initial keywords include at least two initial keywords. Any one; converting the initial keywords of the time element and / or the person element to obtain the time search keyword and / or the person search keyword, including: when the multiple initial keywords in the first initial keyword combination include the time keyword, converting the time keyword to obtain the time search keyword corresponding to the time keyword; when the multiple initial keywords in the first initial keyword combination include the person keyword, converting the person keyword to obtain the person search keyword corresponding to the person keyword; searching in the gallery according to the search keyword to obtain the target gallery that matches the search keyword, including: replacing the corresponding time keyword in the first initial keyword combination with the time search keyword, and / or replacing the corresponding person keyword in the first initial keyword combination with the person search keyword to obtain the first search keyword combination corresponding to the first initial keyword combination, the first search keyword including at least one search keyword; searching in the gallery according to at least one search keyword to obtain the target gallery that matches at least one search keyword.

[0250] The above-mentioned first component is any one of at least one component, and the other components are different from the first component. If the elements in each component can be empty, it can also be one or more. The elements of all components in a recommendation template will not all be empty. The terminal device can first combine different elements in different components to form all combinations of different elements in multiple components, thereby obtaining at least two groups of initial keyword combinations. Taking the first initial keyword combination as an example, the initial keywords therein are converted into keywords to obtain corresponding search keywords.

[0251] It should be noted that the above-mentioned person search keywords may be person feature vectors used to search for people, and the above-mentioned event search keywords may be semantic feature vectors used to search for events.

[0252] The terminal device searches the image library based on the converted search keywords, along with other search keywords corresponding to the same initial keyword combination that do not require conversion, to obtain the corresponding target atlas. Optionally, one or more search keywords in the search keyword combination may be empty. The specific form of the search keyword combination is not limited here.

[0253] Based on this, the terminal device can generate atlases corresponding to a variety of different elements, making the presented atlases diversified and enriching the user experience.

[0254] In some embodiments, at least one component element further includes: a location element, at least one search keyword further includes a location search keyword, and the image element of the first image further includes the location where the first image was generated.

[0255] In some embodiments, the location search keyword is one of a plurality of candidate locations, and the plurality of candidate locations are obtained by clustering the generation locations of a plurality of pictures in the gallery.

[0256] In some embodiments, when the place element includes a first reference address, the method further includes: converting the first reference address into a first map address, and the place search keyword includes a place name corresponding to the first map address.

[0257] In some embodiments, at least one component element also includes: a character element, at least one search keyword also includes a character search keyword, and the picture element of the first picture also includes a character feature vector of the first picture, and the character feature vector of the first picture is used to represent the character appearing in the first picture.

[0258] In some embodiments, the character feature vector of the first image is one or more of a plurality of character feature vectors, and the plurality of character feature vectors are obtained by clustering faces and / or bodies of characters appearing in a plurality of images in the gallery.

[0259] By clustering the images in the gallery, we can obtain the character feature vectors and / or location elements, effectively narrowing the search scope and saving search time and resources. At the same time, we can also achieve personalized recommendations, meet user expectations, and improve the user experience.

[0260] In some embodiments, at least one component element also includes: an event element, at least one search keyword also includes an event search keyword, the picture element of the first picture also includes a semantic feature vector of the first picture, the semantic feature vector of the first picture is used to represent the picture content of the first picture, and the picture content of the first picture includes the event corresponding to the event search keyword.

[0261] For detailed descriptions of the components and image elements, please refer to the previous article and will not be repeated here.

[0262] In some embodiments, the method also includes: obtaining a first recommendation statement template of the target recommendation template, the first recommendation statement template includes at least one element slot, and the at least one element slot corresponds one-to-one to at least one component element; filling the first initial keyword of the first component element corresponding to the first element slot in the first element slot, generating a first recommendation statement corresponding to the first recommendation statement template, the first element slot is any one of the at least one element slot, the first component element is one of the at least one component element corresponding to the first element slot, and the first initial keyword is one of the elements of the first component element.

[0263] The terminal device populates the initial keywords into the corresponding element slots in the recommendation sentence template, generating the corresponding recommendation sentence. A detailed description of this can also be found in the relevant descriptions in the article and will not be repeated here. The initial keywords in this recommendation sentence align with the user's colloquial expression habits, making it easier to read. Furthermore, filling in the sentence based on the element slots provides high flexibility and improves the user experience.

[0264] In some embodiments, the method also includes: receiving a first operation input by the user, the first operation is used to display the interface of the gallery application; in response to the first operation, displaying the first interface of the gallery application, the first interface including a first control; receiving a second operation performed by the user on the first control; in response to the second operation, displaying the second interface, the first recommendation statement is displayed in the second interface; receiving a third operation performed by the user on the second control of the first recommendation statement; in response to the third operation, displaying the third interface, the target gallery is displayed in the third interface.

[0265] The first operation can be an operation for the user to open the gallery application. The terminal device responds to the user's first operation and displays the first interface of the gallery, such as Figure 4 The first interface includes a first control, such as Figure 4 The user can perform a second operation on the first control, such as clicking the smart photo card. The terminal device then activates the smart photo function and displays a second interface including one or more recommendation statements.

[0266] The user performs a third operation on the control of one of the recommendation statements in the second interface, for example, clicking the control of the first recommendation statement (generating a video of Nanjing travel): the terminal device displays the third interface, which includes a thumbnail of the target atlas corresponding to the first recommendation statement, for example Figure 6 As shown in the figure.

[0267] In some embodiments, the third interface also includes a third control, and the method further includes: receiving a fourth operation input by the user for the third control; in response to the fourth operation, generating a video file containing a target atlas according to the target recommendation template; and playing and / or storing the video file.

[0268] The third interface may also include a third control, such as Figure 6 The user performs a fourth operation on the third control, such as clicking "Generate Video", and the terminal device generates a short video from the images in the target collection according to the target recommendation template, achieving personalized display, enriching the image display method, and improving the user experience.

[0269] In some embodiments, before searching the gallery based on the search keyword to obtain the target gallery that matches the search keyword, the process also includes: analyzing the images in the gallery to obtain image elements of each image, where the image elements include at least the generation time.

[0270] In some embodiments, the picture element further includes: a location where the picture is generated, where the location is any one or more of a city, a POI area, and a custom location.

[0271] In some embodiments, the location where the image is generated is converted based on the longitude and latitude at which the image is generated.

[0272] In some embodiments, the picture elements further include: a character feature vector and / or a semantic feature vector of the picture.

[0273] The image analysis and clustering methods can be found in the previous article and will not be repeated here. The above image elements can describe the attributes of the image and facilitate classification.

[0274] The above describes in detail an example of the method provided by the present application. It is understandable that, in order to implement the above functions, the corresponding device includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0275] The present application can divide the atlas acquisition device into functional modules according to the above method examples. For example, each function can be divided into various functional modules, or two or more functions can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in this application is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0276] Figure 9 The schematic diagram of the structure of a device for obtaining an atlas provided by the present application is shown. The device 900 includes:

[0277] The acquisition module 901 is configured to acquire at least one component of a target recommendation template, wherein the at least one component includes at least a time component and / or a person component.

[0278] The conversion module 902 is used to convert the initial keywords of the time element and / or the person element to obtain the time search keyword and / or the person search keyword.

[0279] The search module 903 is used to search in the gallery according to the search keywords to obtain a target gallery that matches the search keywords, and the search keywords include at least a time search keyword and / or a person search keyword; wherein the first picture is any picture in the target gallery, the picture elements of the first picture match at least one component element, and the picture elements of the first picture include at least the generation time of the first picture and / or the person identification.

[0280] In some embodiments, at least one component element includes: a first component element and other components, the first component element is any one of a time element, a person element, a location element, and an event element, the other components include one or more of a time element, a person element, a location element, and an event element, and any one of the first component element and the other components are different;

[0281] When the first component includes at least two initial keywords, the conversion module 902 is specifically configured to combine each of the at least two initial keywords with the initial keywords of the other component to form at least two initial keyword combinations, wherein the first initial keyword combination is any one of the at least two initial keyword combinations, the first initial keyword combination includes multiple initial keywords, the multiple initial keywords include a first initial keyword and a second initial keyword, the component corresponding to the first initial keyword and the component corresponding to the second initial keyword are different components of at least one component, and the multiple initial keywords include any one of the at least two initial keywords; and when the multiple initial keywords in the first initial keyword combination include a time keyword, converting the time keyword to obtain a time search keyword corresponding to the time keyword; when the multiple initial keywords in the first initial keyword combination include a person keyword, converting the person keyword to obtain a person search keyword corresponding to the person keyword;

[0282] The search module 903 is specifically used to replace the corresponding time keyword in the first initial keyword combination with a time search keyword, and / or replace the corresponding person keyword in the first initial keyword combination with a person search keyword, to obtain a first search keyword combination corresponding to the first initial keyword combination, where the first search keyword includes at least one search keyword; and to search the gallery according to the at least one search keyword to obtain a target gallery that matches the at least one search keyword.

[0283] In some embodiments, at least one component element further includes: a location element, at least one search keyword further includes a location search keyword, and the image element of the first image further includes the location where the first image was generated.

[0284] In some embodiments, the location search keyword is one of a plurality of candidate locations, and the plurality of candidate locations are obtained by clustering the generation locations of a plurality of pictures in the gallery.

[0285] In some embodiments, when the place element includes a first reference address, the conversion module 902 is further configured to convert the first reference address into a first map address, and the place search keyword includes a place name corresponding to the first map address.

[0286] In some embodiments, at least one component element also includes: a character element, at least one search keyword also includes a character search keyword, and the picture element of the first picture also includes a character feature vector of the first picture, and the character feature vector of the first picture is used to represent the character appearing in the first picture.

[0287] In some embodiments, the character feature vector of the first image is one or more of a plurality of character feature vectors, and the plurality of character feature vectors are obtained by clustering faces and / or bodies of characters appearing in a plurality of images in the gallery.

[0288] In some embodiments, at least one component element also includes: an event element, at least one search keyword also includes an event search keyword, the picture element of the first picture also includes a semantic feature vector of the first picture, the semantic feature vector of the first picture is used to represent the picture content of the first picture, and the picture content of the first picture includes the event corresponding to the event search keyword.

[0289] In some embodiments, the device 900 also includes: a recommendation module, which is used to obtain a first recommendation statement template of the target recommendation template, the first recommendation statement template includes at least one element slot, and the at least one element slot corresponds one-to-one to at least one component element; fill the first initial keyword of the first component element corresponding to the first element slot in the first element slot, and generate a first recommendation statement corresponding to the first recommendation statement template, the first element slot is any one of the at least one element slot, the first component element is one of the at least one component element corresponding to the first element slot, and the first initial keyword is one of the elements of the first component element.

[0290] In some embodiments, the apparatus 900 further includes: a receiving module configured to receive a first operation input by a user, where the first operation is configured to display an interface of a gallery application.

[0291] The display module is configured to display a first interface of the gallery application in response to a first operation, where the first interface includes a first control.

[0292] The receiving module is further configured to receive a second operation performed by the user on the first control.

[0293] The display module is further configured to display a second interface in response to a second operation, wherein the first recommendation statement is displayed in the second interface.

[0294] The receiving module is further configured to receive a third operation performed by the user on the second control of the first recommendation statement.

[0295] The display module is further configured to display a third interface in response to a third operation, wherein the target atlas is displayed in the third interface.

[0296] In some embodiments, the third interface also includes a third control, and the receiving module is further used to receive a fourth operation input by the user for the third control.

[0297] The display module is further configured to generate a video file containing a target atlas according to the target recommendation template in response to a fourth operation, and to play and / or store the video file.

[0298] In some embodiments, the apparatus 900 further includes: a picture analysis module configured to analyze the pictures in the picture library to obtain picture elements of each picture, where the picture elements at least include a generation time.

[0299] In some embodiments, the picture element further includes: a location where the picture is generated, where the generation location is any one or more of a city, a POI area, and a custom location.

[0300] In some embodiments, the location where the image is generated is converted based on the longitude and latitude at which the image is generated.

[0301] In some embodiments, the picture elements further include: a character feature vector and / or a semantic feature vector of the picture.

[0302] The specific manner in which the apparatus 900 executes the atlas acquisition method and the beneficial effects produced can be found in the relevant description in the method embodiment, which will not be repeated here.

[0303] The embodiment of the present application also provides an electronic device, including the above-mentioned processor. The electronic device provided by this embodiment can be Figure 1 The terminal device 100 shown is used to perform the above-mentioned atlas acquisition method. When integrated, the terminal device may include a processing module, a storage module, and a communication module. The processing module may be used to control and manage the terminal device's operations. For example, it may be used to support the terminal device in executing the steps performed by the display unit, detection unit, and processing unit. The storage module may be used to support the terminal device in executing and storing program code and data. The communication module may be used to support communication between the terminal device and other devices.

[0304] The processing module may be a processor or a controller. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor (DSP) and a microprocessor, and so on. The storage module may be a memory. The communication module may specifically be a device that interacts with other terminal devices, such as a radio frequency circuit, a Bluetooth chip, or a Wi-Fi chip.

[0305] In one embodiment, when the processing module is a processor and the storage module is a memory, the terminal device involved in this embodiment may be a Figure 1 Device with the structure shown.

[0306] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the processor executes the atlas acquisition method described in any of the above embodiments.

[0307] An embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the above-mentioned related steps to implement the atlas acquisition method in the above-mentioned embodiment.

[0308] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.

[0309] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, the replaced units may or may not be physically separated, and the components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed in multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0310] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0311] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0312] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for acquiring an atlas, characterized in that: include: Acquire at least one component element of a target recommendation template, wherein the at least one component element includes at least a time element and / or a person element; Converting the initial keywords of the time element and / or person element to obtain a time search keyword and / or a person search keyword; Searching the gallery according to the search keyword to obtain a target gallery that matches the search keyword, wherein the search keyword includes at least the time search keyword and / or the person search keyword; Among them, the first picture is any picture in the target picture set, the picture elements of the first picture match the at least one component element, and the picture elements of the first picture at least include the generation time of the first picture and / or the character identification.

2. The method according to claim 1, characterized in that The at least one component element includes: a first component element and other components, the first component element is any one of the time element, the person element, the location element, and the event element, the other components include one or more of the time element, the person element, the location element, and the event element, and the first component element and any one of the other components are different; When the first component includes at least two initial keywords, the method further includes: Combining each of the at least two initial keywords with the initial keywords of the other components to form at least two initial keyword combinations, wherein a first initial keyword combination is any one of the at least two initial keyword combinations, the first initial keyword combination includes a plurality of initial keywords, the plurality of initial keywords include a first initial keyword and a second initial keyword, the component corresponding to the first initial keyword and the component corresponding to the second initial keyword are different components of the at least one component, and the plurality of initial keywords include any one of the at least two initial keywords; The converting of the initial keywords of the time element and / or person element to obtain the time search keyword and / or person search keyword includes: When the multiple initial keywords in the first initial keyword combination include a time keyword, converting the time keyword to obtain the time search keyword corresponding to the time keyword; When the multiple initial keywords in the first initial keyword combination include a character keyword, converting the character keyword to obtain the character search keyword corresponding to the character keyword; The step of searching the gallery based on the search keyword to obtain a target gallery that matches the search keyword includes: Replacing the corresponding time keyword in the first initial keyword combination with the time search keyword, and / or replacing the corresponding person keyword in the first initial keyword combination with the person search keyword, to obtain a first search keyword combination corresponding to the first initial keyword combination, wherein the first search keyword includes at least one search keyword; A search is performed in the gallery according to the at least one search keyword to obtain the target gallery that matches the at least one search keyword.

3. The method according to claim 2, characterized in that The at least one component element further includes: a location element, the at least one search keyword further includes a location search keyword, and the image element of the first image further includes a location where the first image is generated.

4. The method according to claim 3, characterized in that The location search keyword is one of a plurality of candidate locations, and the plurality of candidate locations are obtained by clustering the generation locations of a plurality of pictures in the gallery.

5. The method according to claim 3 or 4, characterized in that When the place element includes a first reference address, the method further includes: The first reference address is converted into a first map address, and the location search keyword includes a location name corresponding to the first map address.

6. The method according to any one of claims 2 to 5, characterized in that The at least one component element also includes: a character element, the at least one search keyword also includes a character search keyword, the picture element of the first picture also includes a character feature vector of the first picture, and the character feature vector of the first picture is used to represent the character appearing in the first picture.

7. The method according to claim 6, characterized in that The character feature vector of the first picture is one or more of a plurality of character feature vectors, and the plurality of character feature vectors are obtained by clustering faces and / or bodies of characters appearing in a plurality of pictures in the gallery.

8. The method according to any one of claims 2 to 7, characterized in that The at least one component element also includes: an event element, the at least one search keyword also includes an event search keyword, the picture element of the first picture also includes a semantic feature vector of the first picture, the semantic feature vector of the first picture is used to represent the picture content of the first picture, and the picture content of the first picture includes the event corresponding to the event search keyword.

9. The method according to claim 2, characterized in that The method further comprises: Acquire a first recommendation sentence template of the target recommendation template, wherein the first recommendation sentence template includes at least one element slot, and the at least one element slot corresponds one-to-one to the at least one component element; Fill the first element slot with the first initial keyword of the first component corresponding to the first element slot to generate the first recommendation statement corresponding to the first recommendation statement template, the first element slot is any one of the at least one element slot, the first component is one of the at least one component corresponding to the first element slot, and the first initial keyword is one of the elements of the first component.

10. The method according to claim 9, characterized in that The method further comprises: Receive a first operation input by a user, where the first operation is used to display an interface of a gallery application; In response to the first operation, displaying a first interface of the gallery application, wherein the first interface includes a first control; receiving a second operation performed by a user on the first control; In response to the second operation, displaying a second interface, wherein the first recommendation statement is displayed in the second interface; receiving a third operation performed by a user on a second control of the first recommendation statement; In response to the third operation, a third interface is displayed, in which the target atlas is displayed.

11. The method according to claim 10, characterized in that The third interface also includes a third control, and the method further includes: receiving a fourth operation input by the user for the third control; In response to the fourth operation, generating a video file including the target atlas according to the target recommendation template; Play and / or store the video file.

12. The method according to claim 1, characterized in that Before searching the gallery based on the search keyword to obtain a target gallery that matches the search keyword, the method further includes: The pictures in the picture library are analyzed to obtain picture elements of each picture, where the picture elements at least include generation time.

13. The method according to claim 12, characterized in that The picture element also includes: a location where the picture is generated, where the location is any one or more of a city, a POI area, and a custom location.

14. The method according to claim 13, characterized in that The location where the image is generated is converted based on the longitude and latitude at which the image is generated.

15. The method according to claim 12, characterized in that The picture elements also include: a character feature vector and / or a semantic feature vector of the picture.

16. An electronic device, characterized in that: include: processors, memory, and interfaces; The processor, the memory, and the interface cooperate with each other so that the electronic device executes the method according to any one of claims 1 to 15.

17. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Picture searching method and device, terminal and storage medium

    CN111104536A

  • Search processing method and device, electronic equipment and storage medium

    CN113111248A

  • Retrieval method for information in table picture, electronic equipment and storage medium

    CN115408497A

  • Image retrieval method and related equipment

    CN116431855A

  • Method and system for searching files based on natural language and computing equipment

    CN117235014A