Method and device for generating video album

By filtering and aggregating multimedia resources based on the user's personal attribute tags, the terminal generates video albums related to the topic, solving the problem that video albums have little impact on users' emotions in the prior art, and improving user experience and diversity.

CN120050484APending Publication Date: 2025-05-27HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311591134.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The video albums generated by the prior art have little impact on users' emotions, users are not willing to browse, and their user experience is not high.

Method used

By receiving the user's operation of opening the video album in the gallery, the terminal displays the multimedia resources in the video album in the form of a video, filters and aggregates the multimedia resources in the gallery based on the topics related to the user's personal attribute tags, and generates summary and recallable video albums related to the topic.

Benefits of technology

It improves the impact of video albums on users' emotions, stimulates users' browsing awareness, improves user experience, avoids duplication of video albums, and increases diversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050484A_ABST
    Figure CN120050484A_ABST
Patent Text Reader

Abstract

The invention provides a method and equipment for generating a video album, and relates to the technical field of terminals. The problems that the impact of the generated video album on the emotion of the user is not strong, the browsing willingness of the user is not strong, and the use experience is not high are solved. The terminal receives a first operation of opening a video album in a picture library by a user; in response, the terminal displays a plurality of multimedia resources in the video album in a video form; the multimedia resources comprise pictures and / or videos; the multimedia resources for generating the video album are obtained by screening from a picture library of the terminal based on the target theme, the target theme is related to user information of a terminal user, and the user information comprises one or more of feature information of the user, relation information of the user or record information generated when the user uses an application in the terminal. The feature information is used for representing behavior preferences of the user; the relation information is used for representing the social relation of the user; the record information comprises at least one of the following data: social media data, browsing data or shopping data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of terminals, and in particular, to a method and device for generating a video album. Background Art

[0002] In daily life, people often use their mobile phones to take some photos and videos and store them in the mobile phone's gallery. Currently, the mobile phone can automatically screen and aggregate the excellent photos and videos in the mobile phone's gallery to generate a recollective and summary video (or called a video album), such as Huawei's wonderful moments, which is convenient for users to view. However, in this function, currently, only based on dimensions such as the time and location of photo or video shooting, and preset themes, the screening of photos and videos in the mobile phone's gallery is realized, and relevant videos are generated using these as materials. Since the conditions for screening video materials are relatively fixed and single, the generated video album has a weak impact on the user's emotions, resulting in a weak browsing willingness of users and a low experience. Summary of the Invention

[0003] This application provides a method and device for generating a video album, which solves the problems that the generated video album has a weak impact on the user's emotions, a weak browsing willingness of users, and a low usage experience.

[0004] In the first aspect of this application, a method for generating a video album is provided. The method may include:

[0005] The terminal can receive a first operation by the user to open the video album in the gallery; in response to this first operation, the terminal displays multiple multimedia resources in the video album in the form of a video.

[0006] Among them, the video album is generated based on multiple multimedia resources, and the multimedia resources include pictures and / or videos; moreover, the multimedia resources for generating the video album are screened from the terminal's gallery based on a target theme, and the target theme is related to the user information of the terminal user. The user information includes one or more of the user's characteristic information, the user's relationship information, or the record information generated by the user using the applications in the terminal. The characteristic information is used to characterize the user's behavior preferences; the relationship information is used to characterize the user's social relationships; the record information includes at least one of the following data: social media data, browsing data, or shopping data.

[0007] The method for generating a video album provided by the embodiments of the present application is such that the terminal filters and aggregates the multimedia resources in the terminal gallery based on the theme related to the user's personal attribute tags, and then automatically generates a summary and recollective video album related to the above theme. In this way, the generated video album is related to the user's personal attribute tags, which can allow the user to have a more intuitive emotional resonance, thereby stimulating the user's browsing awareness and improving the experience. In addition, since different user personal attribute tags correspond to different themes, the generated related video albums are also different, and are no longer restricted by the preset themes, thus enhancing the diversity of the generated video albums, avoiding the repetition of video albums, and further improving the user's usage experience.

[0008] In combination with the first aspect, in a possible implementation manner, the terminal displays multiple multimedia resources in the video album in video form, which specifically may include: displaying the multiple multimedia resources in the video album in video form based on the target video style and target video materials; wherein, the target video style and target video materials are also related to the user information; the target video style includes one or more of: filters, color grading, special effects, or transition animations; the target video materials include one or more of: titles, subtitles, or background music. The video style, video materials, etc. of the video album are also related to the user's personal attribute tags, further enhancing the impact of the video album on the user's emotions and improving the experience.

[0009] In combination with the first aspect, in another possible implementation manner, the above method may further include: obtaining the user information; generating a target theme, target video style, and target video materials based on the user information; filtering multimedia resources from the terminal's gallery based on the target theme; generating a video album based on the filtered multimedia resources, target video style, and target video materials. In this way, the terminal can generate the theme, video style, and video materials of the relevant video album according to the user information, no longer restricted by the preset dimensions, enhancing the diversity of video album generation, and improving the user's usage experience. Among them, this process can be executed by the terminal or by the server.

[0010] In combination with the first aspect, in yet another possible implementation manner, after obtaining the user information, the above method may further include: determining a task for triggering the generation of a video album based on the user information. In this way, the diversity of triggering the generation of a video album is improved, and the adhesion between the video album and the user's personal attributes is deepened, improving the user's usage experience. In addition, when the user information meets the conditions, such as through emotional polarity analysis and user portrait construction, it is decided whether to trigger the task of generating a video album, improving the flexibility of the timing of generating a video album. And triggering the task of generating a video album only at an appropriate time can also reduce the device power consumption. Among them, this process can be executed by the terminal or by the server.

[0011] In combination with the first aspect, in another possible implementation, a target theme, a target video style, and target video materials are generated based on user information. Specifically, it may include: extracting features from the user information to obtain a feature vector of the user information. The feature vector is used to represent the features of the user information and hides the private content in the user information; using the feature vector of the user information as input, through an artificial intelligence (AI) model, generating a target theme, a target video style, and target video materials; where the AI model has the function of generating corresponding themes, video styles, and video materials based on user information. In this way, on the basis of protecting the security of user information, a large amount of data can be processed through the AI model, which can reduce labor costs and improve work efficiency. In addition, using Artificial Intelligence Generated Content (AIGC) to generate personalized background music, titles, captions, filters, etc. in combination with user preferences can further enhance the user experience. Among them, this process can be executed by the terminal or by the server.

[0012] In combination with the first aspect, in another possible implementation, determining a task to trigger the generation of a video album based on user information may specifically include: extracting features from the user information to obtain a feature vector of the user information. The feature vector is used to represent the features of the user information and hides the private content in the user information; using the feature vector of the user information as input, through the AI model, outputting a task to trigger the generation of a video album; where the AI model has the function of determining whether to trigger the task of generating a video album based on user information. In this way, on the basis of protecting the security of user information, a large amount of tasks can be processed through the AI model, which can reduce labor costs and improve work efficiency. Using the AI model to decide whether to trigger the task of generating a video album through sentiment polarity analysis and user portrait construction can give a richer timing for pushing video albums. Among them, this process can be executed by the terminal or by the server.

[0013] In combination with the first aspect, in another possible implementation, before screening multimedia resources from the terminal's picture library based on the target theme to generate a video album, the above method may further include: correcting the target theme, the target video style, and the target video materials according to a preset rule. The preset rule is used to correct the emotional tendency of the theme, the video style, and the video materials of the video album. In this way, it can effectively avoid negative results with bad emotional tendencies caused by insufficient discrimination ability of the AI model. Among them, this process can be executed by the terminal or by the server.

[0014] In connection with the first aspect, in another possible implementation, obtaining user information may specifically include: obtaining a user portrait and a personal knowledge graph of the user; determining the characteristic information of the user based on the user portrait, and determining the relationship information of the user based on the personal knowledge graph. In this way, by obtaining user information related to the end user from multiple aspects, the user can be understood more accurately, improving the user experience. In addition, the material screening method based on user emotions and personal knowledge graphs is not restricted by the pre-set theme range, enhancing the diversity of video album generation.

[0015] In a second aspect of the present application, a method for generating a video album is provided, which is applied to a terminal. The method may include:

[0016] The terminal may receive a first operation by the user to open the video album in the gallery; in response to this first operation, the terminal displays multiple multimedia resources in the video album in the form of a video.

[0017] Among them, the video album is generated based on multiple multimedia resources, and the multimedia resources include pictures and / or videos; moreover, the multimedia resources for generating the video album are screened from the gallery of the terminal based on a target theme, and the target theme is related to the user information of the end user. The user information includes one or more of the characteristic information of the user, the relationship information of the user, or the record information generated by the user using the applications in the terminal. The characteristic information is used to characterize the user's behavior preferences; the relationship information is used to characterize the user's social relationships; the record information includes at least one of the following data: social media data, browsing data, or shopping data.

[0018] In connection with the second aspect, in a possible implementation, the terminal displays multiple multimedia resources in the video album in the form of a video, which may specifically include: displaying the multiple multimedia resources in the video album in the form of a video based on a target video style and target video materials; among them, the target video style and target video materials are also related to the user information; the target video style includes one or more of: filters, color grading, special effects, or transition animations; the target video materials include one or more of: titles, subtitles, or background music.

[0019] In combination with the second aspect, in another possible implementation, the above method may further include: The terminal obtains user information, extracts features from the user information to obtain a feature vector of the user information. The feature vector is used to characterize the features of the user information and hides the private content in the user information; The terminal sends the feature vector to the server; The terminal receives a target theme, a target video style, and target video materials from the server; The terminal filters multimedia resources from the terminal's picture library based on the target theme; A video album is generated based on the filtered multimedia resources, the target video style, and the target video materials. One or more of the target video style and the target video materials may not be received from the server and may be pre-configured in the terminal or randomly determined by the terminal.

[0020] In combination with the second aspect, in yet another possible implementation, obtaining user information may specifically include: Obtaining the user's user portrait and personal knowledge graph; Determining the user's feature information based on the user portrait and determining the user's relationship information based on the personal knowledge graph.

[0021] In a third aspect of the present application, a method for generating a video album is provided, which is applied to a server. The method may include:

[0022] The server receives a feature vector of user information from the terminal. The feature vector is used to characterize the features of the user information and hides the private content in the user information. The server determines a target theme based on the feature vector and sends the target theme to the terminal; The target theme is used for the terminal to filter multimedia resources from the terminal's picture library to generate a video album. The user information includes one or more of the user's feature information, the user's relationship information, or record information generated by the user using an application in the terminal. The feature information is used to characterize the user's behavior preferences; The relationship information is used to characterize the user's social relationships; The record information includes at least one of the following data: social media data, browsing data, or shopping data.

[0023] In combination with the third aspect, in a possible implementation, the method may further include: The server determines at least one of a target video style and target video materials based on the feature vector and sends it to the terminal.

[0024] In combination with the third aspect, in another possible implementation, before determining the target theme based on the feature vector, the above method may further include: Determining a task that triggers the generation of a video album based on the feature vector.

[0025] In connection with the third aspect, in another possible implementation, a target theme, a target video style, and target video materials are generated based on the feature vector. Specifically, it may include: using the feature vector as an input, and through an AI model, generating the target theme, the target video style, and the target video materials; wherein, the AI model has the function of generating corresponding themes, video styles, and video materials based on user information.

[0026] In connection with the third aspect, in another possible implementation, a task of triggering the generation of a video album is determined based on the feature vector. Specifically, it may include: using the feature vector as an input, and through an AI model, outputting the task of triggering the generation of the video album; wherein, the AI model has the function of determining whether to trigger the task of generating the video album based on user information.

[0027] In connection with the third aspect, in another possible implementation, before sending the target theme, the target video style, and the target video materials to the terminal, the above method may further include: correcting the target theme, the target video style, and the target video materials according to a preset rule, and the preset rule is used to correct the emotional tendencies of the theme, the video style, and the video materials of the video album.

[0028] Fourth aspect, a device is provided, and the device has the function of implementing the device behaviors in the methods described in the above first aspect or second aspect or third aspect. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example, an input unit or module, a display unit or module, a processing unit or module. This device can also be called an electronic device, such as a terminal or a server.

[0029] Fifth aspect, a device is provided, and the device includes: a processor; a memory; and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the device executes the method described in any one of the above first aspect or second aspect or third aspect. This device can also be called an electronic device, such as a terminal or a server.

[0030] Sixth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium includes a computer program. When the computer program runs on a device, the device can execute the method described in any one of the above first aspect or second aspect or third aspect. This device can also be called an electronic device, such as a terminal or a server.

[0031] In a seventh aspect, there is provided a computer program product containing instructions which, when run on a device, enable the device to execute the method described in any one of the first, second, or third aspects above. The device may also be referred to as an electronic device, such as a terminal or a server.

[0032] In an eighth aspect, an embodiment of the present application provides a chip. The chip includes a processor, and the processor is used to call a computer program in a memory to execute the method described in any one of the first, second, or third aspects.

[0033] It can be understood that for the beneficial effects that can be achieved by the methods described in the second and third aspects provided above, the device described in the fourth aspect, the equipment described in the fifth aspect, the computer-readable storage medium described in the sixth aspect, the computer program product described in the seventh aspect, and the chip described in the eighth aspect, reference may be made to the beneficial effects in the first aspect and any of its possible implementation manners, which will not be elaborated here. Description of the Drawings

[0034] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0035] Figure 2 It is a schematic flowchart of a method for generating a video album provided by an embodiment of the present application;

[0036] Figure 3 It is a schematic interface diagram when generating a video album provided by an embodiment of the present application;

[0037] Figure 4 It is a schematic flowchart of another method for generating a video album provided by an embodiment of the present application;

[0038] Figure 5 It is a schematic diagram of a process for generating a video album provided by an embodiment of the present application;

[0039] Figure 6 It is a schematic interface diagram when generating a video album provided by an embodiment of the present application;

[0040] Figure 7 It is a schematic interface diagram when generating a video album provided by an embodiment of the present application;

[0041] Figure 8 It is a schematic interface diagram when generating a video album provided by an embodiment of the present application. Detailed Embodiments

[0042] Hereinafter, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0043] Currently, terminals such as mobile phones provide a function of generating a recollective and summary video (or called video album) based on the materials in the terminal's picture gallery. Among them, in the related art, photos / videos are mainly selected from the terminal's picture gallery as materials based on dimensions such as time, location, and preset themes to generate a video album. Specifically, based on a preset theme, attributes such as the time, location, and semantic tags of the photos / videos in the picture gallery can be clustered to obtain materials for generating a video album. For example, if the preset theme is travel, the terminal can screen out the photos / videos taken by the user when traveling to a certain place at a certain time in the picture gallery and generate a video album on the theme of travel.

[0044] There are at least the following problems in the related art: Since the conditions for screening video materials are relatively fixed, that is, all materials are screened based on preset themes, the generated video albums are all videos related to these fixed themes. For example, most of the generated video albums are summary videos related to the user's life, behavior, relatives, and friends. In addition, the number of preset themes is generally limited, which will lead to duplication when the number of generated video albums is large. All of these result in that the generated video albums have a weak impact on the user's emotions, thereby leading to a weak browsing willingness of the user and a low experience.

[0045] To solve the above problems, an embodiment of the present application provides a method for generating a video album, which can be applied to the scenario of a terminal generating a video album. Specifically, based on the theme related to the user's personal attribute tags, the multimedia resources (such as photos and / or videos) in the terminal's picture gallery can be screened and aggregated to automatically generate a summary and recollective video album related to the above theme, that is, the generated video album is related to the user's personal attribute tags. The user can view the generated video album in the terminal's picture gallery. For example, after the terminal receives the operation of the user to open a certain video album in the picture gallery, it can display the multiple multimedia resources included in the video album in video form.

[0046] Among them, the above-mentioned user personal attribute tags can also be referred to as user information, which can include one or more of the user's characteristic information, the user's relationship information, or the record information generated by the user using the applications in the terminal. It can be understood that these user information can often reflect the user's personal attributes, such as the user's mood / state of mind, the user's interests / concerns, etc. Then, the multimedia resources selected from the picture library using the themes related to these information can also reflect the user's personal attributes, so that the video album generated using these materials can generate a more intuitive emotional resonance with the user, thereby stimulating the user's browsing awareness and improving the experience. In addition, different user information corresponds to different themes, and the generated related video albums are also different, no longer limited by the preset themes, thus enhancing the diversity of the generated video albums, avoiding the repetition of video albums, and further improving the user's usage experience.

[0047] In some embodiments, the terminal in the present application can be an electronic device such as a mobile phone, a tablet computer, a handheld computer, a personal computer (PC), a cellular phone, a personal digital assistant (PDA), a wearable device (such as a smart watch), a smart home device (such as a television), a vehicle-mounted computer, a game console, and an augmented reality (AR) / virtual reality (VR) device, etc. If the electronic device refers to a device with the storage capacity of multimedia resources such as pictures and videos (or referred to as media files), the specific form of the terminal in the embodiments of the present application is not particularly limited.

[0048] It should be noted that in the embodiments of the present application, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs. For example, in the embodiments of the present application, the acquisition and processing of the user's personal information are carried out under the authorization of the user. This is uniformly stated here and will not be repeated below.

[0049] The implementation manners of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0050] Please refer to Figure 1 , taking the terminal as a mobile phone as an example. Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 1As shown in the figure, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0051] Among them, the sensor module 180 may include a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.

[0052] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0053] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0054] The controller may be the nerve center and command center of the electronic device. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.

[0055] A memory can also be set in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0056] The charging management module 140 is used to receive a charging input from a charger. Herein, the charger can be a wireless charger or a wired charger. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.

[0057] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc.

[0058] The wireless communication function of the electronic device can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0059] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device.

[0060] The modulation and demodulation processor can include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.), or displays an image or video through the display screen 194.

[0061] The wireless communication module 160 may provide solutions for wireless communications applied to an electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc.

[0062] The electronic device implements a display function through a GPU, a display screen 194, an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information. In some embodiments of the present application, the display screen 194 may be used to display multimedia resources in a video album in video display.

[0063] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel may adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include one or N display screens 194, where N is a positive integer greater than 1.

[0064] The electronic device may implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.

[0065] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin tone of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.

[0066] The camera 193 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. For example, in the embodiments of this application, the user can use the camera 193 to take photos / videos. After the user finishes shooting, the taken photos / videos can be saved in the electronic device, such as being saved in the gallery of the electronic device. It should be noted that the photos / videos in the gallery of the electronic device can also be downloaded by the user from a third-party application or the network.

[0067] The external memory interface 120 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card. For instance, in the embodiments of this application, the photos / videos taken by the user can be stored in the external memory card. Of course, this is only an example and does not constitute a limitation to the solution of this application.

[0068] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121. For example, in the embodiments of the present application, the processor 110 can generate a theme, a video style, video materials, etc. for generating a video album based on user information. The processor 110 can also screen out pictures and / or videos that match the theme from the picture library based on the generated theme and aggregate them into a video album for the user to view. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device (such as audio data, phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0069] The electronic device can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.

[0070] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0071] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device can listen to music or hands-free calls through the speaker 170A.

[0072] The receiver 170B, also known as an "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the ear.

[0073] The microphone 170C, also known as a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal.

[0074] The button 190 includes a power-on button, volume buttons, etc. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light, which can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect the SIM card.

[0075] The methods in the following embodiments can all be implemented in an electronic device with the above hardware structure.

[0076] Figure 2 It is a schematic flowchart of a method for generating a video album provided by an embodiment of the present application. As Figure 2 shown, the method may include:

[0077] S201. The terminal receives a first operation by the user to open the video album in the gallery.

[0078] The video album is a recollective and summary album automatically generated by the terminal based on the materials in the terminal's gallery. Generally, a video album may include multiple multimedia resources, such as pictures and / or videos. The terminal can automatically screen the photos / videos in the terminal's gallery and use them as materials to generate a video album. After generating the video album, a corresponding entry can be provided in the terminal's gallery for the user to view the generated video album. That is to say, when the user wants to view the above video album, the user can perform a corresponding operation on the corresponding entry in the gallery to trigger the terminal to display the multimedia resources included in the video album. In this embodiment, the operation on the corresponding entry in the gallery can also be referred to as the operation of opening the video album in the gallery. For the sake of convenience of description, in the embodiments of the present application, it is simply referred to as the first operation.

[0079] Exemplarily, the above first operation can be an operation in any form. For example, taking the terminal as a mobile phone and the display screen of the mobile phone as a touch screen, the first operation can be a touch operation by the user on the corresponding entry of the video album displayed on the display screen of the mobile phone, such as a click operation, a double-click operation, etc. Among them, as an example, the corresponding entry of the video album can be a thumbnail, a folder icon, or an album icon (or called a photo album icon) and other controls displayed in the interface of the gallery. Correspondingly, the first operation can specifically be a touch operation such as a click operation or a double-click operation on the control. The first operation can also be a voice command input by the user, etc. This embodiment does not specifically limit the first operation here.

[0080] For example, in combination with Figure 3 , taking the terminal as a mobile phone and the display screen of the mobile phone as a touch screen, taking the album icon as the corresponding entry of the video album and opening the video album by performing a touch operation on the corresponding entry of the video album displayed on the display screen of the mobile phone. The user can open the gallery in the mobile phone by performing a corresponding operation. For example, as Figure 3As shown in (a) therein, the user can perform a click operation on the icon 301 of the photo gallery displayed on the mobile phone desktop. In response, as Figure 3 shown in (b) therein, the mobile phone can open the photo gallery and display the interface 302 of the photo gallery. Generally, in order to facilitate the user to view the pictures / videos in the photo gallery, the pictures / videos can be classified and managed according to various different classification methods, such as classified and managed according to time, people, scenes, tags, etc. Correspondingly, the photo gallery application provides different sub-interfaces for the user to view the pictures / videos according to different classification methods. The video album is also similar to a classification management method. Therefore, in the photo gallery, a sub-interface is correspondingly provided for the user to view the video album. Among them, these sub-interfaces can be entered through the menu options included in the interface of the photo gallery. For example, as Figure 3 shown in (b) therein, the interface 302 of the photo gallery includes a "Moment" control 303. The "Moment" control 303 is a menu option for entering the corresponding sub-interface of the video album. The user can operate on the "Moment" control 303, such as a click operation. In response, as Figure 3 shown in (c) therein, the mobile phone can display the sub-interface 304. The sub-interface 304 provides an entry for the video album automatically generated by the mobile phone. For example, the sub-interface 304 includes one or more album icons of the video albums automatically generated by the mobile phone. The user can operate on the corresponding entry, such as the album icon, that is, perform the first operation, to view the corresponding video album. For example, the user can operate on the album icon 305, such as a click operation, to view the video album corresponding to the album icon 305.

[0081] S202. In response to the operation in S201, the terminal displays multiple multimedia resources in the video album in video form.

[0082] After receiving the first operation of the user to open the video album in the photo gallery, in response, the terminal can display multiple multimedia resources in the video album in video form for the user to view.

[0083] In the embodiments of the present application, the multiple multimedia resources for generating a video album (or, the multiple multimedia resources included in the video album) are screened from the terminal's picture gallery based on a target theme. Among them, the above-mentioned target theme is related to one or more pieces of user information such as the characteristic information of the user using the terminal, the user's relationship information, and the record information generated by the user using the applications in the terminal. It can be understood that the user's characteristic information can be used to characterize the user's behavior preferences, the user's relationship information can be used to characterize the user's social relationships, and the record information generated by the user using the applications in the terminal can be used to characterize the things that the user has recently paid attention to / cared about. These pieces of user information are related to the user's personal attributes. Therefore, the target theme generated based on the user information related to the user's personal attributes can also reflect the user's personal attributes, and the multimedia resources screened based on this target theme can also reflect the user's personal attributes. Thus, the video album generated using these materials can generate a more intuitive emotional resonance with the user, thereby stimulating the user's browsing awareness and improving the experience.

[0084] It can be understood that the user information of the user using the terminal obtained at different times is not the same. For example, at a certain period of time, the user information obtained is used to characterize that the user has recently made new friends, and at another period of time, the user information obtained is used to characterize that the user often exercises recently. Then, based on different user information, different types of target themes can be generated. In this embodiment, the types of themes can include but are not limited to emotional, blessing, summary of hobbies, personal habits, etc. For example, continuing with the above example, based on the obtained user information used to characterize that the user has recently made a new friend, an emotional target theme can be generated. Based on the obtained user information characterizing that the user often exercises recently, a target theme of the summary of hobbies can be generated. Thus, the terminal can screen photos / videos corresponding to the target theme from the terminal's picture gallery as materials to generate a video album based on different target themes generated in different periods.

[0085] The following is combined with Figure 4 to introduce the process of generating a video album. As Figure 4 shown, the process of generating a video album includes: S401 - S406.

[0086] S401. The terminal obtains user information.

[0087] Among them, the user information can be information that can characterize the user's personality attributes to a certain extent, such as reflecting the user's recent concerns, recent hobbies, recent mood, etc. As described in S202, the user information can include one or more of the user's characteristic information, the user's relationship information, and the record information generated by the user using the applications in the terminal. The record information generated by the user using the applications in the terminal can include one or more of social media data, browsing data, or shopping data.

[0088] In some embodiments of the present application, the terminal can obtain the user portrait of the user and determine the user's characteristic information based on the obtained user portrait. Among them, the user portrait usually can include data of one or more tags characterizing the user's personal basic information, behavior data, social network, consumption habits, or psychological characteristics, etc. For example, the user portrait includes at least one of the user's gender, age, geographical location, occupation, education level, the time or frequency of using a certain application by the user, etc. The user portrait generally can reflect the overall information of the user. The terminal can obtain the characteristics or preferences of the user through the user's user portrait. Or rather, the terminal can obtain the characteristic information used to characterize the user's behavior preferences through the user's user portrait. Among them, the user's user portrait can be obtained by the terminal from its own database.

[0089] In some embodiments of the present application, the terminal can obtain the personal knowledge graph of the user and determine the user's relationship information based on the obtained personal knowledge graph. Among them, the user's personal knowledge graph usually can include information of one or more persons related to the user, and the relationship between the user and the related persons, such as family relationship, friendship, or lover relationship, etc. The personal knowledge graph generally can reflect the social relationship of the user. The terminal can obtain the relationship information used to characterize the relationship between the user and the related persons through the user's personal knowledge graph. Among them, the user's personal knowledge graph can be obtained by the terminal from its own database. Further, the user's user portrait and the user's personal knowledge graph can be stored in the same database or in different databases, and the present application does not make any limitation thereto.

[0090] In some embodiments of the present application, the terminal can obtain the record information generated by the user using the application installed on the terminal. For example, the terminal can obtain the record information generated by the user using the application from the data of the application stored locally on the terminal, or the terminal can obtain the record information generated by the user using the application from the server corresponding to the application. Taking the application as an instant messaging application and the social media data as chat records as an example, the terminal can obtain the user's recent chat records from the terminal locally or from the server of the instant messaging application. The chat records may contain information such as the user's recent mood or emotional relationship. Taking the application as a browser application as an example, the terminal can obtain the user's recent browsing records from the terminal locally or from the server of the browser application. The browsing records may contain things / events that the user has recently paid attention to or is interested in. Taking the application as a shopping application as an example, the terminal can obtain the user's recent shopping data from the terminal locally or from the server of the shopping application, such as purchase records, shopping cart and other data. The shopping data may contain items that the user has recently paid attention to or is interested in, and the user's shopping preferences, etc.

[0091] It should be noted that the user portrait of the user, the personal knowledge graph of the user, and the record information generated by the user using the applications in the terminal are all carried out under the authorization of the user. For example, before the terminal needs to obtain this user information, a prompt interface can be displayed to ask the user whether to allow the terminal to obtain this user information, and after obtaining the user's authorization, the operation of obtaining the user information is performed.

[0092] It should be noted that in some embodiments, the terminal can perform the operation of obtaining user information every predetermined period. The acquisition periods of different types of user information, such as the user's characteristic information, the user's relationship information, or the record information generated by the user using the applications in the terminal, can be the same or different. For example, the frequency of update of the user's characteristic information is generally relatively low, so the acquisition period of the characteristic information can be relatively long, while the record information generated by the user using the applications in the terminal may have a relatively high update frequency, so the acquisition period of the record information can be shorter than that of the characteristic information. In some other embodiments, the terminal can also perform the operation of obtaining user information after receiving certain specific operations of the user, such as power-on, an operation to trigger the generation of a video album. The present application embodiments do not make specific restrictions on the acquisition timing of user information here.

[0093] S402. The terminal extracts features from the user information to obtain a feature vector of the user information, which is used to characterize the features of the user information and hides the privacy content in the user information.

[0094] In some embodiments of the present application, private content refers to the content in user information that represents the user's private information. For example, the user's age, gender, occupation, hobbies, etc. Further, in the embodiments of the present application, hiding the private content in user information can be achieved by the terminal extracting features from the user information. It can be understood that in the present application, the user information is desensitized by means of feature extraction.

[0095] Exemplarily, after obtaining the user information, the terminal can extract features from the obtained user information to obtain the feature vector of the user information. The user information includes the shopping data generated by the user using the shopping application, such as shopping records. The shopping record contains that the user recently purchased a certain fitness equipment. The terminal can extract features from the shopping record to obtain the feature vector of the shopping record, where the feature vector can represent the fitness equipment purchased by the user, but the fitness equipment is not explicitly shown. For another example, the user information includes the browsing data generated by the user using the browsing application, such as browsing records. The browsing record can reflect the things that the user recently pays attention to and is interested in. The terminal can extract features from the browsing record to obtain the feature vector of the browsing record, where the feature vector can represent the things that the user pays attention to and is interested in, but the things that the user pays attention to and is interested in are not directly and explicitly shown. In this way, the user's personal privacy can be protected to the greatest extent, ensuring the security and confidentiality of the user information.

[0096] It should be noted that the above illustration is an example of protecting the user privacy by extracting features from the user information. In some other embodiments, other encryption methods can also be used to protect the user information, and the embodiments of the present application do not make specific limitations here.

[0097] S403. The terminal determines whether to trigger the task of generating a video album based on the feature vector.

[0098] After obtaining the user information and extracting the feature vector from the user information, the terminal can determine whether to trigger the task of generating a video album based on the feature vector. In some embodiments, the terminal can input the feature vector into an Artificial Intelligence (AI) model. The AI model can output indication information for indicating whether to trigger the task of generating a video album. In this embodiment, the AI model has the function of deciding whether to trigger the task of generating a video album according to the feature vector of the user information. The AI model can be pre-trained based on sample data, and the sample data includes sample feature vectors and labels corresponding to the sample feature vectors, and the labels are used to indicate whether the corresponding sample feature vectors are allowed to trigger the task of generating a video album.

[0099] It should be noted that the above embodiments are described by taking the feature vector of the extracted user information as the input of the AI to obtain the result of the task of determining whether to trigger the generation of a video album as an example. In some other embodiments, the terminal may also input the user information or the desensitized user information into the AI model to obtain the result of the task for indicating whether to trigger the generation of a video album. The embodiments of the present application do not make specific limitations here.

[0100] In the above embodiments, the AI model is used to decide whether to trigger the task of generating a video album. In some other embodiments, the terminal may directly determine whether to trigger the task of generating a video album according to the user information, or the desensitized user information, or the feature vector of the user information. Exemplarily, taking the terminal's determination of whether to trigger the task of generating a video album according to the feature vector of the user information as an example. After the terminal obtains the user information, extracts the features of the user information, and obtains the feature vector of the user information, the terminal can decide whether to trigger the task of generating a video album by determining whether the extracted feature vector meets the preset conditions. As an example, the above preset condition may be that the frequency of occurrence of the extracted feature vector is greater than a predetermined frequency. For example, the terminal extracts the features of the user information to obtain a feature vector. When the frequency of occurrence of the obtained feature vector is greater than the preset frequency, it can be determined that it is allowed to generate a video album according to the user information, that is, it is determined to trigger the task of generating a video album. Another example is that when the frequency of occurrence of the obtained feature vector is less than the preset frequency, it can be determined that it is not allowed to generate a video album according to the user information, that is, it is determined not to trigger the task of generating a video album. Of course, it can also be other conditions, and the embodiments of the present application do not make specific limitations thereto.

[0101] If it is determined in S403 to trigger the task of generating a video album, then the following S404 is executed. If it is determined not to trigger the task of generating a video album, then the process ends.

[0102] S404. The terminal generates a target theme, a target video style, and target video materials based on the feature vector.

[0103] Among them, the above target video style may include one or more of filters, color correction, special effects, or transition animations. The above target video materials may include one or more of titles, subtitles, or background music.

[0104] After the terminal determines to trigger the task of generating a video album, the terminal may generate one or more of a target theme, a target video style, and target video materials for generating a video album based on the feature vector. The following takes the target theme, the target video style, and the target video materials all being generated based on the user information as an example for description, but the present application is not limited thereto. For example, the target theme is generated based on the user information, while the target video style and the target video materials may be predefined, etc.

[0105] In some embodiments, after the terminal extracts the feature vector from the obtained user information, the terminal may input the feature vector into the AI model. The AI model may output the target theme, the target video style, and the target video material corresponding to the feature vector. In this embodiment, the AI model has the function of generating the corresponding theme, video style, and video material according to the feature vector of the user information. Similar to the above embodiment, the AI model may be pre-trained based on sample data. The sample data may include sample feature vectors and the corresponding theme, video style, and video material for the sample feature vectors.

[0106] It should be noted that the above embodiments are described by taking the feature vector of the extracted user information as the input of the AI to obtain the target theme, the target video style, and the target video material as an example. In some other embodiments, the terminal may also input the user information or the desensitized user information into the AI model to obtain the target theme, the target video style, and the target video material. The embodiments of the present application do not make specific limitations here. In addition, in the embodiments of generating the target theme, the target video style, and the target video material through the AI model, the AI model may be the same model as the model in S403 or a different model, and the embodiments of the present application do not make specific limitations here. For example, an AI model may be used to implement the decision-making of whether to trigger the task of generating a video album, as well as the generation of the target theme, the target video style, and the target video material.

[0107] The above two embodiments generate the target theme, the target video style, and the target video material through the AI model. In still some other embodiments, the terminal may also directly generate the target theme, the target video style, and the target video material according to the user information, or the desensitized user information, or the feature vector of the user information. Exemplarily, taking the terminal generating the target theme, the target video style, and the target video material according to the feature vector of the user information as an example. After the terminal obtains the user information and extracts the feature vector of the user information through feature extraction, it may analyze the feature vector to obtain information such as the user's recent mood, recent focus, and hobbies, and determine the target theme, the target video style, and the target video material according to the obtained information. Among them, in this embodiment, one or more themes, video styles, and video materials may be pre-configured in the terminal. After analyzing and obtaining the user's recent mood, recent focus, and hobbies, the terminal may select the theme, video style, and video material that are adapted to this information from the pre-configured multiple themes, video styles, and video materials as the target theme, the target video style, and the target video material.

[0108] Exemplarily, take the example of inputting the feature vector of user information into an AI model to obtain the corresponding target theme, target video style, and target video material. For example, user information includes the user's relationship information and social media data generated by the user using an instant messaging application, such as chat records. The chat records contain chat content indicating that the user is recently in a relationship. The relationship information is used to indicate that the user has recently got a boyfriend. The terminal can extract features from the chat records and relationship information, and after obtaining the feature vector, input it into the AI model. The AI model can output a target theme of the emotional type suitable for a relationship. The AI model can also output a target video style suitable for a relationship, such as a pink filter, and output target video material suitable for a relationship, such as the background music being "Just Met You", etc. Another example is that user information includes shopping data generated by the user using a shopping application. The shopping data indicates that the user has recently often purchased / browsed fitness equipment, etc. The terminal can input the feature vector obtained by extracting features from the shopping data into the AI model, and the AI model can output a target theme of the hobby type suitable for fitness. The AI model can also output a target video style suitable for fitness, such as a timeline special effect, etc. It can also output target video material suitable for fitness, such as the subtitle being an explanation of a certain action point, etc.

[0109] In the embodiment of generating the target theme, target video style, and target video material through the AI model, the terminal can also correct the target theme, target video style, and target video material output by the AI model according to the standardization rules. In this way, it can prevent the AI model from generating negative results with bad emotional tendencies (such as: negative, offensive, discriminatory, off-topic, etc.) due to insufficient discrimination ability.

[0110] Specifically, the standardization rules can be preset by developers. By matching the prefabricated standardization rules with the target theme, target video style, and target video material output by the AI model, the target theme, target video style, and target video material that do not conform to the above standardization rules can be corrected to make them standardized.

[0111] Exemplarily, take the example that user information includes social media data generated by the user using an instant messaging application, such as chat records. When the chat records contain chat content indicating that the user is in a low mood, such as "in a bad mood", "sad", "cheer up", "don't be unhappy", etc., the terminal inputs the feature vector of the chat records into the AI model. The AI model may output a target theme, target video style, and target video material corresponding to these chat contents with a low mood, that is, the output target theme, target video style, and target video material may be relatively negative and gloomy. The terminal can correct these negative and gloomy output results through the standardization rules. For example, the negative target theme can be corrected to a positive target theme, etc.

[0112] S405. Screen multimedia resources from the terminal's picture gallery based on the target theme.

[0113] Among them, the multimedia resources can be pictures or videos. The multimedia resources in the picture gallery can be taken by the user using the terminal, downloaded from a third-party application, screenshots, or sent from other devices. The embodiments of the present application do not make specific limitations here.

[0114] In the embodiments of the present application, after determining the target theme, targeted search can be performed in the picture gallery to screen out multimedia resources that match the target theme as materials for generating a video album.

[0115] In some embodiments, the terminal can screen multimedia resources related to (or matching) the target theme from the picture gallery according to one or more of the time, location, semantic tags, or the relationship with the terminal user of each multimedia resource in the picture gallery. For example, the terminal can determine the relevance of each multimedia resource to the target theme according to one or more of the time, location, semantic tags, or the relationship with the terminal user of each multimedia resource, and screen out multimedia resources that match the target theme based on the obtained relevance.

[0116] Among them, the time of the above multimedia resources can refer to the acquisition time of the multimedia resources. For example, if the multimedia resource is a picture taken by the user using the terminal, the time of the picture can refer to the shooting time of the photo. Another example is that if the multimedia resource is a video downloaded by the user from a third-party application, the time of the video can refer to the download time of the video. The location of the multimedia resource (or called Point of Interest (POI), or Area of Interest (AOI)) can refer to the location where the multimedia resource is acquired. For example, if the multimedia resource is a video taken by the user using the terminal, the location of the video can refer to the shooting location of the video. Another example is that if the multimedia resource is a picture and the picture is a screenshot, the location of the picture can refer to the location where the terminal is when the picture is intercepted. It can be understood that the time and location of the multimedia resources can be obtained from the attribute information of the multimedia resources. In some other embodiments, the time and location of the multimedia resources can also refer to the time and location of the actual content displayed by the multimedia resources. For example, if the multimedia resource is a picture and the content shown in the picture is the scenery of Area A in winter, the time of the picture can be winter and the location can be Area A. In this embodiment, the time and location of the multimedia resources can be obtained by analyzing the multimedia resources.

[0117] The semantic tags of the above-mentioned multimedia resources may refer to the tags that mark the actual content displayed by the multimedia resources. For example, if the multimedia resource is a picture, and the picture is a photo of a person and a pet, then the semantic tags of the picture may include: person, pet. Another example, if the multimedia resource is a video, and the video is a video of a snow mountain scenery, then the semantic tags of the video may include: mountain, snow, snow mountain. The semantic tags of the above-mentioned multimedia resources can be obtained by analyzing the actual content displayed by the multimedia resources.

[0118] The relationship between the above-mentioned multimedia resources and the user of the terminal may refer to the relationship between the person or object included in the actual content displayed by the multimedia resources and the user of the terminal. For example: family relationship, lover relationship, friend relationship, etc. As an example, the content included in the actual content displayed by the multimedia resources can be analyzed to obtain the people included in the actual content displayed by the multimedia resources, and based on the relationship information of the user obtained in S401, the relationship between the multimedia resources and the user of the terminal can be determined. For example, if the multimedia resource is a picture, analyzing the content displayed by the picture obtains that the person A is included in the picture, and the relationship information of the user obtained in S401 records that the person A and the user of the terminal are in a family relationship, then it can be determined that the relationship between the multimedia resources and the user of the terminal is a family relationship. It should be noted that the method for determining the relationship between the multimedia resources and the user of the terminal here is only an example and does not constitute a limitation to this application.

[0119] The relevance between the above-mentioned multimedia resource and the target theme may refer to the degree of matching between the content actually presented by the multimedia resource and the target theme. The relevance between the multimedia resource and the target theme can be obtained by matching the content presented by the multimedia resource with the target theme. As an example, one or more of the information such as the semantic tags, time, location, or the relationship with the user of the terminal of the multimedia resource can be matched with the target theme to obtain the relevance between the multimedia resource and the target theme. In some embodiments, for example, taking the matching of the semantic tags of the multimedia resource and the target theme to obtain the above-mentioned relevance as an example. If the target theme is a landscape theme and the semantic tags of the multimedia resource include mountains, snow, and snow-capped mountains, the matching result can be that the multimedia resource is relevant to the target theme. Another example, taking the matching of the semantic tags of the multimedia resource and the target theme to obtain the above-mentioned relevance as an example. If the target theme is a personal habit theme and the semantic tags of the multimedia resource include people and pets, the matching result can be that the multimedia resource is not relevant to the target theme. In some other embodiments, it is also possible to match one or more of the information such as the semantic tags, time, location, or the relationship with the user of the terminal of the multimedia resource with the target theme to obtain a specific matching degree value, and then determine whether the obtained matching degree is greater than a threshold value. If it is greater than the threshold value, it can be considered that the multimedia resource is relevant to the target theme; if it is less than the threshold value, it can be considered that the multimedia resource is not relevant to the target theme. It should be noted that the method for determining the relevance here is only an example and does not constitute a limitation to this application.

[0120] In some embodiments, after screening out the multimedia resources that match the target theme from the picture library according to one or more of the information such as the time, location, semantic tags, or the relationship with the user of the terminal of each multimedia resource in the picture library, the matched multimedia resources can be directly used as the materials for generating the album video.

[0121] In some other embodiments, after obtaining the multimedia resources through matching, the similar or duplicate multimedia resources among the matched multimedia resources can be further screened. For example, for multiple similar multimedia resources among the matched multimedia resources, only one or more of them can be selected as the materials for generating the video album. In this way, the duplication of the materials for generating the video album can be avoided.

[0122] Among them, the selected multimedia resources can be the multimedia resources that meet the preset conditions among multiple similar multimedia resources. Among them, the preset conditions include one or more of the following conditions: the clarity is greater than the threshold value (or the clarity is the highest), the integrity is greater than the threshold value (or the integrity is the highest), or the aesthetic score is greater than the threshold value (or the aesthetic score is the highest).

[0123] The integrity of a multimedia resource may refer to the completeness of the objects included in the multimedia resource. For example, if the multimedia resource is a photo and a person is included in the photo, the integrity of the photo may refer to the completeness of the person in the photo. The above-mentioned aesthetic score of the multimedia resource may be a comprehensive score obtained by evaluating the multimedia resource from multiple dimensions such as objective, subjective, photographic, and portrait aesthetic perspectives. Among them, the terminal can obtain the aesthetic score of the multimedia resource by using a deep learning algorithm, or can also obtain the aesthetic score of the multimedia resource by performing logical judgment and analysis using manually designed rules. This application does not specifically limit the method for obtaining the aesthetic score.

[0124] S406. The terminal generates a video album based on the selected multimedia resources, the target video style, and the target video materials.

[0125] After obtaining the target video style, the target video materials, and selecting the multimedia resources, the terminal can synthesize the target video style, the target video materials, and the selected multimedia resources to obtain the corresponding video album.

[0126] It can be understood that after generating the video album, the terminal can provide an entry for the video album for the user to view the video album. Specifically, it can be implemented through the above Figure 2 illustrated embodiments, which will not be elaborated here in detail.

[0127] Next, in combination with Figure 5 and examples, the process of generating a video album provided by the embodiments of the present application will be illustrated by examples.

[0128] As Figure 5 shown, the terminal can obtain the record information generated by the user using the applications in the terminal. For example, the obtained record information may include: chat records, browsing records (such as articles, news, videos browsed), shopping data (such as purchase records, shopping carts), etc. The terminal can also obtain the user portrait and the personal knowledge graph of the user. After that, the terminal can perform feature extraction on the obtained record information, user portrait, and personal knowledge graph generated by the user using the applications in the terminal to obtain a feature vector. For example, this feature vector can often reflect the user's recent mood, recent focus, interests, etc. After obtaining the feature vector, the terminal can input the feature vector into the AI model so that the AI model can analyze and process the feature vector.

[0129] The AI model can output whether to trigger the task of generating a video album. In the case where the AI model outputs not to trigger the task of generating a video album, the process ends. In the case where the AI model outputs to trigger the task of generating a video album, the AI model can also output the target theme of the video album to be generated, the target video style (e.g., filters, color grading, special effects, etc.), and the target video materials (e.g., background music, titles, captions, etc.). Then, the terminal can further correct the generated target theme, target video style, and target video materials according to preset rules and specifications, so as to standardize the results output by the AI model. Then, the terminal can perform a targeted search in the picture library according to the corrected target theme, filter out the materials that match the target theme, and generate a video album based on the filtered materials, the target video style, and the target video materials.

[0130] The above process will be described below with specific examples.

[0131] Combined with Figure 6 , taking the user information including the chat records generated by the user using the instant messaging application in the terminal as an example. As Figure 6 shown in (a) of Figure 6 , the terminal can obtain the user's recent chat records 601 from the local or the server of the instant messaging application. Then, the terminal can perform feature extraction on the obtained chat records 601 to obtain feature vectors such as "Valentine's Day" and "boyfriend". After obtaining the feature vectors, the terminal can input the feature vectors into the AI model. The AI model outputs to trigger the task of generating a video album, and generates a target theme of the emotional category, a target video style adapted to a pink filter, a heart-shaped special effect, etc., and target video materials adapted to a background music with the song name "Just Met You" and a title of "Fortunately Met You". When the generated target theme, target video style, and target video materials meet the preset rules, the terminal can directly perform a targeted search in the picture library according to the target theme of the emotional category. As Figure 6 shown in (b) of

[0132] Combined with Figure 7 , taking the user information including the shopping data generated by the user using the shopping application and the browser application in the terminal, and the browsing records generated by using the browser application as an example. As Figure 7As shown in (a), the terminal can obtain the user's recent shopping orders 7011 and recently browsed news 7012 from the servers of the shopping application and the browser application. Afterwards, the terminal can perform feature extraction on the obtained shopping orders 7011 and recently browsed news 7012 to obtain feature vectors such as "sports bike", "yoga", and "exercise method". After obtaining the feature vector, the terminal can input the feature vector into the AI ​​model. The AI ​​model output triggers the task of generating a video album, and generates target themes of interests and hobbies, such as fast motion and / or slow motion special effects, target video styles adapted according to timeline transitions, and adapted target video materials such as subtitles explaining the key points of a certain action. When the generated target theme, target video style, and target video material meet the preset rules, the terminal can directly conduct a targeted search in the gallery based on the target theme of interests and hobbies. As shown in FIG. Figure 7 As shown in (b) in FIG. 70, the content of photo B 7021 is "user running", and the content of photo C 7022 is "user playing football". Obviously, both photo B 7021 and photo C 7022 are compatible with the target theme of hobbies. The terminal can filter out these photos as materials matching the target theme. Of course, the terminal can also filter out other pictures / videos that are compatible with the theme of hobbies. Figure 7 As shown in (c) in FIG. 7 , the terminal may generate a video album 703 titled “Sweating Everyday” based on the filtered materials (multimedia resources such as photo B 7021 and photo C 7022 ), the target video style and the target video materials.

[0133] Combination Figure 8 , taking the example of user information including chat records generated by the user using the SMS application in the terminal. Figure 8 As shown in (a) in the figure, the terminal can obtain the user's recent SMS record 801 from the local or SMS server. After that, the terminal can perform feature extraction on the obtained SMS record 801 to obtain feature vectors such as "bad mood", "sad", "think positively", and "don't be unhappy". After obtaining the feature vector, the terminal can input the feature vector into the AI ​​model. The output of the AI ​​model triggers the task of generating a video album, and generates a target theme, such as a target video style adapted by cold tones, soft light filters, etc., and a target video material adapted by a title such as "Life is not as expected". The terminal can correct these target themes, target video styles, and target video materials that do not meet the preset rules to obtain a target theme that meets the preset rules, and the target video style and target video material meet the preset rules. For example, the target theme is corrected to an encouraging target theme, the cold tones are corrected to warm tones, and the title "Life is not as expected" is corrected to "Life is a little bit happy", etc. Then, the terminal can conduct a targeted search in the gallery according to the corrected target theme (i.e., the encouraging target theme). As Figure 8As shown in (b) thereof, the content of photo D802 is "Sunrise at Sea". Obviously, photo D802 is suitable for the target theme of the encouragement category, and the terminal can filter it out as the material matching the target theme. Of course, the terminal can also filter out other pictures / videos suitable for the encouragement theme. For example, Figure 8 As shown in (c) thereof, the terminal can generate a video album 803 titled "Small Joys in Life" based on the filtered materials (such as multimedia resources like photo D802), the target video style, and the target video materials.

[0134] It should be noted that the above embodiments are described by taking the operation of generating a video album as being implemented in the terminal as an example. In some other embodiments, the operation of generating a video album can also be achieved by the cooperation of the terminal and the server. For example, S402, S403, and S404 can be executed by the server, and S401, S405, and S406 can be executed by the terminal. Another example is that S403 and S404 can be executed by the server, and S401, S402, S405, and S406 can be executed by the terminal. In the embodiments where the terminal and the server cooperate to achieve this process, some interactions between the terminal and the server are required. For example, taking S403 and S404 being executed by the server and S401, S402, S405, and S406 being executed by the terminal as an example, after the terminal executes S402 to obtain the feature vector, the feature vector can be sent to the server so that the server can execute S403 - S404. After the server determines to trigger the task of generating a video album and obtains the target theme, the target video style, and the target video materials, the target theme, the target video style, and the target video materials can be sent to the terminal so that the terminal can execute S405 - S406. In the case where S402, S403, and S404 can be executed by the server and S401, S405, and S406 can be executed by the terminal. After the terminal obtains the user information in S401, the user information can be encrypted and sent to the server. After receiving it, the server can decrypt the user information using the corresponding decryption method to execute S402 - S404. In addition, after the server determines to trigger the task of generating a video album and obtains the target theme, the target video style, and the target video materials, the target theme, the target video style, and the target video materials can be sent to the terminal so that the terminal can execute S405 - S406.

[0135] In the technical solution provided by the embodiments of the present application, the terminal screens and aggregates the multimedia resources in the terminal gallery based on the theme related to the user's personal attribute tags, and then automatically generates a summary and recollective video album related to the above theme. In this way, the generated video album is related to the user's personal attribute tags, which can allow the user to have a more intuitive emotional resonance, thereby stimulating the user's browsing awareness and improving the experience. In addition, since different user personal attribute tags correspond to different themes, the generated related video albums are also different, and are no longer limited to the preset themes, thus enhancing the diversity of the generated video albums, avoiding the repetition of video albums, and further improving the user's usage experience.

[0136] In addition, the video style, video materials, etc. of the video album can be generated based on the user information related to the user's personal attributes, further enhancing the impact of the video album on the user's emotions and improving the experience.

[0137] Some other embodiments of the present application further provide a computer storage medium, which may include computer instructions. When the computer instructions run on the terminal, the terminal is caused to execute each step executed by the terminal (such as a mobile phone) or the server in the above embodiments.

[0138] Some other embodiments of the present application further provide a computer program product. When the computer program product runs on a computer, the computer is caused to execute each step executed by the terminal (such as a mobile phone) or the server in the above embodiments.

[0139] Some other embodiments of the present application further provide a device, which has the function of implementing the behavior of the terminal (such as a mobile phone) or the server in the above embodiments. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example, an input unit or module, a display unit or module, an acquisition unit or module, a processing unit or module. As an example, the input unit or module can execute the above S201; the display unit or module can execute the above S202. As another example, the acquisition unit or module can execute the above S401; the processing unit or module can execute the above S402, S403, S404, S405 and S406.

[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0141] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0142] The units described as separate components may or may not be physically separated. The components displayed as units can be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0143] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0144] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical discs that can store program codes.

[0145] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for generating a video album, characterized in that, the method includes: The terminal receives a first operation by the user to open a video album in the gallery, and the video album is generated based on a plurality of multimedia resources, and the multimedia resources include pictures and / or videos; In response to the first operation, the terminal displays the plurality of multimedia resources in the video album in video form; Among them, the multimedia resources for generating the video album are screened from the terminal's gallery based on a target theme, and the target theme is related to the user information of the user of the terminal, and the user information includes at least one of the following information: the characteristic information of the user, the relationship information of the user, or the record information generated by the user using the applications in the terminal. The characteristic information is used to characterize the user's behavior preferences, the relationship information is used to characterize the user's social relationships, and the record information includes at least one of the following data: social media data, browsing data, or shopping data.

2. The method according to claim 1, characterized in that, The terminal displays the plurality of multimedia resources in the video album in video form, including: Based on a target video style and target video materials, the plurality of multimedia resources in the video album are displayed in video form; Among them, the target video style and the target video materials are also related to the user information; the target video style includes one or more of: filters, color grading, special effects or transition animations; the target video materials include one or more of: titles, subtitles, or background music.

3. The method according to claim 2, characterized in that, The method further includes: Obtaining the user information; Generating the target theme, the target video style and the target video materials based on the user information; Screening multimedia resources from the terminal's gallery based on the target theme; Generating the video album based on the screened multimedia resources, the target video style and the target video materials.

4. The method according to claim 3, characterized in that, After obtaining the user information, the method further includes: Determining a task for triggering the generation of a video album based on the user information.

5. The method according to claim 2 or 3, characterized in that, The generating the target theme, the target video style and the target video materials based on the user information includes: Performing feature extraction on the user information to obtain a feature vector of the user information, and the feature vector is used to characterize the features of the user information and hides the privacy content in the user information; Using the feature vector of the user information as an input, and through an artificial intelligence AI model, generating the target theme, the target video style and the target video materials; Among them, the AI model has the function of generating corresponding themes, video styles and video materials based on user information.

6. The method according to claim 4, characterized in that, The determining a task for triggering the generation of the video album based on the user information includes: Extract features from the user information to obtain a feature vector of the user information, where the feature vector is used to characterize the features of the user information and hides the privacy content in the user information; Use the feature vector of the user information as input, and through an AI model, output a task that triggers the generation of the video album; Among them, the AI model has the function of determining whether to trigger the task of generating a video album based on the user information.

7. The method according to any one of claims 3-6, characterized in that, Before screening multimedia resources from the terminal's picture library based on the target theme to generate the video album, the method further includes: Correct the target theme, the target video style, and the target video material according to a preset rule, where the preset rule is used to correct the emotional tendency of the theme, video style, and video material of the video album.

8. The method according to any one of claims 1-7, characterized in that, The obtaining of the user information includes: Obtain the user portrait and personal knowledge graph of the user; Determine the feature information of the user based on the user portrait, and determine the relationship information of the user based on the personal knowledge graph.

9. An electronic device, characterized in that, The electronic device includes: a processor; a memory; and a computer program, where the computer program is stored on the memory, and when the computer program is executed by the processor, the electronic device executes the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the electronic device executes the method according to any one of claims 1-8.

Citation Information

Cited By

  • Method, device, equipment and product for generating video

    CN120935430A