Multimedia file processing method and device and electronic equipment

By performing multi-layer perception processing on the multimedia file set in the multimedia database, text feature vectors and correlation are generated, the problem of inaccurate push of multimedia files is solved, and more accurate user multimedia file recommendations are achieved.

CN120196772APending Publication Date: 2025-06-24VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249982.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing multimedia file push technology has the problem of inaccurate push, making it difficult for users to quickly find the required multimedia files.

Method used

Through the data push model, multi-layer perceptual processing is performed on the multimedia file set in the multimedia database to generate text feature vectors and correlation degrees, and based on these feature vectors and correlation degrees, the recollection video matching the user input is output.

Benefits of technology

It improves the accuracy of multimedia file push, ensures that the pushed memorized video matches user needs, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196772A_ABST
    Figure CN120196772A_ABST
Patent Text Reader

Abstract

The invention discloses a multimedia file processing method and device and electronic equipment, and belongs to the technical field of artificial intelligence. The method comprises the following steps: receiving a first input for a memory creation interface; in response to the first input, performing multi-layer perception processing on a splicing feature vector corresponding to each multimedia file set in a multimedia database through a data pushing model to obtain a text feature vector corresponding to each multimedia file set, the relevancy between each multimedia file in each multimedia file set and the recall theme created by the first input is obtained; and outputting a recall video corresponding to the recall theme created by the first input based on the text feature vector and the relevancy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a method, apparatus, and electronic device for processing multimedia files. Background Art

[0002] With the popularization of electronic devices, the number of multimedia files such as photos and videos taken by people is increasing. These multimedia files have become precious memories and assets of users. However, due to the large number of multimedia files, it is difficult for users to quickly find the multimedia files they want. Therefore, the multimedia file push technology has emerged. The multimedia file push technology is a method of automatically identifying users' past multimedia files by analyzing the user's photo album library through algorithms and pushing them to the users. This method of pushing multimedia files can help users find the multimedia files they want to recall more conveniently. Currently, all the methods of pushing multimedia files push recall videos with fixed themes according to pre-set push rules, resulting in the problem of inaccurate pushing of multimedia files. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a method, apparatus, and electronic device for processing multimedia files to accurately push multimedia files to users.

[0004] In a first aspect, the embodiments of this application provide a method for processing multimedia files, which includes:

[0005] Receiving a first input to a memory creation interface;

[0006] In response to the first input, through a data push model, performing multi-layer perception processing on the splicing feature vectors corresponding to each multimedia file set in the multimedia database to obtain the text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the memory theme created by the first input;

[0007] Based on the text feature vectors and the relevance, outputting a memory video corresponding to the memory theme created by the first input.

[0008] In a second aspect, the embodiments of this application provide a method for processing multimedia files, which includes:

[0009] Obtaining the viewing intention of the user for the multimedia files in the photo album application;

[0010] Through a data push model, performing multi-layer perception processing on the splicing feature vectors corresponding to each multimedia file set in the multimedia database to obtain the text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the theme represented by the viewing intention;

[0011] Output a recalled video corresponding to the theme characterized by the viewing intention based on the text feature vector and the relevance.

[0012] In a third aspect, an embodiment of the present application provides a multimedia file processing device, which includes:

[0013] A receiving module, configured to receive a first input to a recall creation interface;

[0014] A processing module, configured to, in response to the first input, perform a multi-layer perception process on the concatenated feature vectors corresponding to each multimedia file set in a multimedia database through a data push model, to obtain text feature vectors respectively corresponding to each multimedia file set, and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input;

[0015] An output module, configured to output a recalled video corresponding to the recall theme created by the first input based on the text feature vector and the relevance.

[0016] In a fourth aspect, an embodiment of the present application provides a multimedia file processing device, which includes:

[0017] An acquisition module, configured to acquire a user's viewing intention for multimedia files in an album application;

[0018] A processing module, configured to perform a multi-layer perception process on the concatenated feature vectors corresponding to each multimedia file set in a multimedia database through a data push model, to obtain text feature vectors respectively corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the theme characterized by the viewing intention;

[0019] An output module, configured to output a recalled video corresponding to the theme characterized by the viewing intention based on the text feature vector and the relevance.

[0020] In a fifth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect or the steps of the method described in the second aspect are implemented.

[0021] In a sixth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect or the steps of the method described in the second aspect.

[0022] Seventh aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect or the steps of the method described in the second aspect.

[0023] In the embodiment of the present application, when pushing a memory video containing multimedia files, it is pushed according to the first input of the user to the memory creation interface, and the pushed memory video corresponds to the memory theme created by the first input. In this way, according to the user's needs and the timing of pushing desired by the user, the memory video required by the user can be pushed, improving the flexibility of pushing the memory video. In addition, when pushing the memory video, the splicing feature vectors corresponding to each multimedia file set in the multimedia database can be processed by a multi-layer perception through a data push model to obtain the text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the memory theme created by the first input. Then, based on the text feature vectors and the relevance, a memory video corresponding to the memory theme created by the first input is output, rather than pushing it to the user randomly, improving the accuracy of pushing the memory video. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of a multimedia file processing method provided by some embodiments of the present application;

[0025] Figure 2 is a schematic diagram of the interface of an album application provided by some embodiments of the present application;

[0026] Figure 3 is a schematic diagram of a memory creation interface provided by some embodiments of the present application;

[0027] Figure 4 is a schematic diagram of the determination of at least one time multimedia file set provided by some embodiments of the present application;

[0028] Figure 5 is a schematic diagram of the determination of at least one location - to - location multimedia file set provided by some embodiments of the present application;

[0029] Figure 6 is a schematic diagram of the determination of at least one shooting object multimedia file set provided by some embodiments of the present application;

[0030] Figure 7 is a schematic diagram of the determination of at least one shooting content multimedia file set provided by some embodiments of the present application;

[0031] Figure 8 is a schematic diagram of the determination of the image feature vector of the image corresponding to the multimedia file provided by some embodiments of the present application;

[0032] Figure 9 is a schematic flowchart of a multimedia file processing method provided by some embodiments of the present application;

[0033] Figure 10 is a schematic structural diagram of a multimedia file processing apparatus shown by some embodiments of the present application;

[0034] Figure 11 is a schematic structural diagram of a multimedia file processing apparatus shown by some embodiments of the present application;

[0035] Figure 12 is a schematic structural diagram of an electronic device shown by some embodiments of the present application;

[0036] Figure 13 is a schematic hardware structural diagram of an electronic device shown by some embodiments of the present application. Detailed Embodiments

[0037] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.

[0038] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. generally belong to the same category, and the number of objects is not limited. For example, the first object may be one or N. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the associated objects before and after.

[0039] The terms used in the embodiments section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0040] Next, the terms related to the embodiments of the present invention will be explained.

[0041] Multimedia file: Various encoded data of media are stored in the form of files in a computer and are a collection of binary data. The multimedia files in the present application refer to photos and videos.

[0042] The identifiers in this application are texts, symbols, images, etc. used to indicate information, and can use controls or other containers as the carriers for displaying information, including but not limited to text identifiers, symbol identifiers, and image identifiers.

[0043] A control refers to the encapsulation of data and methods. A control can have its own properties and methods, where the properties are simple visitors to the control data, and the methods are some simple and visible functions of the control.

[0044] A loss function is used to measure the difference or error between the model prediction result and the true label in machine learning and deep learning algorithms.

[0045] Artificial Intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0046] The technical solution of the embodiment of this application can be applied to the scenario where a user actively creates a memory video of multimedia files in a photo album application. For example, the user's photo album application has multimedia files such as photos and videos taken by the user, and also has a memory recommendation function for recommending multimedia files to the user. The user wants to create a memory video with the theme they want based on the memory recommendation function.

[0047] Next, in combination with the accompanying drawings, through specific embodiments and their application scenarios, the multimedia file processing method provided by the embodiment of this application will be described in detail.

[0048] Figure 1 It is a schematic flowchart of a multimedia file processing method provided by the embodiment of this application. The execution subject of this multimedia file processing method can be an electronic device, which can be but is not limited to a personal computer (PC), a smart phone, a tablet computer, or a personal digital assistant (PDA), etc.

[0049] As Figure 1 shown, the multimedia file processing method provided by the embodiment of this application can include Step 110 - Step 130.

[0050] Step 110: Receive a first input to the memory creation interface.

[0051] Among them, the first input may be the user's input to the recollection creation interface. The above first input is used to obtain the text feature vectors corresponding to each multimedia file set in the multimedia database and the relevance between each multimedia file in each multimedia file set and the theme represented by the text feature vectors. The first input may be a first operation. Exemplarily, the above first input includes but is not limited to: the user's touch input to the recollection creation interface through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above first input may be: the user's touch input to the recollection creation interface, such as the first input is the user's click input to the recollection creation interface.

[0052] The recollection creation interface may be an interface for creating a recollection video. The recollection creation interface may be obtained based on the user's click operation on the recollection recommendation control in the photo album application. The recollection recommendation control may be used to implement the function of recommending multimedia files to the user.

[0053] In some embodiments of the present application, before step 110, the above-mentioned method may further include:

[0054] Receiving a second input from the user to the recollection recommendation control in the photo album application;

[0055] In response to the second input, displaying the recollection creation interface.

[0056] Among them, the second input can be the input of the user to the memory recommendation control in the photo album application. The above-mentioned second input is used to display the memory creation interface, and the second input can be the second operation. Exemplarily, the above-mentioned second input includes but is not limited to: the touch input of the user to the memory recommendation control in the photo album application through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input. For example, the above-mentioned second input can be: the touch input of the user to the memory recommendation control in the photo album application. For example, the second input is the click input of the user to the memory recommendation control in the photo album application.

[0057] In some embodiments of the present application, when the user wants to create a memory video, the user can click on the memory recommendation control in the photo album application, so that the memory creation interface can be displayed, and then the user can perform the first input in the memory creation interface.

[0058] In one example, referring to Figure 2 , the photo album application interface 21 includes some multimedia files previously taken by the user, such as multimedia file 1, multimedia file 2, multimedia file 3, multimedia file 4, multimedia file 5, multimedia file 6, multimedia file 7, multimedia file 8, multimedia file 9, and multimedia file 10. Among them, multimedia file 1 is a solo photo of Xiaoming taken in Tianjin on September 28, 2023, multimedia file 2 is a solo photo of Xiaoming taken in Beijing on October 1, 2023, multimedia file 3 is a group photo of Xiaoming and Xiaohong taken in Beijing on October 3, 2023, multimedia file 4 is a solo photo of Xiaohong taken in Beijing on October 3, 2023, multimedia file 5 is a landscape photo taken in Beijing on October 3, 2023, multimedia file 6 is a photo of a lily taken in Tangshan on October 6, 2023, multimedia file 7 is a food photo taken in Tangshan on October 6, 2023, multimedia file 8 is a photo of a pet dog taken in Tianjin on October 9, 2023, multimedia file 9 is a solo photo of Xiaoming taken in Tianjin on November 1, 2023, and multimedia file 10 is a video of Xiaoming and food taken in Tianjin on November 1, 2023. The photo album application interface 21 also includes a memory recommendation control 22. When the user clicks on the memory recommendation control 22, the memory creation interface 31 can be displayed as Figure 3 shown.

[0059] In some embodiments of the present application, by responding to a second input from the user to the memory recommendation control in the photo album application, a memory creation interface can be displayed, thus providing a channel for the user to create a memory video and enabling the user to create a memory video in the memory creation interface.

[0060] Step 120: In response to the first input, through a data push model, perform a multi-layer perception process on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database to obtain the text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the memory theme created by the first input.

[0061] Among them, the data push model can be a model pre-trained based on artificial intelligence technology for pushing multimedia files. The data push model can be, but is not limited to, a support vector machine model, a neural network model, a decision tree model, etc. based on deep learning.

[0062] The multimedia database can be a pre-constructed database containing all multimedia files in the photo album application. The multimedia database can include multiple multimedia file sets. The multimedia files in each multimedia file set can be obtained by clustering all multimedia files in the photo album application based on one clustering dimension. The clustering dimension here can include at least one of the following: shooting time, shooting location, shooting object, shooting content. The shooting time here can be the time when the multimedia file was shot, the shooting location can be the location where the multimedia file was shot, the shooting object can be the human object in the multimedia file, and the shooting content can be the content in the multimedia file, such as food, pets, figurines, Lego, concerts, etc.

[0063] It should be noted that the above shooting time and shooting location can be obtained through the attribute information of the multimedia file, specifically through the exif information of the multimedia file.

[0064] The concatenated feature vector of a multimedia file set in the embodiments of the present application can be obtained by concatenating the image feature vector and the text feature vector corresponding to the multimedia file set. The image feature vector of the multimedia file set can be determined based on the feature vectors of each multimedia file in the multimedia file set, and the text feature vector of the multimedia file set is determined based on the shooting information of each multimedia file in the multimedia file set. Specifically, how to determine the image feature vector of the multimedia file set based on the feature vectors of each multimedia file in the multimedia file set, and how to determine the text feature vector of the multimedia file set based on the shooting information of each multimedia file in the multimedia file set will be introduced in detail in subsequent embodiments.

[0065] The text feature vector corresponding to a multimedia file set in an embodiment of the present application can be a vector for characterizing the theme of the multimedia file set, and the text feature vector of the multimedia file set can be obtained after performing multi-layer perception processing on the splicing feature vector corresponding to the multimedia file set.

[0066] Step 130: Output a recall video corresponding to the recall theme created with the first input based on the text feature vector and the relevance.

[0067] Among them, the recall video can be a video corresponding to the recall theme created with the first input output by the data push model.

[0068] In some embodiments of the present application, the recall creation interface may include a recall theme input box, and the recall theme input box can be an input box for inputting a recall theme, such as Figure 3 the input box 32 in

[0069] Optionally, step 110 may specifically include:

[0070] Receive a first input for inputting a recall theme in the recall theme input box.

[0071] Among them, the recall theme input in the input box for the first input is the recall theme created with the first input in step 130 above.

[0072] In some embodiments of the present application, the user can input by himself the recall theme he wants to create in the recall theme input box, and thus a video corresponding to the recall theme input by the user can be output.

[0073] Continue to refer to Figure 3 , if the user wants to create a recall video of himself and his baby, the user can input "Beijing trip during the National Day in 2023" in the input box 32, and a recall video corresponding to "Beijing trip during the National Day in 2023" can be output.

[0074] In some embodiments of the present application, the user can customize the recall theme of the recall video by inputting a recall theme in the recall theme input box according to his own needs, thereby improving the flexibility of creating the recall video.

[0075] In some embodiments of the present application, the recall creation interface may further include at least one recall theme identifier, such as Figure 3 the "baby" recall theme identifier 33, the "National Day tour in 2023" recall theme identifier 34, and the "pet cat" recall theme identifier 35 in

[0076] Each of the above recall theme identifiers can be used to indicate a recall theme. For example, Figure 3The "Baby" memory theme identifier 33 in it is used to indicate the memory theme related to the baby, the "National Day Tour in 2023" memory theme identifier 34 is used to indicate the memory theme related to the National Day tour in 2023, and the "Pet Cat" memory theme identifier 35 is used to indicate the memory theme related to the pet cat.

[0077] To improve the determination efficiency of the memory theme created by the first input, step 110 may specifically include:

[0078] Receive a first input for one of the at least one memory theme identifiers.

[0079] Among them, the memory theme indicated by the one memory theme identifier selected from the at least one memory theme identifier in the first input is the memory theme created by the first input in the above step 130.

[0080] In some embodiments of the present application, when there is an identifier corresponding to the memory theme desired by the user among the at least one memory theme identifiers included in the memory creation interface, the user can directly click on this memory theme identifier, and then a video corresponding to the memory theme indicated by the memory theme identifier clicked by the user can be output. Specifically, multimedia files matching the memory theme indicated by the memory theme identifier clicked by the user can be selected from the album application, and then these multimedia files are displayed.

[0081] Continue to refer to Figure 3 , if the user wants to create a memory video about the baby and there is a memory theme identifier related to the baby in the memory creation interface, the user can directly click on the "National Day Tour in 2023" memory theme identifier 34, and a memory video corresponding to the "National Day Tour in 2023" can be output.

[0082] In some embodiments of the present application, when there is an identifier corresponding to the memory theme desired by the user among the at least one memory theme identifiers included in the memory creation interface, the user can directly click on this memory theme identifier without entering the memory theme in the memory theme input box, which simplifies the user operation and improves the efficiency of creating memory videos.

[0083] In some embodiments of the present application, before performing multi-layer perception processing on the splicing feature vectors corresponding to each multimedia file set in the multimedia database, the multimedia database needs to be generated first. Therefore, before step 120, the above-mentioned method may further include:

[0084] Cluster all multimedia files in the album application according to at least one clustering dimension to obtain at least one multimedia file set corresponding to different clustering dimensions respectively;

[0085] Generate a multimedia database based on at least one multimedia file set corresponding to different clustering dimensions respectively.

[0086] Among them, the clustering dimension can be a dimension preset for clustering all multimedia files in the album application, and at least one clustering dimension includes at least one of the following: shooting time, shooting location, shooting object, shooting content.

[0087] In some embodiments of the present application, there is at least one multimedia file set formed by different clustering dimensions. For at least one multimedia file set corresponding to a clustering dimension, the clustering attributes of the same multimedia file set are the same. The above clustering attributes may include at least one of the following: shooting time, shooting location, shooting object, shooting content, that is, for at least one multimedia file set corresponding to a certain clustering dimension, the clustering attributes of the multimedia files in a multimedia file set are the same.

[0088] For example, in at least one multimedia file set formed by clustering all multimedia files in the album application according to the shooting time dimension, the shooting time of the multimedia files in each multimedia file set is the same. For another example, in at least one multimedia file set formed by clustering all multimedia files in the album application according to the shooting location dimension, the shooting location of the multimedia files in each multimedia file set is the same.

[0089] In some embodiments of the present application, all multimedia files in the album application can be clustered according to at least one clustering dimension, at least one multimedia file set corresponding to different clustering dimensions can be obtained, and then a multimedia database can be generated according to at least one multimedia file set corresponding to different clustering dimensions respectively.

[0090] In the embodiments of the present application, by clustering all multimedia files in the album application according to at least one clustering dimension, at least one multimedia file set of different clustering dimensions can be obtained. In this way, when the user wants to create a memory video with different types of memory themes, a memory video corresponding to the type of memory theme required by the user can be obtained from at least one multimedia file set of different clustering dimensions, improving the recommendation diversity of the memory video.

[0091] In some embodiments of the present application, in order to improve the push accuracy of multimedia files, clustering all multimedia files in the album application according to at least one clustering dimension to generate at least one multimedia file set corresponding to different clustering dimensions respectively may specifically include:

[0092] Cluster each multimedia file in the photo album application according to the shooting time of each multimedia file to obtain at least one set of time multimedia files corresponding to different shooting times;

[0093] Cluster each multimedia file in the photo album application according to the shooting location of each multimedia file to obtain at least one set of location multimedia files corresponding to different shooting locations;

[0094] Cluster the shooting objects in each multimedia file in the photo album application to obtain at least one set of shooting object multimedia files corresponding to different shooting objects;

[0095] Classify each multimedia file in the photo album application according to the shooting content of each multimedia file to obtain at least one set of shooting content multimedia files corresponding to different shooting contents.

[0096] Among them, at least one set of time multimedia files can be obtained by clustering each multimedia file according to different shooting times. The shooting times of the multimedia files in the same set of time multimedia files are the same.

[0097] At least one set of location multimedia files can be obtained by clustering each multimedia file according to different shooting locations. The shooting locations of the multimedia files in the same set of location multimedia files are the same.

[0098] At least one set of shooting object multimedia files can be obtained by clustering each multimedia file according to different shooting objects. The shooting objects of the multimedia files in the same set of shooting object multimedia files are the same.

[0099] At least one set of shooting content multimedia files can be obtained by clustering each multimedia file according to different shooting contents. The shooting contents of the multimedia files in the same set of shooting content multimedia files are the same.

[0100] In some embodiments of the present application, each multimedia file can be clustered according to the shooting time of each multimedia file in the photo album application, so as to obtain at least one set of time multimedia files corresponding to different shooting times.

[0101] Continuing to refer to the above example, cluster Multimedia File 1, Multimedia File 2, Multimedia File 3, Multimedia File 4, Multimedia File 5, Multimedia File 6, Multimedia File 7, Multimedia File 8, Multimedia File 9, and Multimedia File 10 according to the shooting time dimension, and obtain as Figure 4The six sets of time-based multimedia files shown, where multimedia file 1 is the time-based multimedia file set 41 corresponding to the shooting time of September 28, 2023, multimedia file 2 is the time-based multimedia file set 42 corresponding to the shooting time of October 1, 2023, multimedia files 3, 4, and 5 are the time-based multimedia file set 43 corresponding to the shooting time of October 3, 2023, multimedia files 6 and 7 are the time-based multimedia file set 44 corresponding to the shooting time of October 6, 2023, multimedia file 8 is the time-based multimedia file set 45 corresponding to the shooting time of October 9, 2023, and multimedia files 9 and 10 are the time-based multimedia file set 46 corresponding to the shooting time of November 1, 2023.

[0102] In some embodiments of the present application, each multimedia file can be clustered according to the shooting location of each multimedia file in the album application, so that at least one location-based multimedia file set corresponding to different shooting locations can be obtained.

[0103] Continuing to refer to the above example, clustering multimedia files 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 according to the shooting location dimension, we get Figure 5 The three sets of location-based multimedia files shown, where multimedia files 1, 8, 9, and 10 are the location-based multimedia file set 51 corresponding to the shooting location of Tianjin, multimedia files 2, 3, 4, and 5 are the location-based multimedia file set 52 corresponding to the shooting location of Beijing, and multimedia files 6 and 7 are the location-based multimedia file set 53 corresponding to the shooting location of Tangshan.

[0104] In some embodiments of the present application, the shooting objects in each multimedia file can be clustered according to the shooting objects in each multimedia file in the album application, so that at least one shooting object-based multimedia file set corresponding to different shooting objects can be obtained.

[0105] Continuing to refer to the above example, clustering multimedia files 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 according to the shooting object dimension, we get Figure 6The two sets of multimedia files of the photographed objects shown, where multimedia files 1, 2, 3, 5, 9, and 10 are the set of multimedia files of the photographed object Xiaoming, corresponding to the set of multimedia files 61 of the photographed object, and multimedia files 3 and 4 are the set of multimedia files of the photographed object Xiaohong, corresponding to the set of multimedia files 62 of the photographed object.

[0106] In some embodiments of the present application, according to the shooting content in each multimedia file in the album application, each multimedia file can be classified according to the shooting content, so that at least one set of multimedia files corresponding to different shooting contents can be obtained.

[0107] Continuing to refer to the above example, clustering multimedia files 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 according to the shooting content dimension, we get Figure 7 The three sets of multimedia files of the shooting content shown, where multimedia files 5 and 6 are the set of multimedia files of the shooting content landscape, corresponding to the set of multimedia files 71 of the shooting content, multimedia files 7 and 10 are the set of multimedia files of the shooting content food, corresponding to the set of multimedia files 72 of the shooting content, and multimedia file 8 is the set of multimedia files of the shooting content pet dog, corresponding to the set of multimedia files 73 of the shooting content.

[0108] It should be noted that in the case where the multimedia file is a video, if there are multiple shooting dimensions in the video, the video is respectively used as the multimedia file of the shooting dimension. For example, in multimedia file 10 in the above example, there is both Xiaoming and food in this multimedia file 10, so this multimedia file 10 belongs to both the set of multimedia files of the photographed object and the set of multimedia files of the shooting content.

[0109] In the embodiments of the present application, by clustering each multimedia file according to the clustering dimensions of shooting time, shooting location, photographed object, and shooting content, sets of multimedia files corresponding to different clustering dimensions are obtained. In this way, when subsequently recommending recall videos with recall themes based on shooting time, shooting location, photographed object, and shooting content, the recommended multimedia files can be more adapted to the recall theme, improving the accuracy of multimedia file push.

[0110] In some embodiments of the present application, in order to further improve the accuracy of multimedia file push, generating a multimedia database according to at least one set of multimedia files respectively corresponding to different clustering dimensions may specifically include:

[0111] According to the shooting time of at least one set of time multimedia files, clustering at least one set of time multimedia files according to different time ranges to obtain at least one set of continuous time multimedia files;

[0112] Sort all the multimedia files in at least one location multimedia file set according to the shooting date to obtain a time-sorted multimedia file set;

[0113] Based on the shooting locations of each multimedia file in the time-sorted multimedia file set, use the shooting location with the highest frequency as the permanent residence;

[0114] Cluster the multimedia files in the time-sorted multimedia file set whose shooting locations are between the permanent residences to obtain at least one continuous location multimedia file set;

[0115] Generate a multimedia database based on at least one time multimedia file set, at least one location multimedia file set, at least one shooting object multimedia file set, at least one shooting content multimedia file set, at least one continuous time multimedia file set, and at least one continuous location multimedia file set.

[0116] Among them, the time range can be a pre-set time range with special meanings, and this time range can be in days, such as holidays, weekends, etc.

[0117] A continuous time multimedia file set can be obtained by clustering at least one time multimedia file set based on a time range.

[0118] The time-sorted multimedia file set can be the multimedia file set obtained by sorting all the multimedia files in at least one location multimedia file set according to the shooting date.

[0119] A continuous location multimedia file set can be obtained by clustering the multimedia files in the time-sorted multimedia file set whose shooting locations are between the permanent residences.

[0120] In some embodiments of the present application, at least one multi-time multimedia file set can be clustered according to different time ranges to obtain at least one continuous time multimedia file set with special date meanings.

[0121] Continue to refer to Figure 4 , 10.1 - 10.7 is the National Day holiday. Thus, 10.1 - 10.7 can be divided into a time range, and the time multimedia file set 42 corresponding to the shooting time of 2023.10.1, the time multimedia file set 43 corresponding to the shooting time of 2023.10.3, and the time multimedia file set 44 corresponding to the shooting time of 2023.10.6 can be clustered to form a continuous time multimedia file set for the National Day holiday in 2023.

[0122] In some embodiments of the present application, at least one set of location multimedia files can also be clustered according to the shooting location to form a set of multimedia files for consecutive locations. Specifically, all multimedia files in at least one set of location multimedia files can be sorted according to the shooting date first to obtain a time-sorted multimedia file set. As Figure 5 shown, the 3 sets of location multimedia files in Figure 5 are sorted according to the shooting date of each multimedia file to obtain a time-sorted multimedia file set: Multimedia File 1, Multimedia File 2, Multimedia File 3, Multimedia File 4, Multimedia File 5, Multimedia File 6, Multimedia File 7, Multimedia File 8, Multimedia File 9, and Multimedia File 10.

[0123] Then, according to the shooting location of each multimedia file in the time-sorted multimedia file set, the shooting location with the highest frequency of occurrence is taken as the permanent residence.

[0124] As shown in the above example, in the time-sorted multimedia file set: Multimedia File 1, Multimedia File 2, Multimedia File 3, Multimedia File 4, Multimedia File 5, Multimedia File 6, Multimedia File 7, Multimedia File 8, Multimedia File 9, and Multimedia File 10, the shooting location with the highest frequency of occurrence is Tianjin. Thus, it can be known that Tianjin is the permanent residence.

[0125] Then, the multimedia files in the time-sorted multimedia file set whose shooting locations are between the permanent residences can be formed into a cluster.

[0126] As shown in the above example, the multimedia files 2, 3, 4, 5, 6, and 7 in the time-sorted multimedia file set: Multimedia File 1, Multimedia File 2, Multimedia File 3, Multimedia File 4, Multimedia File 5, Multimedia File 6, Multimedia File 7, Multimedia File 8, Multimedia File 9, and Multimedia File 10 whose shooting locations are in Tianjin can be formed into a cluster to obtain a set of multimedia files for consecutive locations.

[0127] It should be noted that when the multimedia files in the time-sorted multimedia file set whose shooting locations are between the permanent residences are formed into a cluster, the multimedia files in the time-sorted multimedia file set whose shooting locations are between adjacent permanent residences are formed into a cluster.

[0128] For example, after the multimedia file 10, the photo album application further includes multimedia files 11, 12, and 13. Among them, multimedia file 11 is a solo photo of Xiaoming taken in Jinan on November 5, 2023, multimedia file 12 is a solo photo of Xiaoming taken in Zhengzhou on November 8, 2023, and multimedia file 13 is a solo photo of Xiaoming taken in Tianjin on November 15, 2023. Then, the multimedia file set sorted by time is: multimedia file 1, multimedia file 2, multimedia file 3, multimedia file 4, multimedia file 5, multimedia file 6, multimedia file 7, multimedia file 8, multimedia file 9, multimedia file 10, multimedia file 11, multimedia file 12, and multimedia file 13. The permanent residence is Tianjin. Multimedia files in the time-sorted multimedia file set whose shooting locations are between adjacent permanent residences form a cluster. That is, multimedia files 2, 3, 4, 5, 6, and 7 form a cluster to obtain a continuous location multimedia file set. In addition, multimedia files 11 and 12 also form a cluster to obtain a continuous location multimedia file set.

[0129] In the embodiments of the present application, by clustering the multimedia files in the time multimedia file set according to continuous time to form a continuous time multimedia file set with continuous time information, and clustering the multimedia files in the location multimedia file set according to continuous location to form a continuous location multimedia file set with continuous location information. In this way, when recommending recall videos of the recall theme based on continuous time or continuous location subsequently, the recommended multimedia files can be more adapted to the recall theme, further improving the accuracy of multimedia file push.

[0130] In some embodiments of the present application, in order to further improve the accuracy of multimedia file push, generating a multimedia database according to at least one time multimedia file set, at least one location multimedia file set, at least one shooting object multimedia file set, at least one shooting content multimedia file set, at least one continuous time multimedia file set, and at least one continuous location multimedia file set may specifically include:

[0131] Clustering the multimedia files in the multimedia file sets that appear in at least two different clustering dimensions to obtain at least one cross-information multimedia file set;

[0132] Generating a multimedia database according to at least one time multimedia file set, at least one location multimedia file set, at least one shooting object multimedia file set, at least one shooting content multimedia file set, at least one continuous time multimedia file set, at least one continuous location multimedia file set, and at least one cross-information multimedia file set.

[0133] Among them, the cross-information multimedia file set can be a multimedia file set formed by clustering the multimedia files that appear in at least two different clustering dimensions. The cross-clustering dimension attributes in the same cross-information multimedia file set are the same, that is, the multimedia files in the same cross-information multimedia file set are clustered based on the same cross-attribute. For example, the time-location cross multimedia file set is formed by crossing the time multimedia file set and the location multimedia file set, and the time-content cross multimedia file set is formed by crossing the time multimedia file set and the shooting content multimedia file set.

[0134] In some embodiments of the present application, the multimedia files of the multimedia file sets that appear in at least two different clustering dimensions can be clustered to obtain at least one cross-information multimedia file set.

[0135] Continuing to refer to the above example, multimedia file 7 appears in both Figure 5 the location multimedia file set 53 and Figure 7 the shooting content multimedia file set 72, then multimedia file 7 can be used as the cross-information multimedia file set of the location multimedia file set and the shooting content multimedia file set, and this multimedia file 7 as the cross-information multimedia file set can be used to represent the food in Tangshan.

[0136] Continuing to refer to the above example, multimedia file 6 appears in both Figure 5 the location multimedia file set 53 and Figure 7 the shooting content multimedia file set 71, then multimedia file 6 can be used as the cross-information multimedia file set of the location multimedia file set and the shooting content multimedia file set, and this multimedia file 6 as the cross-information multimedia file set can be used to represent the scenery in Tangshan.

[0137] Continuing to refer to the above example, multimedia file 3 appears in both Figure 6 the multimedia file set 61 of the shooting object Xiaoming and the multimedia file set 62 of the shooting object Xiaohong shown, then multimedia file 3 can be used as the cross-information multimedia file set of the multimedia file set of the shooting object Xiaoming and the multimedia file set of the shooting object Xiaohong, and this multimedia file 3 as the cross-information multimedia file set can be used to represent the group photo of Xiaoming and Xiaohong.

[0138] After obtaining at least one cross-information multimedia file set, the above-mentioned time multimedia file set, location multimedia file set, shooting object multimedia file set, shooting content multimedia file set, at least one continuous time multimedia file set, at least one continuous location multimedia file set, and at least one cross-information multimedia file set can be combined to generate a multimedia database.

[0139] In an embodiment of the present application, by clustering the multimedia files in the multimedia file sets that appear in at least two different clustering dimensions, at least one cross-information multimedia file set is obtained. Thus, when subsequently recommending recall videos based on the recall theme of the cross-information, the recommended multimedia files can be more adapted to the recall theme, further improving the accuracy of multimedia file push.

[0140] In some embodiments of the present application, in order to make the recall theme of the pushed recall videos clear and the relevance of the multimedia files in the pushed recall videos high, step 120 may specifically include:

[0141] Through a data push model, image encoding is performed on each multimedia file set in the multimedia database to obtain the image feature vector of each multimedia file set;

[0142] Through the data push model, the image feature vector corresponding to the multimedia file set and the text encoding vector are spliced to obtain the spliced feature vector of each multimedia file set;

[0143] Based on the data push model, multi-layer perception processing is performed on the spliced feature vector corresponding to each multimedia file set to obtain the text feature vector corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input.

[0144] Among them, the image feature vector of a multimedia file set may be the feature vector obtained after image encoding of the multimedia file set.

[0145] In some embodiments of the present application, through the image encoding layer in the data push model, image encoding may be performed on each multimedia file set in the multimedia database, so as to obtain the image feature vector of each multimedia file set.

[0146] Then, based on the splicing layer in the data push model, the image feature vector corresponding to the multimedia file set and the text encoding vector are spliced to obtain the spliced feature vector of each multimedia file set.

[0147] Furthermore, based on the multi-layer perception layer in the data push model, multi-layer perception processing is performed on the spliced feature vector corresponding to each multimedia file set to obtain the text feature vector corresponding to each multimedia file set respectively and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input.

[0148] In an embodiment of the present application, each multimedia file set in the multimedia database is subjected to image encoding through a data push model to obtain an image feature vector of each multimedia file set. Then, through the data push model, the image feature vector corresponding to the multimedia file set and the text encoding vector are spliced to obtain a spliced feature vector of each multimedia file set. Furthermore, based on the data push model, a multi-layer perception process is performed on the spliced feature vector corresponding to each multimedia file set, so as to obtain a text feature vector corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the memory theme created by the first input. In this way, the memory theme of the pushed memory video is clear, and the relevance of the multimedia files in the pushed memory video is high.

[0149] In some embodiments of the present application, in order to improve the pushing efficiency of the memory video, the process of performing image encoding on each multimedia file set in the multimedia database through the data push model to obtain an image feature vector of each multimedia file set may specifically include:

[0150] According to the reference segmentation size, the image corresponding to each multimedia file in each multimedia file set is segmented to obtain at least two image blocks of each image;

[0151] The encoding vector of each image block corresponding to each image and the first position encoding vector are spliced to obtain a block encoding vector of each image block;

[0152] The block encoding vectors of each image block are fused to obtain an image encoding vector of each image;

[0153] According to the image encoding vector of each image, the features of each image are extracted to obtain a feature vector of each image;

[0154] The feature vector of the image corresponding to each multimedia file in the multimedia file set is spliced with the second position encoding vector to obtain an image feature vector of each multimedia file set.

[0155] Wherein, the reference segmentation size may be a size preset for segmenting each multimedia file set. The reference segmentation size may be a size of 10*10*3 - 18*18*3. For example, the reference segmentation size may be 16*16*3 or 14*14*3.

[0156] It should be noted that in the embodiments of the present application, the reference segmentation size used for segmenting each multimedia file set is the same. Of course, the reference segmentation size used for segmenting each multimedia file set may also be different. Whether the reference segmentation size used for segmenting each multimedia file set is the same specifically can be set according to user needs and is not limited in the embodiments of the present application.

[0157] An image block can be an image block obtained by splitting the image corresponding to a multimedia file according to a reference splitting size.

[0158] It should be noted that in the case where the multimedia file is a photo file, the image corresponding to the multimedia file is the photo file. In the case where the multimedia file is a video file, the image corresponding to the multimedia file is a preset number of image frames selected from the video file. The preset number here can be determined according to the number of all image frames included in the video file. For example, the preset number can take a value of 10%-15% of all image frames in the video file. For example, the preset number can be 10% of all image frames in the video file.

[0159] For any image block, the encoding vector of the image block can be a vector obtained by encoding it.

[0160] For any image block, the first position encoding vector of the image block can be the position vector of the image block in its corresponding image.

[0161] For any image block, the block feature vector of the image block can be a vector obtained by concatenating the encoding vector of the image block and its first position encoding vector.

[0162] For the image corresponding to any multimedia file, its corresponding image encoding vector can be a vector obtained by fusing the block feature vectors of at least one image block that makes up the image.

[0163] For the image corresponding to any multimedia file, the feature vector corresponding to it can be a vector obtained by extracting the features of the image.

[0164] For any multimedia file in a certain multimedia file set, its corresponding second position encoding vector can be the position vector of it in the multimedia file set.

[0165] For any multimedia file set, its corresponding image feature vector can be a vector obtained by concatenating the feature vectors of the images corresponding to each multimedia file in the multimedia file set, and the second position encoding vectors of each multimedia file.

[0166] In some embodiments of the present application, the image corresponding to each multimedia file in each multimedia file set can be segmented according to the reference splitting size to obtain at least two image blocks of each image.

[0167] In one example, taking a certain multimedia file in any multimedia file set as an example, refer to Figure 8, taking the multimedia file 3 in the time multimedia file set 43 in the above example and taking the reference segmentation size of 16*16*3 as an example, the multimedia file 3 is segmented according to the reference segmentation size of 16×16×3, and 14×14 image blocks can be obtained.

[0168] It should be noted that for the sake of simplicity, Figure 8 not all the image blocks of the multimedia file 3 are shown, only 4 image blocks, namely image block 81, image block 82, image block 83 and image block 84, are shown. In the following examples, Figure 8 the 4 image blocks, namely image block 81, image block 82, image block 83 and image block 84 shown are taken as examples for description.

[0169] Then, the encoding vector of each image block corresponding to each image is concatenated with the first position encoding vector to obtain the block encoding vector of each image block, that is, the encoding vector of each image block is concatenated with its position encoding vector to obtain the block encoding vector of this image block.

[0170] Continue to refer to Figure 8 , the encoding vector corresponding to image block 81 is encoding vector 1, the first position encoding vector corresponding to image block 81 is position vector 1, the encoding vector corresponding to image block 82 is encoding vector 2, the first position encoding vector corresponding to image block 82 is position vector 2, the encoding vector corresponding to image block 83 is encoding vector 3, the first position encoding vector corresponding to image block 83 is position vector 3, the encoding vector corresponding to image block 84 is encoding vector 4, and the first position encoding vector corresponding to image block 84 is position vector 4. Then, the encoding vector 1 corresponding to image block 81 and position vector 1 can be concatenated to obtain the block encoding vector K1 of image block 81, the encoding vector 2 corresponding to image block 82 and position vector 2 can be concatenated to obtain the block encoding vector K2 of image block 82, the encoding vector 3 corresponding to image block 83 and position vector 3 can be concatenated to obtain the block encoding vector K3 of image block 83, and the encoding vector 4 corresponding to image block 84 and position vector 4 can be concatenated to obtain the block encoding vector K4 of image block 84.

[0171] Then, the block encoding vectors of each image block are fused to obtain the image encoding vector of each image. Continue to refer to Figure 8 , the block encoding vector K1 of image block 81, the block encoding vector K2 of image block 82, the block encoding vector K3 of image block 83 and the block encoding vector K4 of image block 84 are fused to obtain the image feature vector T1 of the multimedia file 1.

[0172] Then, the features of each image can be extracted according to the image encoding vector of each image to obtain the feature vector of each image. As Figure 8As shown, according to the image feature vector T1 of the multimedia file 3, the features of the multimedia file 3 are extracted to obtain the feature vector Z1 of the multimedia file 3.

[0173] In this way, the images corresponding to each multimedia file in the time multimedia file set 43 are all processed according to the processing process of the multimedia file 3 shown in Figure 8 so as to obtain the image feature vectors of the images corresponding to each multimedia file in the time multimedia file set 43.

[0174] Then, the feature vectors of the images corresponding to each multimedia file in the multimedia file set are concatenated with the second position encoding vector to obtain the image feature vector of each multimedia file set. Specifically, the feature vector of the image corresponding to each multimedia file and the second position encoding vector can be concatenated first to obtain the file feature vector of the multimedia file, and then the file feature vectors of each multimedia file are fused to obtain the image feature vector of each multimedia file set.

[0175] Continuing to refer to the above example, if the feature vector of the multimedia file 3 is Z1 and its second position encoding vector is S1. The image feature vector of the multimedia file 4 in the time multimedia file set 43 is Z2 and its second position encoding vector is S2. The image feature vector of the multimedia file 5 in the time multimedia file set 43 is Z3 and its second position encoding vector is S3.

[0176] Then, the feature vector Z1 of the multimedia file 3 and the second position encoding vector S1 are concatenated to obtain the file feature vector R1 of the multimedia file 3. The feature vector Z2 of the multimedia file 4 and the second position encoding vector S2 are concatenated to obtain the file feature vector R2 of the multimedia file 4. The feature vector Z3 of the multimedia file 5 and the second position encoding vector S3 are concatenated to obtain the file feature vector R3 of the multimedia file 5.

[0177] Then, the file feature vector R1 of the multimedia file 3, the file feature vector R2 of the multimedia file 4, and the file feature vector R3 of the multimedia file 5 are fused to obtain the image feature vector of the time multimedia file set 43.

[0178] For each multimedia file set, its image feature vector can be calculated according to the calculation method of the image feature vector of the time multimedia file set 43 in the above example, which will not be elaborated here.

[0179] In the embodiments of the present application, by quantifying each multimedia file set, an image feature vector of each multimedia file set is obtained. When processing the multimedia file set subsequently, the image feature vector of the multimedia file set can be processed, without the need to process each multimedia file in the multimedia file set, reducing the processing amount of the multimedia file set, improving the processing efficiency of the multimedia file set, and further improving the pushing efficiency of the recalled video.

[0180] In some embodiments of the present application, in order to improve the processing efficiency of each multimedia file, before segmenting the image corresponding to each multimedia file in each multimedia file set according to the reference segmentation size to obtain at least two image blocks corresponding to each image, the methods involved above may further include:

[0181] Adjust the size of the image corresponding to each multimedia file in the multimedia file set according to the reference size;

[0182] Perform normalization processing on each image whose size is adjusted to the reference size to obtain each normalized image;

[0183] The step of segmenting the image corresponding to each multimedia file in each multimedia file set according to the reference segmentation size to obtain at least two image blocks corresponding to each image may specifically include:

[0184] Perform image segmentation processing on each normalized image in the multimedia file set according to the reference segmentation size to obtain at least two image blocks corresponding to each normalized image.

[0185] Among them, the reference size may be a size preset for adjusting the image corresponding to each multimedia file. Specifically, the reference size may be determined based on the size of the image that the data pushing model can process, that is, the reference size is the size of the image that the data pushing model can process. For example, the reference size may be 224×224×3.

[0186] In some embodiments of the present application, since the sizes of each multimedia file may be different and the size of each multimedia file may not conform to the size of the image that the data pushing model can process, before segmenting the image corresponding to each multimedia file in each multimedia file set according to the reference segmentation size, it is first necessary to adjust the size of each multimedia file to the reference size.

[0187] Continue to refer to Figure 8, taking the size of multimedia file 3 as 512*512*3 and the reference size as 224*224*3 as an example, before splitting multimedia file 3 according to the reference splitting size of 16*16*3, first adjust the size of multimedia file 3 to 224*224*3.

[0188] Then, perform normalization processing on each image whose size is adjusted to the reference size to obtain each normalized image, and then perform image segmentation processing on each normalized image in the multimedia file set according to the reference splitting size to obtain at least two image blocks corresponding to each normalized image.

[0189] In some embodiments of the present application, performing normalization processing on each image whose size is adjusted to the reference size may be to normalize the pixel value of each pixel point of each image whose size is adjusted to the reference size to between 0 and 1. For example, it may be to divide the pixel value of each pixel point of each image whose size is adjusted to the reference size by 255, so as to normalize the pixel value of each pixel point of each image whose size is adjusted to the reference size to between 0 and 1.

[0190] Continue to refer to Figure 8 , after adjusting the size of multimedia file 3 from 512*512*3 to 224*224*3, divide the pixel value of each pixel point in multimedia file 3 of 224*224*3 by 255 to obtain the normalized multimedia file 3, and then split the normalized multimedia file 3 according to the reference splitting size of 16*16*3.

[0191] In the embodiments of the present application, by adjusting the size of the image corresponding to each multimedia file to the reference size matching the data push model, it is convenient for the data push model to process the image corresponding to each multimedia file. In addition, performing normalization processing on each image whose size is adjusted to the reference size can simplify the processing amount of each image whose size is adjusted to the reference size, reduce the calculation amount of each multimedia file, and improve the processing efficiency of each multimedia file.

[0192] In some embodiments of the present application, in order to be able to recommend memory videos with high user perception, before splicing the image feature vector and text encoding vector corresponding to the multimedia file set through the data push model to obtain the spliced feature vector of the multimedia file set, the methods involved above may further include:

[0193] Generate text description information for each multimedia file set according to the shooting information of each multimedia file in each multimedia file set;

[0194] Convert the text description information into text features to obtain the text encoding vector of each multimedia file set.

[0195] Among them, the shooting information of the multimedia file may include the shooting content, shooting time, shooting location, etc. of the multimedia file.

[0196] For any multimedia file set, its corresponding text description information may be information generated according to the shooting information of each multimedia file in the multimedia file set for describing the multimedia file set. For example, for the continuous-time multimedia file set of the National Day holiday composed of multimedia file 2, multimedia file 3, multimedia file 4, multimedia file 5, multimedia file 6, and multimedia file 7 in the above example, if this continuous-time multimedia file set is a picture set of the National Day in 2023, then the text description information of this continuous-time multimedia file set may be "Picture set of the National Day in 2023".

[0197] In some embodiments of the present application, for each multimedia file set, the text description information of the multimedia file set can be generated according to the shooting information of each multimedia file in the multimedia file set, and then the text description information of the multimedia file set is converted into text features to obtain the text encoding vector of the multimedia file set.

[0198] In the embodiments of the present application, for each multimedia file set, the text encoding vector of the multimedia file set is determined according to the shooting information of each multimedia file in the multimedia file set. The text encoding vector of the multimedia file set determined in this way is more in line with the shooting information of each multimedia file in the multimedia file set. Therefore, the recalled video pushed based on the text encoding vector of the multimedia file set will also be more in line with the shooting information of the multimedia file, that is, the pushed recalled video is in line with the shooting information of the multimedia file. That is to say, through the solution of the present application, a recalled video with a high user perception can be recommended.

[0199] In some embodiments of the present application, before using the data push model, the reference data push model needs to be trained first to obtain the data push model. Therefore, before step 120, the above-mentioned method may further include:

[0200] Obtain at least two training samples;

[0201] Input at least two training samples into the reference data push model, and output the predicted text feature vector of each sample multimedia file set and the predicted relevance of each sample multimedia file in each sample multimedia file set;

[0202] The reference data push model is trained according to the degree of difference between the predicted text feature vectors and the sample text feature vectors of each sample multimedia file set, and the degree of difference between the predicted relevance and the sample relevance of each sample multimedia file in each sample multimedia file set, to obtain a data push model.

[0203] Among them, each training sample may include the sample text feature vectors of each sample multimedia file set in the sample multimedia database, and the sample relevance of each sample multimedia file in each sample multimedia file set.

[0204] The above-mentioned sample multimedia database may be a pre-constructed multimedia database for training the reference data push model. The construction process of this sample multimedia database is the same as that of the multimedia database in the above embodiment, and will not be elaborated here.

[0205] The sample multimedia file set may be a multimedia file set included in the sample multimedia database.

[0206] The sample text feature vector is the text feature vector of the sample multimedia file set. The determination of the sample text feature vector of each sample multimedia file set here is the same as the determination method of the text feature vector of each multimedia file set in the above embodiment, and will not be elaborated here.

[0207] The sample relevance of each sample multimedia file may be the relevance between the sample multimedia file and the theme of each sample multimedia file set. The theme of each sample multimedia file set here is determined based on the sample file feature vector of the sample multimedia file set.

[0208] The reference data push model may be the model before training the data push model.

[0209] For each sample multimedia file set, the predicted text feature vector of the sample multimedia file set may be the text feature vector of the sample multimedia file set output after inputting the sample multimedia file set into the reference data push model.

[0210] For each sample multimedia file set, the predicted relevance measure of each sample multimedia file in the sample multimedia file set may be the relevance between each sample multimedia file in the sample multimedia file set and the theme represented by the sample multimedia file set output after inputting the sample multimedia file set into the reference data push model.

[0211] In some embodiments of the present application, at least two obtained training samples can be input into a reference data push model. Based on the reference data push model, the at least two training samples are processed to output the predicted text feature vectors of each sample multimedia file set and the predicted relevance of each sample multimedia file in each sample multimedia file set.

[0212] Then, based on the degree of difference between the predicted text feature vectors of each sample multimedia file set and the sample text feature vectors, and the degree of difference between the predicted relevance of each sample multimedia file in each sample multimedia file set and the sample relevance, the reference data push model is trained to obtain a data push model.

[0213] It should be noted that the process of the above reference data push model processing at least two training samples to output the predicted text feature vectors of each sample multimedia file set and the predicted relevance of each sample multimedia file in each sample multimedia file set is the same as the process in step 120 of the above embodiment, where the data push model performs multi-layer perception processing on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database to obtain the text feature vectors corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the theme represented by the text feature vectors. This will not be elaborated here.

[0214] In the embodiments of the present application, the reference data push model is trained with at least two training samples to obtain a data push model. In this way, the recall videos corresponding to the recall theme created with the first input can be directly pushed to the user based on this data push model, improving the push efficiency of the recall videos corresponding to the recall theme created with the first input.

[0215] In some embodiments of the present application, in order to improve the robustness of the data push model, the training of the reference data push model to obtain a data push model according to the degree of difference between the predicted text feature vectors of each sample multimedia file set and the sample text feature vectors, and the degree of difference between the predicted relevance of each sample multimedia file in each sample multimedia file set and the sample relevance may specifically include:

[0216] Based on the degree of difference between the predicted text feature vectors of each sample multimedia file set and the sample text feature vectors, the first loss function value of the reference data push model is obtained;

[0217] Based on the degree of difference between the predicted relevance of each sample multimedia file in each sample multimedia file set and the sample relevance, the second loss function value of the reference data push model is obtained;

[0218] Train a reference data push model according to the first loss function value and the second loss function value to obtain a data push model.

[0219] Among them, the first loss function value can be used to characterize the distance between the predicted text feature vector and the sample text feature vector. Specifically, the first loss function value can be obtained according to the difference degree between the predicted text feature vector and the sample text feature vector of each sample multimedia file set.

[0220] The second loss function value can be used to characterize the distance between the sample multimedia file and the theme corresponding to each sample multimedia file set. Specifically, the second loss function value can be obtained according to the difference degree between the predicted relevance and the sample relevance of each sample multimedia file in each sample multimedia file set.

[0221] In some embodiments of the present application, for each sample multimedia file set, according to the difference degree between the predicted text feature vector and the sample text feature vector of the sample multimedia file set, the first loss function value of the reference data push model can be obtained according to the following formula (1).

[0222]

[0223] In the above formula (1), i is the i-th sample multimedia file set, N is the number of sample multimedia file sets, j is the j-th predicted sample multimedia file set, K is the length of the predicted text feature vector of the sample multimedia file set, G is the sample text feature vector, and P is the predicted text feature vector of the sample multimedia file set.

[0224] In some embodiments of the present application, for each sample multimedia file set, the second loss function value of the reference data push model can be obtained according to the difference degree between the predicted relevance and the sample relevance of each sample multimedia file in the sample multimedia file set according to the following formula (2).

[0225]

[0226] In the above formula (2), M is the number of sample multimedia files in a sample multimedia file set, i is the i-th sample multimedia file set, N is the number of sample multimedia file sets, and j is the j-th sample multimedia file in a sample multimedia file set.

[0227] It should be noted that when training the reference data push model, the number of sample multimedia files in each sample multimedia file set in the training samples can be the same or different. To improve the training efficiency of the reference data push model, in the embodiments of the present application, the number of sample multimedia files in each sample multimedia file set is the same, for example, all are 50. If the number of sample multimedia files in a certain sample multimedia file set is less than 50, it can be directly filled with images with pixel values of 0.

[0228] In the embodiments of the present application, the reference data push model is trained by using the first loss function value representing the distance between the predicted text feature vector and the sample text feature vector, and the second loss function value representing the distance between the sample multimedia file and the theme corresponding to each sample multimedia file set, rather than training the reference data push model only with a single loss function value, which improves the robustness of the data push model.

[0229] In some embodiments of the present application, to improve the push accuracy of the recalled video, step 130 may specifically include:

[0230] According to the text feature vector corresponding to each multimedia file set, obtain the text push information of each multimedia file set respectively;

[0231] According to the recalled theme name corresponding to each multimedia file set, select at least one multimedia file set that matches the recalled theme created by the first input from all multimedia file sets in the multimedia database as the alternative multimedia file set;

[0232] According to the relevance between each multimedia file in the alternative multimedia file set and the recalled theme created by the first input, obtain the alternative multimedia files;

[0233] Output the recalled video according to the alternative multimedia files.

[0234] Among them, for any multimedia file set, its text push information can be obtained by parsing according to its corresponding text feature vector. For example, the text decoder in the data push model can be used to parse the text feature vector to obtain the text push information.

[0235] It should be noted that the text push information of a multimedia file set may include the recalled theme name corresponding to the multimedia file set.

[0236] Continuing to refer to the above example, taking the continuous-time multimedia file set of the National Day holiday in 2023, which consists of multimedia file 2, multimedia file 3, multimedia file 4, multimedia file 5, multimedia file 6, and multimedia file 7, as an example, after parsing the text feature vectors of this multimedia file set based on the text decoder, the text push information of this multimedia file set can include "Memory theme: National Day journey in 2023".

[0237] Similarly, for other multimedia file sets in the multimedia database, the text push information of each multimedia file set can be obtained according to the processing method of the continuous-time multimedia file set of the National Day holiday in 2023, which consists of multimedia file 2, multimedia file 3, multimedia file 4, multimedia file 5, multimedia file 6, and multimedia file 7.

[0238] For example, Figure 4 the text push information of the time multimedia file set 41 is "Memory theme: Play file set on September 28, 2023", Figure 4 the text push information of the time multimedia file set 42 is "Memory theme: Play file set on October 1, 2023", Figure 4 the text push information of the time multimedia file set 43 is "Memory theme: Play file set on October 3, 2023", Figure 4 the text push information of the time multimedia file set 44 is "Memory theme: Play file set on October 6, 2023", Figure 4 the text push information of the time multimedia file set 45 is "Memory theme: Play file set on October 9, 2023", Figure 4 the text push information of the time multimedia file set 46 is "Memory theme: Play file set on November 1, 2023".

[0239] Figure 5 the text push information of the location multimedia file set 51 is "Memory theme: Tianjin play file set", Figure 5 the text push information of the location multimedia file set 52 is "Memory theme: Beijing play file set", Figure 5 the text push information of the location multimedia file set 53 is "Memory theme: Tangshan play file set".

[0240] Figure 6 the text push information of the shooting object multimedia file set 61 is "Memory theme: Xiaoming file set", Figure 6 the text push information of the shooting object multimedia file set 62 is "Memory theme: Xiaohong file set".

[0241] Figure 7The text push information of the multimedia file set 71 of the shooting content is "Memory theme: Landscape file set", Figure 7 The text push information of the multimedia file set 72 of the shooting content is "Memory theme: Food file set", Figure 7 The text push information of the multimedia file set 73 of the shooting content is "Memory theme: Pet dog file set".

[0242] The text push information of the continuous-time multimedia file set formed by multimedia file 2, multimedia file 3, multimedia file 4, multimedia file 5, multimedia file 6 and multimedia file 7 is "Memory theme: 2023 National Day travel file set".

[0243] The text push information of the continuous-location multimedia file set formed by multimedia file 2, multimedia file 3, multimedia file 4, multimedia file 5, multimedia file 6 and multimedia file 7 is "Memory theme: 2023 National Day trip to Beijing and Tangshan".

[0244] The text push information of the cross-information multimedia file set formed by multimedia file 3 is "Memory theme: Group photo of Xiaoming and Xiaohong".

[0245] The alternative multimedia file set can be at least one multimedia file set selected from all multimedia atlases in the multimedia database and matching the memory theme created by the first input.

[0246] The alternative multimedia file can be at least one multimedia file in the alternative multimedia file set with a relevance greater than the relevance threshold.

[0247] The relevance threshold can be a threshold of the relevance between each multimedia file in the pre-set alternative multimedia file set and the memory theme of the alternative multimedia file set. The value range of the relevance threshold can be 70%-100%. For example, the relevance threshold can be set to 80%. The specific value of the relevance threshold can be set according to user needs and is not limited in the embodiments of the present application.

[0248] In some embodiments of the present application, the text push information of each multimedia file set can be obtained respectively according to the text feature vector corresponding to each multimedia file set, and then at least one multimedia file set matching the memory theme created by the first input can be selected from all multimedia atlases in the multimedia database as the alternative multimedia file set according to the memory theme name in the text push information of each multimedia file set.

[0249] Continuing to refer to the above example, taking the memory theme created by the first input as "2023 National Day trip to Beijing" as an example, according to Figure 4 the text push information of the at least one time multimedia file set shown, Figure 5The text push information of at least one set of location multimedia files shown, Figure 6 the text push information of at least one set of subject multimedia files shown, and Figure 7 the text push information of at least one set of shooting content multimedia files shown, as well as the text push information of the continuous location multimedia file set formed by multimedia files 2, 3, 4, 5, 6, and 7, and the text push information of the cross-information multimedia file set of the group photo of Xiaoming and Xiaohong formed by multimedia file 3, from Figure 4 at least one set of time multimedia files shown, Figure 5 at least one set of location multimedia files shown, Figure 6 at least one set of subject multimedia files shown, and Figure 7 at least one set of shooting content multimedia files shown, as well as the continuous location multimedia file set formed by multimedia files 2, 3, 4, 5, 6, and 7, and the cross-information multimedia file set of the group photo of Xiaoming and Xiaohong formed by multimedia file 3, select the multimedia file sets that match "Travel in Beijing during the National Day in 2023", that is, Figure 5 select the location multimedia file set 52 in it, and the continuous location multimedia file set formed by multimedia files 2, 3, 4, 5, 6, and 7 as the alternative multimedia file sets.

[0250] Then, according to the relevance between each multimedia file in the alternative multimedia file sets and the memory theme created by the first input, obtain the alternative multimedia files, that is, select the multimedia files with a relevance greater than the similarity threshold between the memory theme created by the first input from all the multimedia files in the alternative multimedia file sets as the alternative multimedia files. Then, the memory video can be output according to the alternative multimedia files.

[0251] Continuing to refer to the above example, taking the relevance threshold as 80% as an example, Figure 5After selecting the medium location multimedia file set 52 and the continuous location multimedia file sets formed by multimedia file 2, multimedia file 3, multimedia file 4, multimedia file 5, multimedia file 6, and multimedia file 7 as the alternative multimedia file sets, if the relevance between multimedia file 2 and the recall theme created by the first input is 90%, the relevance between multimedia file 3 and the recall theme created by the first input is 85%, the relevance between multimedia file 4 and the recall theme created by the first input is 88%, the relevance between multimedia file 5 and the recall theme created by the first input is 90%, the relevance between multimedia file 6 and the recall theme created by the first input is 71%, and the relevance between multimedia file 7 and the recall theme created by the first input is 75%, then it can be determined that multimedia file 2, multimedia file 3, multimedia file 4, and multimedia file 5 are the alternative multimedia files.

[0252] In an embodiment of the present application, by selecting at least one multimedia file set that matches the recall theme created by the first input from all multimedia file sets in the multimedia database as the alternative multimedia file set, and thus selecting alternative multimedia files from the alternative multimedia files, rather than selecting alternative multimedia files from all multimedia file sets in the multimedia database, the selection time of alternative multimedia files is saved, and the multimedia files in the recall video are all multimedia files with a high relevance to the recall theme created by the first input. In this way, the recalled video pushed is more in line with the recall theme created by the user's first input, improving the accuracy of the recalled video push.

[0253] In some embodiments of the present application, in order to increase the viewing interest of alternative multimedia files, the outputting of a recall video according to the alternative multimedia files may specifically include:

[0254] Based on at least one reference transition graph, splicing every two adjacent multimedia files in the alternative multimedia files to output a recall video.

[0255] Among them, the reference transition graph may be an image inserted between the previous multimedia file and the subsequent multimedia file when every two adjacent multimedia files are displayed, which is pre-set. Different transition effects can be presented for every two adjacent multimedia files based on different reference transition graphs. The different transition effects presented by the reference transition graph may include, but are not limited to: fade-in and fade-out, switching, sliding, zooming, etc.

[0256] In some embodiments of the present application, every two adjacent multimedia files in the alternative multimedia files can be spliced based on at least one reference transition graph. In this way, transition effects can be added between every two adjacent multimedia files to obtain a recall video.

[0257] In some embodiments of the present application, after splicing every two adjacent multimedia files in the alternative multimedia files based on at least one reference transition map, an audio can be added to the alternative multimedia files with added transition effects, so that the resulting memory video has background audio.

[0258] In the embodiments of the present application, by splicing every two adjacent multimedia files in the alternative multimedia files based on at least one reference transition map, a memory video with transition effects can be obtained, which can increase the viewing interest of the alternative multimedia files.

[0259] The technical solution of the embodiments of the present application can also be applied to the scenario of actively pushing memory videos in the photo album application according to the user's intention. For example, in the user's photo album application, there are multimedia files such as photos and videos taken by the user. If it is now the National Day period, it can be recognized that the user's intention is to view the multimedia file set of previous National Days, and then a memory video of previous National Day trips can be actively generated in the photo album application according to the user's intention.

[0260] The following will combine the drawings and explain in detail the multimedia file processing method provided by the embodiments of the present application through specific embodiments and their application scenarios.

[0261] Figure 9 It is a schematic flowchart of a multimedia file processing method provided by the embodiments of the present application. The execution subject of this multimedia file processing method can be an electronic device, which can be but is not limited to a PC, a smart phone, a tablet computer, a PDA, etc.

[0262] As Figure 9 shown, the multimedia file processing method provided by the embodiments of the present application may include step 910-step 930.

[0263] Step 910, obtain the viewing intention of the user for the multimedia files in the photo album application.

[0264] Among them, the viewing intention may be the intention of the user to view the multimedia files in the photo album application.

[0265] Step 920, through a data push model, perform multi-layer perception processing on the splicing feature vectors corresponding to each multimedia file set in the multimedia database to obtain the text feature vectors corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the theme represented by the viewing intention.

[0266] Step 930, based on the text feature vectors and the relevance, output a memory video corresponding to the theme represented by the viewing intention.

[0267] In an embodiment of the present application, when pushing a recall video containing multimedia files, the push can be performed according to the user's viewing intention for the multimedia files. In this way, according to the user's needs, the recall videos required by the user can be actively pushed according to the user's intention, improving the flexibility of pushing recall videos. In addition, when pushing recall videos, the splicing feature vectors corresponding to each multimedia file set in the multimedia database can be processed by a multi-layer perception through a data push model to obtain the text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the theme represented by the viewing intention. Then, based on the text feature vectors and relevance, recall videos corresponding to the theme represented by the viewing intention are output, rather than randomly pushing them to the user, improving the accuracy of pushing recall videos.

[0268] In some embodiments of the present application, the implementation manners of step 920-step 930 are the same as those of step 120-step 130 in the above embodiments, and will not be elaborated here.

[0269] In some embodiments of the present application, in order to improve the flexibility of obtaining the user's viewing intention for multimedia files, step 910 may specifically include:

[0270] When the system time information of the electronic device matches the reference date type, obtain the user's viewing intention for the multimedia files according to the system time information;

[0271] Or,

[0272] Obtain the user's viewing intention for the multimedia files according to the user's behavior data within the reference time period.

[0273] Among them, the reference date type may be a special date type such as a holiday, birthday, anniversary, etc.

[0274] The reference time period may be a time period set in advance before the current time, for example, it may be within 1 month before the current time.

[0275] In some embodiments of the present application, when the system time information of the electronic device matches the reference date type, the user's viewing intention for the multimedia files can be obtained according to the system time information.

[0276] In one example, if the system time information of the electronic device is October 1st, which is the National Day holiday, the system infers that the user may want to view multimedia files of previous National Day holidays, that is, the user's viewing intention for the multimedia files is to view multimedia files of previous National Day holidays.

[0277] In another example, if the system time information of the electronic device is March 5th, which is the user's birthday at this time, the system infers that the user may want to view the multimedia files on their birthday in previous years, that is, the user's viewing intention for the multimedia files is to view the multimedia files on their birthday in previous years.

[0278] In some embodiments of the present application, the viewing intention of the user for the multimedia files can be obtained according to the behavior data of the user within the reference time period.

[0279] In one example, if the user frequently views the multimedia files of themselves and their baby within the past month, it can be inferred that the user wants to view the multimedia files of themselves and their baby, that is, the user's viewing intention for the multimedia files is that the user wants to view the multimedia files of themselves and their baby.

[0280] In the embodiments of the present application, by according to the system time information of the electronic device or according to the behavior data of the user within the reference time period, the viewing intention of the user for the multimedia files can be obtained, which improves the flexibility of obtaining the viewing intention of the user for the multimedia files.

[0281] In some embodiments of the present application, in order to further improve the flexibility of obtaining the viewing intention of the user for the multimedia files, the determining the viewing intention of the user for the multimedia files according to the behavior data of the user within the reference time period may specifically include:

[0282] Obtaining the viewing intention of the user for the multimedia files according to the multimedia files taken by the user within the first reference time period;

[0283] Or,

[0284] Obtaining the viewing intention of the user for the multimedia files according to the operation information of the multimedia files in the album application by the user within the second reference time period.

[0285] Wherein, the first reference time period may be a time period before the preset current time, for example, it may be within 1 month before the current time.

[0286] The second reference time period may be a time period before the preset current time, for example, it may be within 1.5 months before the current time.

[0287] It should be noted that the first reference time period and the second reference time period may be the same or different, and specifically can be set according to the user's needs, which is not limited in the embodiments of the present application.

[0288] The operation information of the multimedia files in the album application may be at least one of viewing, sharing, collecting, editing, etc. of the multimedia files.

[0289] In some embodiments of the present application, the viewing intention of the user for the multimedia file can be obtained according to the multimedia files taken by the user within the first reference time period.

[0290] In one example, taking the first reference time period as within one month before the current time, if the user has taken many photos of their baby in the recent month, it can be inferred that the user wants to view the multimedia files of their baby, that is, the viewing intention of the user for the multimedia file is to view the multimedia files of their baby.

[0291] In some embodiments of the present application, the viewing intention of the user for the multimedia file can be obtained according to the operation information of the multimedia files in the album application by the user within the second reference time period.

[0292] In one example, taking the second reference time period as within one month before the current time, if the user frequently views the photos of their baby in the recent month, it can be inferred that the user wants to view the multimedia files of their baby, that is, the viewing intention of the user for the multimedia file is to view the multimedia files of their baby.

[0293] In the embodiments of the present application, by according to the multimedia files taken by the user within the first reference time period, or according to the operation information of the multimedia files in the album application by the user within the second reference time period, the viewing intention of the user for the multimedia file can be obtained, improving the flexibility of obtaining the viewing intention of the user for the multimedia file.

[0294] For the multimedia file processing method provided by the embodiments of the present application, the execution subject may be a multimedia file processing device. In the embodiments of the present application, taking the multimedia file processing device executing the multimedia file processing method as an example, the multimedia file processing device provided by the embodiments of the present application is described.

[0295] Figure 10 It is a schematic structural diagram of a multimedia file processing device shown according to an exemplary embodiment. As Figure 10 shown, the multimedia file processing device 1000 may include:

[0296] A receiving module 1010, configured to receive a first input to the memory creation interface;

[0297] A processing module 1020, configured to, in response to the first input, perform a multi-layer perception process on the spliced feature vectors corresponding to each multimedia file set in the multimedia database through a data push model, to obtain the text feature vectors respectively corresponding to each multimedia file set, and the relevance between each multimedia file in each multimedia file set and the memory theme created by the first input;

[0298] An output module 1030, configured to output a recalled video corresponding to the recalled theme created based on the text feature vector and the relevance, for the first input.

[0299] In an embodiment of the present application, when pushing a recalled video including a multimedia file, the push is performed according to the first input of the user to the recall creation interface, and the pushed recalled video corresponds to the recalled theme created by the first input. In this way, according to the user's needs, the recalled video required by the user can be pushed at the timing desired by the user, improving the flexibility of pushing the recalled video. In addition, when pushing the recalled video, a multi-layer perception process can be performed on the splicing feature vectors corresponding to each multimedia file set in the multimedia database through a data push model to obtain the text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the recalled theme created by the first input. Then, based on the text feature vector and the relevance, a recalled video corresponding to the recalled theme created by the first input is output, rather than pushing it to the user randomly, improving the accuracy of pushing the recalled video.

[0300] In some embodiments of the present application, the recall creation interface includes an input box; the receiving module 1010 is specifically configured to:

[0301] Receive a first input of entering a recalled theme in the recalled theme input box;

[0302] Wherein, the recalled theme created by the first input is the recalled theme entered by the first input in the input box.

[0303] In some embodiments of the present application, the recall creation interface includes at least one recalled theme identifier, and the recalled theme identifier is used to indicate a recalled theme; the receiving module 1010 is specifically configured to:

[0304] Receive a first input to one of the at least one recalled theme identifier;

[0305] Wherein, the recalled theme created by the first input is the recalled theme indicated by one of the recalled theme identifiers selected by the first input from the at least one recalled theme identifier.

[0306] In some embodiments of the present application, the output module 1030 is specifically configured to:

[0307] According to the text feature vectors corresponding to each multimedia file set, obtain the text push information of each multimedia file set respectively, and the text push information of each multimedia file set includes the recalled theme name corresponding to each multimedia file set;

[0308] Select at least one multimedia file set that matches the memory theme created by the first input from all the multimedia file sets in the multimedia database according to the memory theme name corresponding to each multimedia file set as the alternative multimedia file set;

[0309] Obtain alternative multimedia files according to the relevance between each multimedia file in the alternative multimedia file set and the memory theme created by the first input, where the alternative multimedia files are at least one multimedia file in the alternative multimedia file set with a relevance greater than the relevance threshold;

[0310] Output a memory video according to the alternative multimedia files.

[0311] In some embodiments of the present application, the output module 1030 is specifically configured to:

[0312] Based on at least one reference transition diagram, splice every two adjacent multimedia files in the alternative multimedia files and output a memory video.

[0313] In some embodiments of the present application, the device may further include:

[0314] A clustering module, configured to cluster all the multimedia files in the album application according to at least one clustering dimension to obtain at least one multimedia file set corresponding to different clustering dimensions respectively; where, for at least one multimedia file set corresponding to one clustering dimension, the clustering attributes of the same multimedia file set are the same; where, the clustering attributes include at least one of the following: shooting time, shooting location, shooting object, shooting content;

[0315] A determination module, configured to generate a multimedia database according to at least one multimedia file set corresponding to different clustering dimensions respectively.

[0316] In some embodiments of the present application, the clustering module is specifically configured to:

[0317] Cluster each multimedia file according to the shooting time of each multimedia file in the album application to obtain at least one time multimedia file set corresponding to different shooting times;

[0318] Cluster each multimedia file according to the shooting location of each multimedia file in the album application to obtain at least one location multimedia file set corresponding to different shooting locations;

[0319] Cluster the shooting objects in each multimedia file according to the shooting objects in each multimedia file in the album application to obtain at least one shooting object multimedia file set corresponding to different shooting objects;

[0320] According to the shooting content in each multimedia file in the photo album application, each multimedia file is classified according to the shooting content, and at least one multimedia file set corresponding to different shooting contents is obtained.

[0321] In some embodiments of the present application, the determining module is specifically configured to:

[0322] According to the shooting time of the at least one time multimedia file set, the at least one time multimedia file set is clustered according to different time ranges to obtain at least one continuous time multimedia file set;

[0323] All multimedia files in the at least one location multimedia file set are sorted according to the shooting date to obtain a time-sorted multimedia file set;

[0324] According to the shooting location of each multimedia file in the time-sorted multimedia file set, the shooting location with the highest frequency of occurrence is used as the permanent residence;

[0325] The multimedia files in the time-sorted multimedia file set whose shooting locations are between the permanent residences are clustered to obtain at least one continuous location multimedia file set;

[0326] According to the at least one time multimedia file set, the at least one location multimedia file set, the at least one shooting object multimedia file set, the at least one shooting content multimedia file set, the at least one continuous time multimedia file set, and the at least one continuous location multimedia file set, a multimedia database is generated.

[0327] In some embodiments of the present application, the determining module is specifically configured to:

[0328] The multimedia files in the multimedia file sets that appear in at least two different clustering dimensions are clustered to obtain at least one cross-information multimedia file set;

[0329] According to the at least one time multimedia file set, the at least one location multimedia file set, the at least one shooting object multimedia file set, the at least one shooting content multimedia file set, the at least one continuous time multimedia file set, the at least one continuous location multimedia file set, and the at least one cross-information multimedia file set, a multimedia database is generated.

[0330] In some embodiments of the present application, the processing module 1020 is specifically configured to:

[0331] Through a data push model, image coding is performed on each multimedia file set in the multimedia database to obtain the image feature vector of each multimedia file set;

[0332] Through the data push model, splice the image feature vectors and text encoding vectors corresponding to the multimedia file sets to obtain the spliced feature vectors of each multimedia file set;

[0333] Based on the data push model, perform multi-layer perception processing on the spliced feature vectors corresponding to each multimedia file set to obtain the text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input.

[0334] In some embodiments of the present application, the processing module 1020 is specifically configured to:

[0335] According to the reference segmentation size, perform image segmentation on the images corresponding to each multimedia file in each multimedia file set to obtain at least two image blocks for each image; wherein, when the multimedia file is a video file, the image corresponding to the multimedia file is a preset number of image frames selected from the video file;

[0336] Splice the encoding vectors of each image block corresponding to each image and the first position encoding vector to obtain the block encoding vector of each image block;

[0337] Fuse the block encoding vectors of each image block to obtain the image encoding vector of each image;

[0338] Extract the features of each image according to the image encoding vector of each image to obtain the feature vector of each image;

[0339] Splice the feature vectors of the images corresponding to each multimedia file in the multimedia file set with the second position encoding vector to obtain the image feature vector of each multimedia file set.

[0340] In some embodiments of the present application, the processing module 1020 is further configured to:

[0341] Before performing image segmentation on the images corresponding to each multimedia file in each multimedia file set according to the reference segmentation size to obtain at least two image blocks corresponding to each image, adjust the size of the images corresponding to each multimedia file in the multimedia file set according to the reference size; perform normalization processing on each image whose size is adjusted to the reference size to obtain each normalized image;

[0342] The processing module 1020 is specifically configured to:

[0343] Perform image segmentation processing on each normalized image in the multimedia file set according to the reference segmentation size to obtain at least two image blocks corresponding to each normalized image.

[0344] In some embodiments of the present application, the processing module 1020 is further configured to: before splicing the image feature vectors and text encoding vectors corresponding to the multimedia file sets through the data push model to obtain the spliced feature vectors of each multimedia file set, generate text description information for each multimedia file set according to the shooting information of each multimedia file in each multimedia file set; convert the text description information into text features to obtain the text encoding vectors of the multimedia file sets.

[0345] In some embodiments of the present application, the apparatus may further include:

[0346] An acquisition module, configured to acquire at least two training samples, each training sample including the sample text feature vectors of each sample multimedia file set in the sample multimedia database, and the sample relevance of each sample multimedia file in each sample multimedia file set, where the sample relevance of each sample multimedia file is the relevance between the sample multimedia file and the themes of each sample multimedia file set, and the theme of each sample multimedia file set is determined based on the sample file feature vectors of the sample multimedia file set;

[0347] An input module, configured to input at least two training samples into a reference data push model, and output the predicted text feature vectors of each sample multimedia file set and the predicted relevance of each sample multimedia file in each sample multimedia file set;

[0348] A model training module, configured to train the reference data push model according to the difference degree between the predicted text feature vectors and the sample text feature vectors of each sample multimedia file set, and the difference degree between the predicted relevance and the sample relevance of each sample multimedia file in each sample multimedia file set, to obtain the data push model.

[0349] In some embodiments of the present application, the model training module is specifically configured to:

[0350] Obtain a first loss function value of the reference data push model according to the difference degree between the predicted text feature vectors and the sample text feature vectors of each sample multimedia file set;

[0351] Obtain a second loss function value of the reference data push model according to the difference degree between the predicted relevance and the sample relevance of each sample multimedia file in each sample multimedia file set;

[0352] Train the reference data push model according to the first loss function value and the second loss function value to obtain the data push model.

[0353] Figure 11 It is a schematic structural diagram of another multimedia file processing device shown according to an exemplary embodiment. As Figure 11 shown, the multimedia file processing device 1100 may include:

[0354] An acquisition module 1110, configured to acquire a user's viewing intention for multimedia files in an album application;

[0355] A processing module 1120, configured to perform multi-layer perception processing on the spliced feature vectors corresponding to each multimedia file set in a multimedia database through a data push model, to obtain text feature vectors respectively corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the theme represented by the viewing intention;

[0356] An output module 1130, configured to output a recall video corresponding to the theme represented by the viewing intention based on the text feature vectors and the relevance.

[0357] In the embodiments of the present application, when pushing a recall video containing multimedia files, it can be pushed according to the user's viewing intention for the multimedia files. In this way, according to the user's needs, the recall video required by the user can be actively pushed according to the user's intention, improving the flexibility of pushing the recall video. In addition, when pushing the recall video, the spliced feature vectors corresponding to each multimedia file set in the multimedia database can be subjected to multi-layer perception processing through a data push model to obtain text feature vectors respectively corresponding to each multimedia file set, and the relevance between each multimedia file in each multimedia file set and the theme represented by the viewing intention. Then, based on the text feature vectors and the relevance, a recall video corresponding to the theme represented by the viewing intention is output, rather than pushing it to the user randomly, improving the accuracy of pushing the recall video.

[0358] The multimedia file processing device in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0359] The multimedia file processing device in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0360] The multimedia file processing device provided by the embodiments of the present application can implement Figure 1 or Figure 9 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0361] Optionally, as Figure 12 shown, the embodiments of the present application further provide an electronic device 1200, including a processor 1201 and a memory 1202. A program or instruction that can run on the processor 1201 is stored on the memory 1202. When the program or instruction is executed by the processor 1201, it implements each step of the above-mentioned multimedia file processing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0362] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0363] Figure 13 Schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0364] The electronic device 1300 includes, but is not limited to, components such as a radio frequency unit 1301, a network module 1302, an audio output unit 1303, an input unit 1304, a sensor 1305, a display unit 1306, a user input unit 1307, an interface unit 1308, a memory 1309, and a processor 1310.

[0365] Those skilled in the art can understand that the electronic device 1300 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1310 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 13 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0366] Among them, the user input unit 1307 is used to receive a first input to the recollection creation interface;

[0367] The processor 1310 is configured to, in response to the first input, perform multi-layer perception processing on the splicing feature vectors corresponding to each multimedia file set in the multimedia database through a data push model, to obtain text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the recollection theme created by the first input; and output a recollection video corresponding to the recollection theme created by the first input based on the text feature vectors and the relevance.

[0368] In this way, when pushing a recollection video containing multimedia files, it is pushed according to the user's first input to the recollection creation interface, and the pushed recollection video corresponds to the recollection theme created by the first input. Thus, according to the user's needs, the recollection video required by the user can be pushed at the timing the user wants, improving the flexibility of pushing the recollection video. In addition, when pushing the recollection video, the multi-layer perception processing can be performed on the splicing feature vectors corresponding to each multimedia file set in the multimedia database through the data push model to obtain text feature vectors corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the recollection theme created by the first input, and then output a recollection video corresponding to the recollection theme created by the first input based on the text feature vectors and the relevance, rather than pushing it to the user randomly, improving the accuracy of pushing the recollection video.

[0369] Optionally, the recollection creation interface includes a recollection theme input box; the user input unit 1307 is further configured to receive a first input of entering a recollection theme in the recollection theme input box; wherein, the recollection theme created by the first input is the recollection theme entered in the input box by the first input.

[0370] In this way, the user can, according to their own needs, enter a recollection theme in the recollection theme input box, and thus can customize the recollection theme of the recollection video with the recollection theme entered in the input box, thereby improving the flexibility of recollection video creation.

[0371] Optionally, the recollection creation interface includes at least one recollection theme identifier for indicating a recollection theme; the user input unit 1307 is further configured to receive a first input of one of the at least one recollection theme identifiers; wherein, the recollection theme created by the first input is the recollection theme indicated by one of the recollection theme identifiers selected from the at least one recollection theme identifiers by the first input.

[0372] In this way, in the case where there is an identifier corresponding to the recollection theme desired by the user among the at least one recollection theme identifiers included in the recollection creation interface, the user can directly click on the recollection theme identifier without entering a recollection theme in the recollection theme input box, simplifying the user operation and improving the efficiency of recollection video creation.

[0373] Optionally, the processor 1310 is further configured to respectively obtain text push information for each multimedia file set according to the text feature vector corresponding to each multimedia file set, where the text push information for each multimedia file set includes the recollection theme name corresponding to each multimedia file set; select at least one multimedia file set that matches the recollection theme created by the first input from all the multimedia file sets in the multimedia database according to the recollection theme name corresponding to each multimedia file set as an alternative multimedia file set; obtain alternative multimedia files according to the relevance between each multimedia file in the alternative multimedia file set and the recollection theme created by the first input, where the alternative multimedia files are at least one multimedia file in the alternative multimedia file set with a relevance greater than a relevance threshold; and output a recollection video according to the alternative multimedia files.

[0374] In this way, by selecting at least one set of multimedia files that match the memory theme created from the first input from all the multimedia collections in the multimedia database as the alternative multimedia file set, and then selecting alternative multimedia files from the alternative multimedia file set instead of selecting alternative multimedia files from all the multimedia collections in the multimedia database, the time for selecting alternative multimedia files is saved. Moreover, the multimedia files in the memory video are all multimedia files with a high degree of relevance to the memory theme created from the first input. Thus, the pushed memory video is more in line with the memory theme created by the user's first input, improving the accuracy of pushing the memory video.

[0375] Optionally, the processor 1310 is further configured to splice every two adjacent multimedia files in the alternative multimedia files based on at least one reference transition graph and output a memory video.

[0376] In this way, by splicing every two adjacent multimedia files in the alternative multimedia files based on at least one reference transition graph, a memory video with transition effects can be obtained, which can increase the viewing interest of the alternative multimedia files.

[0377] Optionally, the processor 1310 is further configured to cluster all the multimedia files in the album application according to at least one clustering dimension to obtain at least one set of multimedia files corresponding to different clustering dimensions respectively; where the clustering attributes of the same set of multimedia files corresponding to one clustering dimension are the same; where the clustering attributes include at least one of the following: shooting time, shooting location, shooting object, shooting content; and generate a multimedia database according to the at least one set of multimedia files corresponding to different clustering dimensions respectively.

[0378] In this way, by clustering all the multimedia files in the album application according to at least one clustering dimension, at least one set of multimedia files with different clustering dimensions can be obtained. Thus, when the user wants to create a memory video with different types of memory themes, a memory video corresponding to the memory theme of the type required by the user can be obtained from the at least one set of multimedia files with different clustering dimensions, improving the recommendation diversity of the memory video.

[0379] Optionally, the processor 1310 is further configured to cluster each multimedia file according to the shooting time of each multimedia file in the album application, so as to obtain at least one time multimedia file set corresponding to different shooting times; cluster each multimedia file according to the shooting location of each multimedia file in the album application, so as to obtain at least one location multimedia file set corresponding to different shooting locations; cluster the shooting objects in each multimedia file according to the shooting objects in each multimedia file in the album application, so as to obtain at least one shooting object multimedia file set corresponding to different shooting objects; classify each multimedia file according to the shooting content in each multimedia file in the album application, so as to obtain at least one shooting content multimedia file set corresponding to different shooting contents.

[0380] In this way, by clustering each multimedia file according to the clustering dimensions of shooting time, shooting location, shooting object, and shooting content, at least one multimedia file set corresponding to different clustering dimensions is obtained. Subsequently, when recommending a memory video with a memory theme based on shooting time, shooting location, shooting object, and shooting content, the recommended multimedia files can be more suitable for the memory theme, improving the accuracy of multimedia file push.

[0381] Optionally, the processor 1310 is further configured to cluster the at least one time multimedia file set according to the shooting time of the at least one time multimedia file set, so as to obtain at least one continuous time multimedia file set; sort all the multimedia files in the at least one location multimedia file set according to the shooting date, so as to obtain a time-sorted multimedia file set; use the shooting location with the highest frequency of occurrence as the permanent residence according to the shooting location of each multimedia file in the time-sorted multimedia file set; cluster the multimedia files with shooting locations between the permanent residences in the time-sorted multimedia file set, so as to obtain at least one continuous location multimedia file set; generate a multimedia database according to the at least one time multimedia file set, the at least one location multimedia file set, the at least one shooting object multimedia file set, the at least one shooting content multimedia file set, the at least one continuous time multimedia file set, and the at least one continuous location multimedia file set.

[0382] In this way, by clustering the multimedia files in the time multimedia file set according to continuous time, a continuous time multimedia file set with continuous time information is formed, and by clustering the multimedia files in the location multimedia file set according to continuous location, a continuous location multimedia file set with continuous location information is formed. Subsequently, when making recall video recommendations based on continuous time or continuous location for the recall theme, the recommended multimedia files can be more suitable for the recall theme, further improving the accuracy of multimedia file push.

[0383] Optionally, the processor 1310 is further configured to cluster the multimedia files in the multimedia file sets that appear in at least two different clustering dimensions to obtain at least one cross-information multimedia file set; generate a multimedia database according to the at least one time multimedia file set, the at least one location multimedia file set, the at least one shooting object multimedia file set, the at least one shooting content multimedia file set, the at least one continuous time multimedia file set, the at least one continuous location multimedia file set, and the at least one cross-information multimedia file set.

[0384] In this way, by clustering the multimedia files in the multimedia file sets that appear in at least two different clustering dimensions to obtain at least one cross-information multimedia file set, subsequently, when making recall video recommendations based on the cross-information for the recall theme, the recommended multimedia files can be more suitable for the recall theme, further improving the accuracy of multimedia file push.

[0385] Optionally, the processor 1310 is further configured to perform image encoding on each multimedia file set in the multimedia database through a data push model to obtain the image feature vector of each multimedia file set; splice the image feature vector corresponding to the multimedia file set and the text encoding vector through the data push model to obtain the spliced feature vector of each multimedia file set; perform multi-layer perception processing on the spliced feature vector corresponding to each multimedia file set based on the data push model to obtain the text feature vector corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input.

[0386] In this way, each multimedia file set in the multimedia database is subjected to image encoding through the data push model to obtain the image feature vector of each multimedia file set. Then, through the data push model, the image feature vector corresponding to the multimedia file set and the text encoding vector are concatenated to obtain the concatenated feature vector of each multimedia file set. Furthermore, based on the data push model, multi-layer perception processing is performed on the concatenated feature vector corresponding to each multimedia file set, so as to obtain the text feature vector corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input. In this way, the recall theme of the pushed recall video can be made clear, and the relevance of the multimedia files in the pushed recall video is high.

[0387] Optionally, the processor 1310 is further configured to perform image segmentation on the image corresponding to each multimedia file in each multimedia file set according to the reference segmentation size to obtain at least two image blocks of each image; wherein, when the multimedia file is a video file, the image corresponding to the multimedia file is a preset number of image frames selected from the video file; concatenate the encoding vector of each image block corresponding to each image and the first position encoding vector to obtain the block encoding vector of each image block; fuse the block encoding vectors of each image block to obtain the image encoding vector of each image; extract the features of each image according to the image encoding vector of each image to obtain the feature vector of each image; concatenate the feature vectors of the images corresponding to each multimedia file in the multimedia file set with the second position encoding vector to obtain the image feature vector of each multimedia file set.

[0388] In this way, by quantifying each multimedia file set, the image feature vector of each multimedia file set is obtained. When processing the multimedia file set subsequently, the image feature vector of the multimedia file set can be processed, rather than processing each multimedia file in the multimedia file set, reducing the processing amount of the multimedia file set, improving the processing efficiency of the multimedia file set, and further improving the pushing efficiency of the recall video.

[0389] Optionally, the processor 1310 is further configured to adjust the size of the image corresponding to each multimedia file in the multimedia file set according to the reference size; perform normalization processing on each image whose size is adjusted to the reference size to obtain each normalized image; perform image segmentation processing on each normalized image in the multimedia file set according to the reference segmentation size to obtain at least two image blocks corresponding to each normalized image.

[0390] In this way, by adjusting the size of the image corresponding to each multimedia file to a reference size that matches the data push model, it is convenient for the data push model to process the image corresponding to each multimedia file. In addition, each image with its size adjusted to the reference size is normalized, which simplifies the processing amount of each image with its size adjusted to the reference size, reduces the calculation amount for each multimedia file, and improves the processing efficiency of each multimedia file.

[0391] Optionally, the processor 1310 is further configured to generate text description information for each multimedia file set according to the shooting information of each multimedia file in each multimedia file set; convert the text description information into text features to obtain a text encoding vector for each multimedia file set.

[0392] In this way, for each multimedia file set, the text encoding vector of the multimedia file set is determined according to the shooting information of each multimedia file in the multimedia file set. The text encoding vector of the multimedia file set determined in this way is more in line with the shooting information of each multimedia file in the multimedia file set. The recalled video pushed based on the text encoding vector of the multimedia file set will also be more in line with the shooting information of the multimedia file. That is to say, the recalled video pushed is in line with the shooting information of the multimedia file. That is, through the solution of the present application, a recalled video with a high user perception can be recommended.

[0393] Optionally, the processor 1310 is further configured to obtain at least two training samples. Each training sample includes a sample text feature vector of each sample multimedia file set in the sample multimedia database, and the sample relevance of each sample multimedia file in each sample multimedia file set. The sample relevance of each sample multimedia file is the relevance between the sample multimedia file and the theme of each sample multimedia file set, and the theme of each sample multimedia file set is determined based on the sample file feature vector of the sample multimedia file set; input the at least two training samples into the reference data push model, and output a predicted text feature vector of each sample multimedia file set and a predicted relevance of each sample multimedia file in each sample multimedia file set; train the reference data push model according to the difference degree between the predicted text feature vector and the sample text feature vector of each sample multimedia file set, and the difference degree between the predicted relevance and the sample relevance of each sample multimedia file in each sample multimedia file set, to obtain the data push model.

[0394] In this way, the reference data push model is trained through at least two training samples to obtain the data push model. In this way, the recalled video corresponding to the recall theme created with the first input can be directly pushed to the user based on this data push model, improving the pushing efficiency of the recalled video corresponding to the recall theme created with the first input.

[0395] Optionally, the processor 1310 is further configured to obtain a first loss function value of the reference data push model according to the difference degree between the predicted text feature vector of each sample multimedia file set and the sample text feature vector; obtain a second loss function value of the reference data push model according to the difference degree between the predicted relevance of each sample multimedia file in each sample multimedia file set and the sample relevance; and train the reference data push model according to the first loss function value and the second loss function value to obtain the data push model.

[0396] In this way, by using the first loss function value representing the distance between the predicted text feature vector and the sample text feature vector, and the second loss function value representing the distance between the sample multimedia file and the theme corresponding to each sample multimedia file set to train the reference data push model, rather than using a single loss function value to train the reference data push model, the robustness of the data push model is improved.

[0397] Optionally, the processor 1310 is configured to obtain the viewing intention of the user for the multimedia files in the album application; perform multi-layer perception processing on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database through the data push model to obtain the text feature vector corresponding to each multimedia file set respectively and the relevance between each multimedia file in each multimedia file set and the theme represented by the viewing intention; and output a recall video corresponding to the theme represented by the viewing intention based on the text feature vector and the relevance.

[0398] In this way, when pushing the recall video containing multimedia files, it can be pushed according to the user's viewing intention for the multimedia files. In this way, according to the user's needs, the recall video required by the user can be actively pushed according to the user's intention, improving the flexibility of pushing the recall video. In addition, when pushing the recall video, the multi-layer perception processing can be performed on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database through the data push model to obtain the text feature vector corresponding to each multimedia file set respectively, and the relevance between each multimedia file in each multimedia file set and the theme represented by the viewing intention, and then based on the text feature vector and the relevance, output the recall video corresponding to the theme represented by the viewing intention, rather than pushing it to the user randomly, improving the accuracy of pushing the recall video.

[0399] It should be understood that in the embodiments of the present application, the input unit 1304 may include a Graphics Processing Unit (GPU) 13041 and a microphone 13042. The graphics processor 13041 processes the image data of static pictures or videos obtained by an image capture device (such as a color camera) in a video capture mode or an image capture mode. The display unit 1306 may include a display panel 13061, and the display panel 13061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1307 includes at least one of a touch panel 13071 and other input devices 13072. The touch panel 13071 is also referred to as a touch screen. The touch panel 13071 may include two parts: a touch detection device and a touch controller. The other input devices 13072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated herein.

[0400] The memory 1309 can be used to store software programs and various data. The memory 1309 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1309 may include a volatile memory or a non-volatile memory, or the memory 1309 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1309 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0401] The processor 1310 may include one or more processing units; optionally, the processor 1310 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1310.

[0402] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiments of the multimedia file processing method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0403] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs.

[0404] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is configured to run programs or instructions to implement each process of the above embodiment of the multimedia file processing method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0405] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip.

[0406] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above embodiment of the multimedia file processing method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0407] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0408] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0409] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A multimedia file processing method, characterized in that: The method comprises: Receiving a first input to a memory creation interface; In response to the first input, a multi-layer perception processing is performed on the concatenated feature vector corresponding to each multimedia file set in the multimedia database through a data push model to obtain a text feature vector corresponding to each multimedia file set and a relevance between each multimedia file in each multimedia file set and a recall theme created by the first input; Based on the text feature vector and the relevance, a recollection video corresponding to the recollection theme created by the first input is output.

2. The method according to claim 1, characterized in that The memory creation interface includes a memory theme input box; The receiving a first input to the memory creation interface includes: receiving a first input of inputting a recollection theme in a recollection theme input box; Among them, the memory theme created by the first input is the memory theme entered by the first input in the input box.

3. The method according to claim 1, characterized in that The memory creation interface includes at least one memory theme identifier, and the memory theme identifier is used to indicate a memory theme; The receiving a first input to the memory creation interface includes: receiving a first input of a recollection theme identifier among the at least one recollection theme identifier; The recollection theme created by the first input is a recollection theme indicated by a recollection theme identifier selected by the first input from among the at least one recollection theme identifiers.

4. The method according to claim 1, characterized in that: The step of outputting a memory video corresponding to a memory theme created by the first input based on the text feature vector and the relevance comprises: According to the text feature vector corresponding to each multimedia file set, the text push information of each multimedia file set is obtained respectively, and the text push information of each multimedia file set includes the name of the recollection theme corresponding to each multimedia file set; According to the name of the recollection theme corresponding to each multimedia file set, selecting at least one multimedia file set matching the recollection theme created by the first input from all multimedia atlases in the multimedia database as a candidate multimedia file set; Obtaining candidate multimedia files according to the relevance between each multimedia file in the candidate multimedia file set and the recollection theme created by the first input, wherein the candidate multimedia file is at least one multimedia file in the candidate multimedia file set whose relevance is greater than a relevance threshold; Output a recollection video according to the alternative multimedia file.

5. The method according to claim 4, characterized in that The step of outputting a recall video according to the candidate multimedia file comprises: Based on at least one reference transition graph, each two adjacent multimedia files in the candidate multimedia files are spliced ​​to output a recall video.

6. The method according to claim 1, characterized in that Before performing multi-layer perception processing on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database through the data push model to obtain the text feature vectors corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input, the method further includes: Clustering all multimedia files in the album application according to at least one clustering dimension to obtain at least one multimedia file set corresponding to different clustering dimensions; wherein the clustering attributes of the same multimedia file set in at least one multimedia file set corresponding to one clustering dimension are the same; wherein the clustering attributes include at least one of the following: shooting time, shooting location, shooting object, shooting content; A multimedia database is generated according to at least one multimedia file set corresponding to different clustering dimensions.

7. The method according to claim 6, characterized in that The step of clustering all multimedia files in the album application according to at least one clustering dimension to generate at least one multimedia file set corresponding to different clustering dimensions includes: According to the shooting time of each multimedia file in the album application, clustering each multimedia file according to the shooting time to obtain at least one time multimedia file set corresponding to different shooting times; According to the shooting location of each multimedia file in the album application, clustering each multimedia file according to the shooting location to obtain at least one location multimedia file set corresponding to different shooting locations; According to the photographed objects in each multimedia file in the album application, clustering the photographed objects in each multimedia file to obtain at least one photographed object multimedia file set corresponding to different photographed objects; According to the shooting content in each multimedia file in the album application, each multimedia file is classified according to the shooting content to obtain at least one shooting content multimedia file set corresponding to different shooting contents.

8. The method according to claim 7, characterized in that The step of generating a multimedia database according to at least one multimedia file set corresponding to different clustering dimensions includes: According to the shooting time of the at least one temporal multimedia file set, clustering the at least one temporal multimedia file set according to different time ranges to obtain at least one continuous temporal multimedia file set; Sorting all multimedia files in the at least one location multimedia file set according to shooting dates to obtain a time-sorted multimedia file set; According to the shooting location of each multimedia file in the time-sorted multimedia file set, the shooting location with the highest frequency is used as the permanent location; Clustering the multimedia files in the time-sorted multimedia file set whose shooting locations are between the permanent locations to obtain at least one continuous location multimedia file set; A multimedia database is generated according to the at least one time multimedia file set, the at least one location multimedia file set, the at least one shooting object multimedia file set, the at least one shooting content multimedia file set, the at least one continuous time multimedia file set and the at least one continuous location multimedia file set.

9. The method according to claim 8, characterized in that The step of generating the multimedia database according to the at least one time multimedia file set, the at least one location multimedia file set, the at least one shooting object multimedia file set, the at least one shooting content multimedia file set, the at least one continuous time multimedia file set, and the at least one continuous location multimedia file set comprises: Clustering multimedia files that appear in multimedia file sets in at least two different clustering dimensions to obtain at least one cross-information multimedia file set; A multimedia database is generated based on the at least one time multimedia file set, the at least one location multimedia file set, the at least one shooting object multimedia file set, the at least one shooting content multimedia file set, the at least one continuous time multimedia file set, the at least one continuous location multimedia file set and the at least one cross-information multimedia file set.

10. The method according to claim 1, characterized in that The method of performing multi-layer perception processing on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database through the data push model to obtain the text feature vectors corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the recollection theme created by the first input includes: Through the data push model, image encoding is performed on each multimedia file set in the multimedia database to obtain the image feature vector of each multimedia file set; By using the data push model, the image feature vector and the text encoding vector corresponding to the multimedia file set are concatenated to obtain a concatenated feature vector for each multimedia file set; Based on the data push model, multi-layer perception processing is performed on the concatenated feature vector corresponding to each multimedia file set to obtain the text feature vector corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input.

11. The method according to claim 10, characterized in that The data push model is used to perform image encoding on each multimedia file set in the multimedia database to obtain an image feature vector of each multimedia file set, including: According to the reference segmentation size, image segmentation is performed on the image corresponding to each multimedia file in each multimedia file set to obtain at least two image blocks of each image; wherein, when the multimedia file is a video file, the image corresponding to the multimedia file is a preset number of image frames selected from the video file; Concatenate the coding vector of each image block corresponding to each image and the first position coding vector to obtain a block coding vector of each image block; The block coding vectors of each image block are fused to obtain the image coding vector of each image; According to the image coding vector of each image, the features of each image are extracted to obtain the feature vector of each image; The feature vector of the image corresponding to each multimedia file in the multimedia file set is concatenated with the second position coding vector to obtain the image feature vector of each multimedia file set.

12. The method according to claim 11, characterized in that Before performing image segmentation on the image corresponding to each multimedia file in each multimedia file set according to the reference segmentation size to obtain at least two image blocks corresponding to each image, the method further includes: adjusting the size of the image corresponding to each multimedia file in the multimedia file set according to the reference size; Performing normalization processing on each image whose size is adjusted to the reference size to obtain each normalized image; The step of performing image segmentation on the image corresponding to each multimedia file in each multimedia file set according to the reference segmentation size to obtain at least two image blocks corresponding to each image includes: According to the reference segmentation size, image segmentation processing is performed on each normalized image in the multimedia file set to obtain at least two image blocks corresponding to each normalized image.

13. The method according to claim 10, characterized in that Before the image feature vector and the text encoding vector corresponding to the multimedia file set are spliced ​​by the data push model to obtain the spliced ​​feature vector of the multimedia file set, the method further includes: Generate text description information of each multimedia file set according to the shooting information of each multimedia file in each multimedia file set; The text description information is converted into text features to obtain a text encoding vector for each multimedia file set.

14. The method according to claim 1, characterized in that Before performing multi-layer perception processing on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database through the data push model to obtain the text feature vectors corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the recollection theme created by the first input, the method further includes: Acquire at least two training samples, each training sample comprising a sample text feature vector of each sample multimedia file set in a sample multimedia database, and a sample relevance of each sample multimedia file in each sample multimedia file set, wherein the sample relevance of each sample multimedia file is a relevance between the sample multimedia file and a theme of each sample multimedia file set, and the theme of each sample multimedia file set is determined based on the sample file feature vector of the sample multimedia file set; Inputting at least two training samples into a reference data push model, outputting a predicted text feature vector for each sample multimedia file set, and a predicted relevance of each sample multimedia file in each sample multimedia file set; The reference data push model is trained according to the difference between the predicted text feature vector and the sample text feature vector of each sample multimedia file set, and the difference between the predicted relevance and the sample relevance of each sample multimedia file in each sample multimedia file set to obtain the data push model.

15. The method according to claim 14, characterized in that The method of training the reference data push model according to the difference between the predicted text feature vector and the sample text feature vector of each sample multimedia file set, and the difference between the predicted relevance and the sample relevance of each sample multimedia file in each sample multimedia file set, to obtain the data push model, comprises: Obtaining a first loss function value of the reference data push model according to the degree of difference between the predicted text feature vector and the sample text feature vector of each sample multimedia file set; Obtaining a second loss function value of the reference data push model according to a difference between a predicted relevance and a sample relevance of each sample multimedia file in each sample multimedia file set; The reference data push model is trained according to the first loss function value and the second loss function value to obtain the data push model.

16. A multimedia file processing method, characterized in that: The method comprises: Get the user's viewing intention for multimedia files in the photo album application; Through the data push model, multi-layer perception processing is performed on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database to obtain the text feature vectors corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the subject represented by the viewing intention; Based on the text feature vector and the relevance, a recall video corresponding to the theme represented by the viewing intention is output.

17. A multimedia file processing device, characterized in that: The device comprises: A receiving module, configured to receive a first input to the memory creation interface; a processing module, configured to, in response to the first input, perform multi-layer perception processing on the concatenated feature vector corresponding to each multimedia file set in the multimedia database through a data push model to obtain a text feature vector corresponding to each multimedia file set, and a relevance between each multimedia file in each multimedia file set and a recall theme created by the first input; An output module is used to output a memory video corresponding to the memory theme created by the first input based on the text feature vector and the relevance.

18. The device according to claim 17, characterized in that The memory creation interface includes an input box; the receiving module is specifically used to: receiving a first input of inputting a recollection theme in a recollection theme input box; Among them, the memory theme created by the first input is the memory theme entered by the first input in the input box.

19. The device according to claim 17, characterized in that The memory creation interface includes at least one memory theme identifier, and the memory theme identifier is used to indicate the memory theme; the receiving module is specifically used to: receiving a first input of a recollection theme identifier among the at least one recollection theme identifier; The recollection theme created by the first input is a recollection theme indicated by a recollection theme identifier selected by the first input from among the at least one recollection theme identifiers.

20. The device according to claim 17, characterized in that The output module is specifically used for: According to the text feature vector corresponding to each multimedia file set, the text push information of each multimedia file set is obtained respectively, and the text push information of each multimedia file set includes the name of the recollection theme corresponding to each multimedia file set; According to the name of the recollection theme corresponding to each multimedia file set, selecting at least one multimedia file set matching the recollection theme created by the first input from all multimedia atlases in the multimedia database as a candidate multimedia file set; Obtaining candidate multimedia files according to the relevance between each multimedia file in the candidate multimedia file set and the recollection theme created by the first input, wherein the candidate multimedia file is at least one multimedia file in the candidate multimedia file set whose relevance is greater than a relevance threshold; Output a recollection video according to the candidate multimedia file.

21. The device according to claim 17, characterized in that The device also includes: A clustering module, configured to cluster all multimedia files in the album application according to at least one clustering dimension, and obtain at least one multimedia file set corresponding to different clustering dimensions; wherein the clustering attributes of the same multimedia file set in at least one multimedia file set corresponding to one clustering dimension are the same; wherein the clustering attributes include at least one of the following: shooting time, shooting location, shooting object, and shooting content; The determination module is used to generate a multimedia database according to at least one multimedia file set corresponding to different clustering dimensions.

22. The device according to claim 21, characterized in that The clustering module is specifically used for: According to the shooting time of each multimedia file in the album application, clustering each multimedia file according to the shooting time to obtain at least one time multimedia file set corresponding to different shooting times; According to the shooting location of each multimedia file in the album application, clustering each multimedia file according to the shooting location to obtain at least one location multimedia file set corresponding to different shooting locations; According to the photographed objects in each multimedia file in the album application, clustering the photographed objects in each multimedia file to obtain at least one photographed object multimedia file set corresponding to different photographed objects; According to the shooting content in each multimedia file in the album application, each multimedia file is classified according to the shooting content to obtain at least one shooting content multimedia file set corresponding to different shooting contents.

23. The device according to claim 17, characterized in that The processing module is specifically used for: Through the data push model, image encoding is performed on each multimedia file set in the multimedia database to obtain the image feature vector of each multimedia file set; By using the data push model, the image feature vector and the text encoding vector corresponding to the multimedia file set are concatenated to obtain a concatenated feature vector for each multimedia file set; Based on the data push model, multi-layer perception processing is performed on the concatenated feature vector corresponding to each multimedia file set to obtain the text feature vector corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the recall theme created by the first input.

24. A multimedia file processing device, characterized in that: The device comprises: The acquisition module is used to obtain the user's viewing intention for the multimedia files in the album application; A processing module, used to perform multi-layer perception processing on the concatenated feature vectors corresponding to each multimedia file set in the multimedia database through a data push model, to obtain text feature vectors corresponding to each multimedia file set and the relevance between each multimedia file in each multimedia file set and the subject represented by the viewing intention; An output module is used to output a recall video corresponding to the theme represented by the viewing intention based on the text feature vector and the relevance.

25. An electronic device, characterized in that: It includes a processor and a memory, the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, it implements the steps of the multimedia file processing method according to any one of claims 1 to 15, or the steps of the multimedia file processing method according to claim 16.