An image processing method, user interface, and related devices
By receiving the guidance parameters selected by the user, identifying the image scene type, and filtering out images that meet the user's needs, the system solves the problem of cumbersome operations when users are browsing video materials to extract specific images, and achieves fast and personalized image filtering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2023-12-05
- Publication Date
- 2026-05-19
AI Technical Summary
The process of manually extracting specific images when browsing video footage is cumbersome and fails to meet personalized and diverse needs.
By receiving the user-selected guidance parameters, including guidance text, guidance audio, and guidance images, the system identifies the image scene type and filters out images that meet the user's needs based on similarity and aesthetic scores.
It enables the rapid selection of images that meet the guidance parameters from videos based on users' personalized and diverse needs, improving user experience and computational efficiency.
Smart Images

Figure CN120144821B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an image processing method, user interface and related apparatus. Background Technology
[0002] In addition to storing numerous images, users often also store video clips in their electronic device's gallery. Some users may want to select specific images from multiple video clips. Currently, users need to manually extract the desired images while browsing video clips, a cumbersome process. Summary of the Invention
[0003] This application provides an image processing method, user interface, and related apparatus, which enables the selection of images that meet the user's personalized and diverse needs from video frames of a first video selected by the user based on the user's selected guidance parameters.
[0004] Firstly, this application provides an image processing method applied to an electronic device. The method includes: receiving a first video selected by a user and guidance parameters, the guidance parameters including one or more of guidance text, guidance audio, and guidance images; splitting the video frames of the first video into multiple image subsets, the number of images in each subset being less than or equal to a preset number of images; identifying the scene type to which the multiple image subsets belong; determining a target image subset with the same scene type as the guidance parameters; and, based on the guidance parameters, determining a result image from the target image subset whose similarity to the guidance parameters is greater than a preset similarity. In this way, the electronic device can process the first video selected by the user according to the guidance parameters selected by the user, filtering images from the first video that meet the user's personalized and diverse needs. Since the guidance parameters are added by the user, the electronic device can filter images that better meet the user's needs, making it applicable to a wider range of scenarios.
[0005] In one possible implementation, based on guiding parameters, the result image most similar to the guiding parameters is determined from a subset of target images. Specifically, this includes: determining a candidate image set from the subset of target images based on the guiding parameters; the candidate image set includes multiple images from the subset of target images whose similarity to the guiding parameters is greater than a preset similarity; determining the aesthetic score of each image in the candidate image set, where the aesthetic score evaluates the image's composition, color, lighting, exposure, and theme; and determining the result image from the candidate image set whose aesthetic score is greater than a preset aesthetic score. In this way, the electronic device can filter the images most similar to the guiding parameters based on their aesthetic appeal, selecting the image that best matches the user's aesthetic preferences.
[0006] In one possible implementation, a candidate image set is determined from a subset of target images based on guiding parameters. Specifically, this includes: encoding a segment vector corresponding to the target image subset, where the segment vector includes features from all images in the target image subset; encoding the guiding parameters to obtain a guiding vector; determining the similarity between the target image subset and the guiding parameters based on the segment vector and the guiding vector; and identifying a candidate image set whose similarity is greater than a preset similarity from the candidate image set. In this way, the electronic device can quickly process a large number of images, even when the number of images is large, and identify images similar to the guiding parameters from a large pool of images.
[0007] In one possible implementation, based on a subset of target images, a segment vector corresponding to the subset is encoded. Specifically, this includes: encoding the images of the subset into image vectors; and calculating the average of the image vectors of all images in the subset to obtain the segment vector. In this way, the electronic device can calculate the similarity between multiple images and the guiding parameters at once, saving computational resources and energy consumption, and improving the filtering speed of the electronic device.
[0008] In one possible implementation, the guidance parameters include multiple items from the guidance image, guidance text, and guidance audio. The guidance vector includes multiple items from the guidance image vector, guidance text vector, and guidance audio vector. The guidance image corresponds to a first weight, the guidance text to a second weight, and the guidance audio to a third weight. Based on the fragment vector and the guidance vector, the similarity between the target image subset and the guidance parameters is determined. Specifically, this involves determining the similarity between the target image subset and the guidance parameters based on the similarity between the guidance image vector and the fragment vector, the first weight, the similarity between the guidance text vector and the fragment vector, the second weight, the similarity between the guidance audio vector and the fragment vector, and the third weight. In this way, the electronic device can assign different weights to different guidance parameters, and even when the importance of different guidance parameters varies, it can still obtain result images that better meet the user's needs.
[0009] In one possible implementation, determining the subset of images with the same scene type as the guidance parameter specifically includes: if the guidance parameter includes a guidance image, identifying the scene to which the guidance image belongs based on the text description of the scene type, and determining a target subset of images with the same scene type as the guidance image from multiple image subsets; if the guidance parameter includes guidance audio, identifying the scene to which the guidance audio belongs based on key information of the guidance audio, and determining a target subset of images with the same scene type as the guidance audio from multiple image subsets; if the guidance parameter includes guidance text, identifying the scene to which the guidance text belongs based on key information of the guidance text, and determining a target subset of images with the same scene type as the guidance text from multiple image subsets. In this way, the electronic device can identify the scene type of the guidance parameter and select the corresponding target subset of images.
[0010] In one possible implementation, when multiple images have aesthetic scores higher than a preset aesthetic score, the resulting image achieves the highest aesthetic score. This allows the electronic device to acquire a final image that better aligns with the user's aesthetic preferences.
[0011] In one possible implementation, the aesthetic score of each image in the candidate image set is determined, specifically by determining the aesthetic score of all images in the candidate image set based on an aesthetic evaluation model.
[0012] In one possible implementation, when multiple images have a similarity greater than a preset similarity, the resulting image has the highest similarity. This allows the electronic device to acquire the resulting image that best matches the guidance parameters.
[0013] In one possible implementation, before receiving the first video selected by the user and the guiding parameters, the method further includes: displaying a first interface, the first interface including a gallery application icon; receiving first input for the gallery application icon; in response to the first input, displaying one or more video options, the one or more video options including a video option for the first video; receiving the first video selected by the user, specifically including: receiving second input for the video option of the first video; in response to the second input, displaying a second interface and selecting the first video, the second interface being used to play the first video. In this way, the electronic device can receive the user's second input and select the first video.
[0014] In one possible implementation, after displaying the second interface and selecting the first video, the method further includes: receiving a third input to a first control on the second interface; and in response to the third input, displaying a second control, a third control, and a fourth control, wherein the second control is used to add guide text, the third control is used to add guide audio, and the fourth control is used to add guide image. In this way, the electronic device can provide the user with controls for adding guide parameters, making it easier for the user to determine the guide parameters.
[0015] In a second aspect, this application provides an electronic device including: a display screen, one or more processors and one or more memories coupled to a plurality of processors, the one or more memories being used to store an executable program, and when the one or more processors are executing the executable program, causing the electronic device to perform any of the possible implementation methods in the first aspect.
[0016] Thirdly, this application provides a readable storage medium that stores a program, which, when run on the processor of an electronic device, causes the electronic device to perform any of the possible implementations described in the first aspect. Attached Figure Description
[0017] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;
[0018] Figure 2 A schematic flowchart of an image processing method provided in an embodiment of this application;
[0019] Figure 3 A flowchart illustrating another image processing method provided in an embodiment of this application;
[0020] Figure 4 A schematic diagram of a text-image matching model provided in an embodiment of this application;
[0021] Figure 5 A schematic diagram of a training aesthetic evaluation model provided in an embodiment of this application;
[0022] Figure 6 A schematic diagram of an aesthetic evaluation model provided in an embodiment of this application;
[0023] Figures 7A-7K A set of interface schematic diagrams provided for embodiments of this application;
[0024] Figure 8 This is a schematic diagram of a similarity curve provided for an embodiment of this application. Detailed Implementation
[0025] The technical solutions in the embodiments of this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; the word "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.
[0026] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0027] In one possible implementation, the electronic device can provide a snapshot function for specified actions. Once this function is enabled, the electronic device can capture and save images of the subject performing a specified action within a designated scene during video recording. The designated scene can be a preset scene, such as a sports scene like skiing or swimming. The specified action can be a preset action, such as jumping or spinning. In this way, the electronic device can acquire images of the specified action within the designated scene during video recording without requiring user intervention. However, this method is limited to the video recording process and requires the electronic device to capture images within preset scenes and actions. The resulting image scenes and types are very limited, lacking versatility and failing to meet the personalized and diverse needs of a large number of users.
[0028] This application provides an image processing method. An electronic device, after determining a user-selected image set and guidance parameters, can determine a candidate image set from the image set based on the guidance parameters. The image set may include video frames from a specified video, or multiple images. Guidance parameters may include, but are not limited to, one or more of guidance text, guidance images, and guidance audio. The electronic device can then filter the image set to obtain the result image most similar to the guidance parameters. In this way, the electronic device can process the user-selected image set according to the user-selected guidance parameters, filtering out images that meet the user's personalized and diverse needs, thereby obtaining images with a greater variety of scenes and types.
[0029] When the guidance parameters include guidance text, the electronic device can select one or more images from the image set that match the guidance text description. When the guidance parameters include guidance images, the electronic device can select one or more images from the image set that are most similar to the guidance image. When the guidance parameters include guidance audio, the electronic device can select one or more images from the image set that match the guidance audio description. When the guidance parameters include multiple elements such as guidance text, guidance images, and guidance audio, the electronic device can calculate the similarity between the images in the image set and the guidance parameters based on each guidance parameter. The electronic device can also assign different weights to each guidance parameter. Based on the similarity and weight of each guidance parameter for the images in the image set, the electronic device can obtain a comprehensive score for the images. The electronic device can select one or more images with the highest comprehensive score, that is, one or more images that are most similar to the guidance image. Alternatively, the electronic device can select one or more images with a comprehensive score greater than a preset similarity score as one or more images that are most similar to the guidance image.
[0030] In some examples, when an electronic device selects multiple images from an image set that are most similar to the guiding parameters based on these parameters, it can use these images as a candidate image set. Then, the electronic device can use an aesthetic evaluation algorithm to select the result image from the candidate image set. The result image is the one or more images with the highest aesthetic scores in the candidate image set. The aesthetic score represents the visual appeal of the image, and the electronic device can determine the aesthetic score based on multiple metrics (e.g., composition, color, exposure, etc.). In this way, the electronic device can select the most aesthetically pleasing image from multiple images similar to the guiding parameters and provide that image to the user, resulting in a better user experience. It should be noted that the selection of the result image from the candidate image set is not limited to using an aesthetic evaluation algorithm. The electronic device can determine the result image in other ways. For example, it can also display images from the candidate image set, receive user selection input, and determine the result image. Another example is that the electronic device can randomly select the result image from the candidate image set. Yet another example is that the electronic device can use a saliency detection algorithm to select the result image from the candidate image set. For example, when an electronic device detects that an image in the candidate image set is a person image, the electronic device can also use a smile detection algorithm to filter the result image from the candidate image set. This application embodiment does not limit this.
[0031] In some examples, after determining the user-selected image set and guidance parameters, the electronic device can split the image set into one or more image subsets. The electronic device can then classify these image subsets according to scene, obtaining the scene type for each subset. Based on the scene type, the electronic device can determine the image subsets that match the scene type of the guidance parameters. Based on the image subsets with the same scene type as the guidance parameters, the electronic device can determine a candidate image set. This allows the electronic device to filter out image subsets with the same scene type as the guidance parameters from the image subsets, and then determine the candidate image set from these subsets, thereby reducing the computational workload of similarity calculations between image subsets and guidance parameters, saving computational resources and energy. In other examples, the electronic device can classify each image in the image set according to scene, obtaining the scene type for each image. Based on the scene type, the electronic device can determine images with the same scene type as the guidance parameters. Based on the images with the same scene type as the guidance parameters, the electronic device can determine a candidate image set. The scene type can be a preset scene type, such as, but not limited to, flowing hair, hearty laughter, splashing water, looking back, couples gazing lovingly at each other, raising a glass, celebrating with arms raised, throwing a hat at graduation, lifting someone up high, parent-child kissing, parent-child holding hands, childlike fun, horizon, vanishing line, street photography, cycling street photography, fireworks, lightning, badminton smash, skateboarding, skiing, shooting, layup, dunking, surfing, etc. Optionally, when the electronic device identifies that the scene type of an image subset does not belong to a preset scene type, it can use an image recognition algorithm to identify the text description of the images in the image subset and determine a new scene type based on the text description. In this way, the electronic device can add scene types during the scene classification process, allowing it to handle images from a wider range of scenes rather than being limited to a few.
[0032] Figure 1 A schematic diagram of the electronic device is shown.
[0033] The electronic device can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, in-vehicle device, smart home device and / or smart city device. The embodiments of this application do not impose any special restrictions on the specific type of the electronic device.
[0034] like Figure 1 As shown, the electronic device may include, but is not limited to, a processor 11, a memory 12, a display screen 13, a camera 14, etc.
[0035] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0036] Processor 11 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0037] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.
[0038] The processor 11 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 11 is a cache memory. This memory can store instructions or data that the processor 11 has just used or that are used repeatedly. If the processor 11 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 11, and thus improves the efficiency of the system.
[0039] In some embodiments, the processor 11 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0040] The MIPI interface can be used to connect the processor 11 to peripheral devices such as the display screen 13 and the camera 14. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 11 and the camera 14 communicate via the CSI interface to enable the electronic device to capture images. The processor 11 and the display screen 13 communicate via the DSI interface to enable the electronic device to display images.
[0041] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments of this application, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0042] The memory 12 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM). The RAM may include static random-access memory (SRAM), dynamic random-access memory (DRAM), synchronous dynamic random-access memory (SDRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM, such as fifth-generation DDR SDRAM, generally referred to as DDR5 SDRAM), etc.; the NVM may include disk storage devices and flash memory. Flash memory can be classified according to its operating principle, including NOR FLASH, NAND FLASH, 3D NAND FLASH, etc.; according to the level of its storage cells, including single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc.; and according to its storage specification, including universal flash storage (UFS) and embedded multimedia card (eMMC), etc. Random access memory (RAM) can be directly read and written by the processor 11. It can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data. Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into RAM for direct reading and writing by the processor 11.
[0043] The electronic device implements display functions through a GPU, a display screen 13, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 13 and the application processor. The GPU performs mathematical and geometric calculations and is used for graphics rendering. The processor 11 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0044] The display screen 13 is used to display images, videos, etc. The display screen 13 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include one or N displays 13, where N is a positive integer greater than 1.
[0045] Electronic devices can achieve shooting functions through ISP, camera 14, video codec, GPU, display 13 and application processor.
[0046] The ISP is used to process data fed back from the camera 14. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 14.
[0047] Camera 14 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device may include one or N cameras 14, where N is a positive integer greater than 1.
[0048] In this embodiment, the processor 11 can be used to acquire a user-selected image set and guidance parameters. The image set may include video frames of a specified video, or multiple images. The guidance parameters may include, but are not limited to, one or more of guidance text, guidance images, and guidance audio. The guidance parameters can be used to acquire the image most similar to the guidance parameters from the image set. The processor 11 can also acquire the video stream being captured by the user from the camera 14 and use the images of the video stream as the image set. Based on the guidance parameters, the processor 11 can determine a candidate image set from the image set, which includes the image most similar to the guidance parameters. The processor 11 can use an aesthetic evaluation algorithm to filter the candidate image set to obtain a result image. The result image is the image with the highest aesthetic score in the candidate image set, or the result image can be an image in the candidate image set with an aesthetic score greater than a preset aesthetic score. The processor 11 can also send the result image to the display screen 13.
[0049] The display screen 13 can be used as an interface to display the result image.
[0050] Camera 14 can capture the video stream that the user is shooting.
[0051] The following is a flowchart illustrating an image processing method provided in an embodiment of this application.
[0052] For example, such as Figure 2 As shown, the image processing method includes the following steps:
[0053] S201. The electronic device determines the set of images selected by the user and the guiding parameters.
[0054] The image set can be video frames from a video or multiple images; this application embodiment does not limit this. The multiple images can be multiple images selected by the user from the electronic device's image library. For example, the electronic device can receive input from the user selecting a video and determine the image set, which includes the video frames of that video. The image processing method provided in this application embodiment will be described below using the example of the image set being video frames from a video.
[0055] The guidance parameters may include, but are not limited to, one or more of the following: guidance text, guidance image, and guidance audio. The guidance parameters can be used to obtain the image most similar to the guidance parameters from the image set.
[0056] The guiding text can be one or more of the following: template text, user-defined input text, text input locally on the electronic device, and text obtained from a web search. Template text can include one or more text examples provided by the electronic device. User-defined input text can be text typed or voice-input by the user. Text input locally on the electronic device can be text stored on the electronic device (e.g., text from a memo app, text from a file app, etc.). Text obtained from a web search can be text retrieved by the electronic device based on user-provided search terms, and so on. For example, the guiding text can be descriptive statements in styles such as Tang poetry, Song lyrics, Yuan drama, modern poetry, and prose. The electronic device can filter images from an image set that match the description in the guiding text. For example, when the guiding text includes "turn head and smile," the electronic device can filter images from the image set that include a subject turning their head and smiling.
[0057] The guide image can be one or more of the following: a template image, a user-defined input image, an image input locally from the electronic device, or an image obtained from a web search. A template image can include one or more image examples provided by the electronic device. An image input locally from the electronic device can be an image stored on the electronic device (e.g., an image from a gallery application, a video from a video application, etc.). A user-input image can be an image drawn by the user and received by the electronic device via a touchpad or input device (e.g., a stylus, a graphics tablet, etc.). Images obtained from a web search can include images retrieved by the electronic device based on user-provided search terms. The electronic device can filter images from an image set that are similar to the guide image. For example, when the guide image includes "a subject turning and smiling," the electronic device can filter images from the image set that include "a subject turning and smiling."
[0058] The guiding audio can be one or more of the following: template audio, user-input audio, audio input locally from the electronic device, and audio obtained from a web search. For example, the text corresponding to the speech in the template audio could include "smile." The template audio can include one or more audio examples provided by the electronic device. User-input audio can be user audio captured by the electronic device through a microphone. Audio input locally from the electronic device can be audio stored on the electronic device (e.g., audio from a recorder application). Audio obtained from a web search can include audio retrieved by the electronic device based on user-provided search terms. The guiding audio can also be audio derived from text conversion, and so on. The electronic device can use the guiding audio to filter images from an image set that match the description in the guiding audio. For example, when the guiding audio includes the phrase "turn head and smile," the electronic device can filter images from the image set that include a subject turning their head and smiling.
[0059] S202. The electronic device splits the image set into one or more image subsets.
[0060] In some examples, the electronic device can split the image set into one or more image subsets according to a preset number of images (e.g., 16). For example, when the image set includes M images, where M is a positive integer, the number of image subsets is... The number of images in the image subset is less than or equal to the preset number of images.
[0061] S203. The electronic device classifies the image subsets according to the scene to obtain the scene type of each image subset.
[0062] Electronic devices can use image scene classification algorithms to determine the scene type to which a subset of images belongs. For example, scene type can include, but is not limited to, a combination of one or more parameters such as the subject in the image, the subject's actions, the subject's emotions, the subject's facial expressions, and the shooting scene. For instance, scene type could be "people cheering on a football field," "people turning around and smiling," etc.
[0063] For example, an electronic device can use an image scene classification algorithm to average the feature information of all images in an image subset to obtain the average feature information of that image subset. The average feature information includes the features of all images in the image subset. The electronic device can use the image scene classification algorithm to determine the scene type corresponding to the average feature information and use that scene type as the scene type of the image subset. It should be noted that the determination of the scene type of the image subset is not limited to the average feature information. For example, the electronic device can also determine the scene type of the first image in the image subset and use that scene type as the scene type of the image subset. Alternatively, the electronic device can determine the scene type of a randomly selected image in the image subset and use that scene type as the scene type of the image subset, and so on. This application embodiment does not limit this.
[0064] In some examples, the electronic device can also use an image encoder to encode the images in a subset of images, obtaining a fragment vector for the subset. For instance, the electronic device can average the image vectors of all images in the subset to obtain the fragment vector. In some examples, the electronic device can use pooling operations to obtain fragment vectors based on the images in the subset. The fragment vector can then represent the features of all images in the subset. For another example, the electronic device can use an image encoder to encode the Xth image (where X is a positive integer) in the subset, obtaining an image vector, and use this image vector as the fragment vector of the subset, and so on. Then, the electronic device can determine the scene type of the subset based on the fragment vector using a scene classification algorithm. In this way, since some images in the subset may have different sizes, calculating the average of the image vectors can both preserve the features of each image and reduce the number of calculations required by the electronic device.
[0065] It should be noted that electronic devices can also determine the scene type of an image subset based on fragment vectors using other methods, and this application embodiment does not limit this. For example, the electronic device can also use a text encoder to obtain guiding text vectors corresponding to the text descriptions of each scene type. The electronic device can calculate the similarity between the fragment vector and the text vectors of each scene based on the fragment vector to determine the scene type corresponding to the fragment vector, that is, to determine the scene type to which the image subset belongs. For example, the electronic device can calculate the cosine similarity between the fragment vector and the text vector. The closer the cosine similarity value is to 1, the higher the similarity between the fragment vector and the text vector; the closer the cosine similarity value is to 0, the lower the similarity between the fragment vector and the text vector. As another example, the electronic device can also use the absolute value of the difference obtained by subtracting the fragment vector from the text vector as the similarity value. The smaller the absolute value of the difference between the fragment vector and the text vector, the higher the similarity between the fragment vector and the text vector; the larger the absolute value of the difference between the fragment vector and the text vector, the lower the similarity between the fragment vector and the text vector, and so on. Electronic devices can determine the scene type of an image subset by identifying the text description of the scene corresponding to the text vector most similar to the fragment vector.
[0066] S204. The electronic device determines a subset of target images that are the same as the scene type in the guidance parameters based on the scene type.
[0067] After determining the scene type of the image subset, the electronic device can determine the scene type of the guidance parameters. When the guidance parameters include a guidance image, the electronic device can determine the scene type of the guidance image through a scene recognition algorithm or an image encoder, etc. For details, please refer to the description of step S203, which will not be repeated here.
[0068] When the guidance parameters include guidance text, the electronic device can extract key information (token) from the guidance text and determine the scene type of the guidance text based on this key information. Similarly, when the guidance parameters include guidance audio, the electronic device can also extract key information from the guidance audio and determine the scene type of the guidance audio based on this key information.
[0069] The electronic device can compare the scene type of an image subset with the scene type of the guiding parameters to determine the target image subset that has the same scene type as the guiding parameters. Specifically, when the electronic device determines that the scene type of the image subset is the same as or similar to the scene type of any parameter in the guiding parameters, it determines that the image subset is a target image subset with the same scene type as the guiding parameters. When the electronic device determines that the scene type of the image subset is neither the same as nor similar to the scene type of any parameter in the guiding parameters, it determines that the image subset is not a target image subset. In this way, the electronic device can obtain a target image subset with the same scene type as the guiding parameters.
[0070] In other examples, the electronic device can use an image encoder to encode each image in the image set, obtaining an image vector corresponding to each image. Based on the image vectors, the electronic device then calculates the similarity between the image vectors and various scenes to determine the scene type corresponding to each image vector, i.e., to determine the scene type to which each image in the image set belongs. Based on the scene type to which each image in the image set belongs, the electronic device then splits the image set into one or more image subsets, where the images in each subset belong to the same scene type. Based on the scene types of the one or more image subsets, the electronic device determines a target image subset that matches the scene type specified in the guidance parameters. In some examples, the number of images in the target image subset is less than or equal to a preset number of images.
[0071] In some examples, the electronic device can determine the scene type to which each image in the image set belongs. Based on the scene type of each image in the image set, the electronic device then determines images with the same scene type as the guiding parameters, and splits these images with the same scene type as the guiding parameters into one or more target image subsets according to the temporal order of the images (e.g., the playback order of video frames, the capture time of multiple images). In some examples, the number of images in the target image subsets is less than or equal to a preset number of images.
[0072] In some examples, the scene types of the guidance parameters determined by the electronic device are different. After obtaining the scene types of the image subsets, the electronic device can filter out one or more target image subsets that are the same as the scene types of the guidance parameters. The electronic device can also use guidance parameters with the same scene types as the target image subsets to calculate the similarity between the target image subsets and the guidance parameters. Based on the similarity between the one or more target image subsets and the guidance parameters with the same scene types, the electronic device can select one or more target image subsets with the highest similarity as a candidate image set. Specifically, if a target image subset corresponds to one of the guidance text, guidance image, and guidance audio, the similarity with that guidance parameter can be used as the comprehensive score of the target image subset. If a target image subset corresponds to one of the guidance text, guidance image, and guidance audio, the sum of the products of the similarity of the target image subset with multiple guidance parameters and the weight values of multiple guidance parameters can be used as the comprehensive score of the target image subset. In this way, when the scene types of multiple guidance parameters added by the user are different, the electronic device can perform scene matching on the images in the image set, calculate the similarity of some images using guidance parameters with the same scene types, and obtain a candidate image set.
[0073] S205. The electronic device determines a candidate image set based on a subset of target images, the candidate image set including the subset of target images most similar to the guiding parameters.
[0074] The electronic device can, based on guiding parameters, select the subset of images most similar to the guiding parameters from multiple subsets of images of the same scene type; this subset serves as a candidate image set. Alternatively, the electronic device can, based on guiding parameters, select a target subset of images from multiple subsets of images of the same scene type whose similarity to the guiding parameters is greater than a preset similarity; this target subset serves as a candidate image set. It is understood that when the number of image subsets most similar to the guiding parameters obtained by the electronic device is greater than one, the electronic device can use these multiple image subsets as the candidate image set.
[0075] For example, the electronic device can encode the guiding parameters to obtain guiding vectors corresponding to the guiding parameters, whereby the guiding vectors represent the features of the guiding parameters. The electronic device can also encode images from a subset of images of the same scene type to obtain fragment vectors for that subset. Based on the guiding vectors and fragment vectors, the electronic device can determine the subset of images corresponding to the fragment vectors most similar to the guiding vectors, and use this subset as a candidate image set. The fragment vector of the image subset can be the average of the image vectors of all images in the subset.
[0076] The electronic device can calculate the cosine similarity between the guide vector and the fragment vector. The numerical value of the cosine similarity represents the similarity between the fragment vector and the guide vector. Specifically, the formula for the similarity is as follows:
[0077]
[0078] Here, A can represent the guiding vector, and B can represent the fragment vector. Similarity is the cosine similarity between the guiding vector and the fragment vector. The closer the cosine similarity value is to 1, the higher the similarity between the guiding vector and the fragment vector. The closer the cosine similarity value is to 0, the lower the similarity between the guiding vector and the fragment vector. It should be noted that the similarity between the guiding vector and the fragment vector is not limited to cosine similarity. For example, electronic devices can also determine the similarity between the guiding vector and the fragment vector by subtracting the guiding vector from the fragment vector and obtaining the absolute value of the difference. Here, the smaller the absolute value of the difference, the higher the similarity between the guiding vector and the fragment vector; the larger the absolute value of the difference, the lower the similarity between the guiding vector and the fragment vector.
[0079] When the guidance parameters include multiple guidance images, the electronic device can calculate the similarity between the fragment vectors of a subset of images and the guidance vector of each guidance image. The electronic device can use the average (or maximum, minimum, etc.) similarity between the fragment vectors of the subset of images and each guidance image as the value of the similarity between that subset of images and the guidance images. Similarly, when the guidance parameters include multiple guidance texts, the electronic device can calculate the similarity between the fragment vectors of a subset of images and the guidance vector of each guidance text. The electronic device can use the average (or maximum, minimum, etc.) similarity between the fragment vectors of the subset of images and each guidance text as the value of the similarity between that subset of images and the guidance text. When the guidance parameters include multiple guidance audios, the electronic device can calculate the similarity between the fragment vectors of a subset of images and the guidance vector of each guidance audio. The electronic device can use the average (or maximum, minimum, etc.) similarity between the fragment vectors of the subset of images and each guidance audio as the value of the similarity between that subset of images and the guidance audio.
[0080] When the guidance parameters include multiple items such as guidance text, guidance image, and guidance audio, the electronic device can set weighting coefficients for the similarity between the fragment vector and the guidance text vector, the fragment vector and the guidance image vector, and the fragment vector and the guidance audio vector. The similarity of each guidance parameter is then weighted and averaged to calculate a comprehensive score. The electronic device can then select the image subset with the highest comprehensive score from the image subset to obtain a candidate image set. It should be noted that the comprehensive score is not limited to setting weighting coefficients for different guidance parameters. The electronic device can also use the average of the similarity between the fragment vector and the guidance text vector, the fragment vector and the guidance image vector, and the fragment vector and the guidance audio vector as the comprehensive score. Alternatively, the electronic device can use the maximum (or minimum) value among the similarity between the fragment vector and the guidance text vector, the fragment vector and the guidance image vector, and the fragment vector and the guidance audio vector as the comprehensive score for the image subset, etc. This embodiment of the application does not limit this approach.
[0081] In other examples, the electronic device can also encode image vectors based on each image in the image subset. Then, based on the image vectors and the guiding vector, a candidate image set is obtained, which includes the image corresponding to the input image vector with the highest similarity to the guiding vector from each image subset. In this way, the electronic device can filter out the images from all image subsets that are most similar to the guiding parameters.
[0082] In some examples, the electronic device can set a preset similarity level, allowing it to filter candidate images from a subset of images with a comprehensive score greater than the preset similarity level. The preset similarity level can be set based on the similarity values of each image subset or individual image. For example, the electronic device can use the median or average of the similarity scores of the image subsets or images as the preset similarity level. Alternatively, the electronic device can set a fixed initial value for the preset similarity level. The electronic device can lower the preset similarity level when the number of images filtered according to the fixed value is less than a preset number 1. Conversely, the electronic device can increase the preset similarity level when the number of images filtered according to the fixed value is greater than a preset number 2. This dynamic adjustment of the preset similarity level allows the electronic device to filter a subset of images from the image set based on the guidance parameters, even when the similarity between the images and the guidance parameters is low. Optionally, when the electronic device determines that the similarity between all image subsets and the guidance parameters is less than the preset similarity level, it can display a specified prompt message. This specified prompt message informs the user that the image set does not include images that meet the guidance parameters. For example, the specified prompt message could be a text message such as: "No matching images were found from the selected image set."
[0083] Optionally, when the electronic device determines that the similarity between all image subsets and the guiding parameters is less than a preset similarity, the electronic device may also display a prompt box. The prompt box may include the image subset most similar to the guiding parameters, as well as a specified control. The specified control may be used to trigger the electronic device to execute step S206.
[0084] It should also be noted that the determination of the similarity between the image subset and the guidance parameters is not limited to using guidance vectors and fragment vectors. For example, when the guidance parameters include a guidance image, the electronic device can also use image processing algorithms (e.g., convolutional neural network algorithms) to calculate the similarity between the guidance image and the image subset. For instance, the electronic device can calculate the similarity between the features of the guidance image and the average of the image features of all images in the image subset, and use this similarity as the similarity between the guidance image and the image subset. Alternatively, the electronic device can calculate the similarity between the guidance image and all images in the image subset, and use the average (or maximum, minimum, etc.) of the similarity of all images as the similarity between the guidance image and the image subset, and so on. As another example, when the guidance parameters include guidance audio, the electronic device can use an audio-to-text algorithm to convert the guidance audio into guidance text. The electronic device can also use an image recognition algorithm to identify the descriptive text for each image in the image subset. The electronic device can compare the similarity between the guidance text and the descriptive text, and use this similarity as the similarity between the guidance audio and the image subset. As also exemplarily, the electronic device may use key information (token) of guiding audio or guiding text, and calculate the similarity between images of a subset of images and key information, etc., which is not limited in the embodiments of this application.
[0085] In some examples, after performing step S202, the electronic device can determine a candidate image set based on the image subsets. The candidate image set includes one or more image subsets that are most similar to the guiding parameters. In this way, the electronic device does not need to perform scene type detection; that is, the electronic device can directly calculate the similarity between the guiding parameters and the image subsets to obtain the candidate image set without performing steps S203 and S204.
[0086] In some examples, when the electronic device detects that the number of images in the image set is less than a preset threshold (e.g., 16), after executing step S201, the electronic device can directly calculate the similarity between each image in the image set and the guiding parameter, and select the one or more images most similar to the guiding parameter as a candidate image set. For example, the electronic device can obtain an image vector for each image in the image set. The electronic device can obtain a guiding vector from the guiding parameter. Based on the guiding vector and the image vector of each image, the electronic device determines the similarity between the guiding vector and the image vector. Specifically, please refer to the above embodiments, which will not be repeated here. The electronic device can determine the one or more images corresponding to the image vector most similar to the guiding vector as a candidate image set. In this way, the electronic device can reduce the amount of computation required to obtain the image scene type, and since the number of images in the image set is small, the electronic device does not need to split the image set into one or more image subsets, saving resources.
[0087] In some examples, not limited to the above description, the electronic device can also obtain a candidate image set through multiple filtering based on guiding parameters. For example, if the image set includes M images, the electronic device can split the image set into X image subsets according to a preset number of images 21, where X equals... The number of images in each image subset is less than or equal to a preset number of images, 21. The electronic device can, based on each image subset and the guiding parameters, obtain M image subsets most similar to the guiding parameters, where M is a positive integer. Then, the electronic device can split the M image subsets into Y image subsets according to a preset number of images, 22, where the preset number of images, 22, is less than the preset number of images, 21, and Y is less than or equal to... The electronic device can obtain N image subsets most similar to the guiding parameters based on each image subset, where N is a positive integer. This process continues until the electronic device, based on the guiding parameters, selects a specified number (e.g., 6) of images from the image set. Specifically, the description of how the electronic device obtains the candidate image set based on the fragment vectors of the image subsets and the guiding parameters can be found in the above embodiments and will not be repeated here. Thus, when there are many images in the image set, the electronic device can obtain the candidate image set through multiple selections.
[0088] In some examples, the electronic device can receive multiple guide images. An image encoder obtains guide vectors for each guide image. The electronic device can then perform sampling operations on these guide vectors. For example, it can use pooling operations such as max pooling, average pooling, k-max pooling, and chunk-max pooling to sample multiple guide vectors and obtain dimensionality-reduced guide vectors. Based on these dimensionality-reduced guide vectors, the electronic device can determine a set of candidate images. In this way, the electronic device can retain the main features of the images while reducing computational parameters and computational cost, thus preventing overfitting.
[0089] S206. Electronic devices use an aesthetic evaluation algorithm to select the result image from the candidate image set. The result image is the image with the highest aesthetic score in the candidate image set.
[0090] The electronic device uses an aesthetic evaluation algorithm to score each image in the candidate image set, obtaining an aesthetic score for each image. The electronic device then selects the image with the highest aesthetic score from the image set as the result image. Thus, after obtaining the image set with the highest similarity to the guiding parameters, the electronic device finds that each image in this set also has high similarity, making it impossible to select a result image from this set. Instead, the aesthetic evaluation algorithm obtains the image with the highest aesthetic score, and this image is used as the result image.
[0091] Optionally, the electronic device can set a preset aesthetic score, where the aesthetic score of the selected image from the candidate image set is greater than the preset aesthetic score. When the aesthetic score of the selected image is less than the preset aesthetic score, the electronic device can display a specified prompt message, which indicates to the user that the image set does not include images that meet the aesthetic evaluation criteria. For example, the specified prompt message can be a text message: "No image meeting the requirements was selected from the image set." Optionally, when the electronic device determines that the aesthetic scores of all images in the candidate image set are less than the preset aesthetic score, the electronic device can also display a prompt box, which can include the image with the highest aesthetic score and a prompt message. The prompt message indicates to the user that the image is the one with the highest aesthetic score in the candidate image set. The prompt box can also include one or more images from the candidate image set for the user to select. Specifically, the description of training the aesthetic evaluation model and the aesthetic evaluation model can be found in the following embodiments and will not be repeated here.
[0092] In some examples, the electronic device can determine a preset aesthetic score based on the aesthetic score of each image in the candidate image set. For example, the electronic device can use the median or average aesthetic score of the images as the preset aesthetic score. Alternatively, the electronic device can set a fixed initial value for the preset aesthetic score. The electronic device can lower the preset aesthetic score when the number of images filtered according to the fixed value is less than a preset number 1. The electronic device can raise the preset aesthetic score when the number of images filtered according to the fixed value is greater than a preset number 2. In this way, the electronic device dynamically adjusts the preset aesthetic score, allowing it to filter a subset of images from the candidate image set even when the aesthetic scores of all images are low.
[0093] In some examples, if the candidate image set obtained by the electronic device in step S205 contains only one image, the electronic device can use that image as the result image without scoring it using an aesthetic evaluation algorithm.
[0094] In this way, electronic devices can process the set of images selected by the user based on the guidance parameters selected by the user, and filter out images that meet the user's personalized and diverse needs from the set of images, thereby obtaining images with a more diverse number and types of scenes.
[0095] The following is a flowchart illustrating another image processing method provided by an embodiment of the present invention.
[0096] For example, such as Figure 3 As shown, the electronic device includes multiple modules, which can implement the image processing method provided in this application embodiment. These multiple modules may include, but are not limited to, one or more guidance modules, one or more encoders, a scene detection module, one or more scene matching modules, one or more matching models, a key segment selection module, and a key image selection module.
[0097] One or more guidance modules can be used to determine guidance parameters and send them to the encoder. For example, these guidance modules may include, but are not limited to, text guidance (guided_texts) modules, image guidance (guided_images) modules, and audio guidance (guided_audio) modules. The text guidance module can be used to determine the guidance text. The image guidance module can be used to determine the guidance image. The audio guidance module can be used to determine the guidance audio. The electronic device can obtain the guidance text, guidance image, and guidance audio through a web search and input them into the corresponding guidance module. For example, Figure 3The arrow pointing from the Chinese text guidance module to the image guidance module indicates that the text guidance module receives user-input keywords, searches for images based on those keywords, and sends the found images to the image guidance module. Similarly, the text guidance module receives user-input keywords, searches for audio based on those keywords, and sends the found audio to the audio guidance module.
[0098] Optionally, the text guidance module can receive text input by the user (e.g., typed via an electronic device or voice input) and transmit the input text to the image guidance module. The image guidance module uses the input text as keywords to perform a web search, searching for images that match the text description, and uses the image selected by the user as the guidance image. Alternatively, the text guidance module can receive text input by the user and transmit the input text to the audio guidance module. The audio guidance module uses the input text as keywords to perform a web search, searching for audio that matches the text description, and uses the audio selected by the user as the guidance audio.
[0099] One or more encoders can be used to convert guiding parameters into guiding vectors. For example, one or more encoders may include, but are not limited to, a text encoder, an image encoder, and an audio encoder. The text encoder can process the guiding text to obtain a corresponding guiding text vector (guided_text_embedding), which can be used to represent the features of the guiding text. The image encoder can obtain vectors corresponding to a set of images. For example, an electronic device can decode a video file or a video stream acquired by a camera, perform frame extraction, etc., to obtain a set of images, which includes one or more image subsets. The image encoder can obtain clip vectors (clip_embeddings) corresponding to the image subsets, which can be used to represent the features of the images in the image subsets. The image encoder can also obtain a guiding image vector (guided_image_embedding) based on the guiding image, which can be used to represent the features of the guiding image. The audio encoder can process the guiding audio to obtain a corresponding guiding audio vector (guided_audio_embedding), which can be used to represent the features of the guiding audio.
[0100] The scene detection module calculates the similarity between the input video segment vectors and various scenes, determining the scene type corresponding to each segment vector. This identifies the scene type of the image subset, such as splashing water, flying hair, or flapping wings. The scene detection module can then send the scene type of the image subset to one or more matching modules.
[0101] One or more matching modules can be used to filter image subsets that share the same scene type as the guidance parameters. These matching modules may include, but are not limited to, text strategy matching modules, image strategy matching modules, and audio strategy matching modules. Specifically, the text strategy matching module identifies target image subsets from one or more image subsets that share the same scene type as the guidance text; the image strategy matching module identifies target image subsets from one or more image subsets that share the same scene type as the guidance image; and the audio strategy matching module identifies target image subsets from one or more image subsets that share the same scene type as the guidance audio. It is understood that after the electronic device matches image subsets with guidance parameters using the matching modules, the electronic device can calculate the similarity between the guidance parameters and the matched target image subsets using one or more matching models.
[0102] In some examples, one or more encoders can directly send vectors to the corresponding matching model. Alternatively, one or more encoders can directly send vectors to the corresponding matching model when the number of images in the image set is less than a preset threshold (e.g., 16). In this way, the electronic device can directly calculate similarity without performing scene matching.
[0103] The matching model may include, but is not limited to, an image-text matching model (ITMoel), an image-image matching model (IIMoel), and an image-audio matching model (IAMoel). Specifically, when the text matching strategy module sends a guiding text vector to the image-text matching model, the scene detection module may send fragment vectors of the target image subset that matches the guiding text to the image-text matching model. Similarly, when the image matching strategy module sends a guiding image vector to the image-image matching model, the scene detection module may send fragment vectors of the target image subset that matches the guiding image to the image-image matching model. Likewise, when the audio matching strategy module sends guiding audio to the image-audio matching model, the scene detection module may send fragment vectors of the target image subset that matches the guiding audio to the image-audio matching model.
[0104] Specifically, ITMoel processes the fragment vectors provided by the scene detection module with the guiding text vectors to obtain an image-text matching score. The image-text matching score represents the similarity between the fragment vectors and the guiding text vectors; a higher score indicates a higher similarity. IIModel processes the fragment vectors provided by the scene detection module with the image vectors to obtain an image-image matching score. The image-image matching score also represents the similarity between the fragment vectors and the image vectors; a higher score indicates a higher similarity. IAModel processes the fragment vectors provided by the scene detection module with the guiding audio vectors to obtain an image-audio matching score. The image-audio matching score also represents the similarity between the fragment vectors and the guiding audio vectors; a higher score indicates a higher similarity.
[0105] The key clip selection module (key_clip_selector) can obtain a comprehensive score for a subset of target images based on the scores for image-text matching, image-image matching, and image-audio matching. Based on this comprehensive score, the module can also determine a candidate image set. The candidate image set is the subset of target images with the highest comprehensive score. For example, the module can assign different weight coefficients to the image-text matching, image-image matching, and image-audio matching scores to calculate the comprehensive score. For instance, it could assign a weight coefficient of 0.5 to the image-text matching score, 0.25 to the image-image matching score, and 0.25 to the image-audio matching score. Since users tend to obtain information visually, assigning a higher weight to the image-image matching score can provide users with result images that better meet their requirements. In some examples, the module can also set a preset similarity level, filtering out target image subsets with comprehensive scores lower than the preset similarity level. Alternatively, the key segment selection module can use a subset of target images with a comprehensive score higher than the preset similarity as a candidate image set.
[0106] The key frame selector module can obtain the result image based on the candidate image set provided by the key segment selection module. Specifically, the key frame selector module can use an aesthetic evaluation model to score all images in the candidate image set, obtaining an aesthetic score for each image. The key frame selector module can select the image with the highest aesthetic score as the result image. In some examples, the key frame selector module can set a preset aesthetic score (e.g., half of the maximum aesthetic score). The key frame selector module can then determine the result image with the highest aesthetic score from images whose aesthetic scores are greater than the preset aesthetic score. For example, when the electronic device uses decimals between 0 and 1 to represent aesthetic scores, the preset aesthetic score can be 0.5. For details regarding the training of the aesthetic evaluation model and the description of the aesthetic evaluation model, please refer to the subsequent embodiments; they will not be repeated here.
[0107] In this embodiment, the electronic device can process the image set of the input video using an image encoder to obtain one or more segment vectors. The electronic device can then determine the scene type corresponding to one or more image subsets using a scene detection module. For details, please refer to... Figure 2 The embodiments shown are not described in detail here.
[0108] The electronic device can receive guiding text through a text guidance module, and then process the guiding text through a text encoder to obtain a guiding text vector. The electronic device can use a text strategy matching module to identify target image subsets from one or more image subsets that have the same scene type as the guiding text, and then output fragment vectors of these target image subsets with the same scene type as the guiding text through a scene detection module. The electronic device can use an image-text matching model to obtain an image-text matching score, which represents the similarity between the fragment vector and the guiding text vector, i.e., the similarity between the target image subset and the guiding text.
[0109] The electronic device can receive a guidance image through an image guidance module, and obtain a guidance image vector corresponding to the guidance image through an image encoder. The electronic device can also use an image strategy matching module to identify target image subsets from one or more image subsets that have the same scene type as the guidance image, and output fragment vectors of these target image subsets with the same scene type as the guidance image through a scene detection module. Finally, the electronic device can use an image-to-image matching model to obtain an image-to-image matching score, which represents the similarity between the fragment vector and the guidance image vector, i.e., the similarity between the target image subset and the guidance image.
[0110] The electronic device can also receive guiding audio through an audio guidance module, and obtain a guiding audio vector corresponding to the guiding audio through an audio encoder. The electronic device can use an audio strategy matching module to identify target image subsets from one or more image subsets that have the same scene type as the guiding audio, and output fragment vectors of the target image subsets with the same scene type as the guiding audio through a scene detection module. The electronic device can obtain an image-audio matching score through an image-audio matching model. The image-audio matching score is used to represent the similarity between the fragment vector and the guiding audio vector, that is, the similarity between the target image subset and the guiding audio.
[0111] Electronic devices can use a key segment selection module to perform weighted calculations on image-text matching scores, image-image matching scores, and image-audio matching scores to obtain a comprehensive score. The electronic device then selects one or more target image subsets with the highest comprehensive scores as candidate image sets. Furthermore, the electronic device can use a key image selection module to perform aesthetic evaluations on each image in the candidate image set, selecting the one or more images with the highest aesthetic scores as the result images.
[0112] It should be noted that the number and function of the above modules are merely examples. Electronic devices may also implement the image processing method provided in this application embodiment with more or fewer modules, and this application embodiment does not limit this.
[0113] In this way, through the processing of various modules, the electronic device can process the input video based on the user's input guidance text, guidance images and guidance audio, and filter out the result images that meet the user's personalized and diverse needs from the video, further making the number and types of scenes captured more diverse.
[0114] The following is a schematic diagram of a training encoder provided in an embodiment of this application.
[0115] First, the cloud server can acquire multiple sets of training data, each set including an image and its corresponding text. The cloud server can then input the text from these training sets into a text encoder and the images into an image encoder. The image encoder processes the input images to obtain corresponding image vectors, and the text encoder processes the input text to obtain corresponding guiding text vectors. Next, the cloud server can calculate the similarity between the image and the text based on the image vectors and the guiding text vectors. Finally, the cloud server can modify the parameters of the text encoder and image encoder based on the training data to maximize the similarity within the same training set while minimizing the similarity between different training sets, resulting in well-trained text and image encoders.
[0116] The cloud server can crawl images and matching text from web pages and video sharing websites containing various images, and perform cleaning, deduplication, and filtering operations on the crawled images and text to obtain training data. And / or the cloud server can randomly generate text and then use the text as search terms to search for images corresponding to the text, obtaining training data. And / or, the cloud server can search a large number of images and generate a large amount of training data using a large language model. It should be noted that, not limited to cloud servers, electronic devices can train encoders based on training data; this application embodiment does not limit this. It should also be noted that the cloud server can obtain image data through manual annotation. Specifically, the cloud server can search a large number of images and manually annotate the descriptive text of the images to obtain training data.
[0117] For example, cloud servers can be based on Figure 4 The model shown trains an image encoder and a text encoder. The cloud server can obtain N sets of training data. The cloud server can then use the text (Text1, ..., Text...) from these N sets of training data... N This yields N guiding text vectors (T1, T2, T3, ..., T...). N The cloud server can use the images (Image1, ..., Image2) from the N sets of training data. N This yields N image vectors (I1, I2, I3, ..., I...).N A cloud server can use N image vectors as column vectors and N guiding text vectors as row vectors, and then perform matrix multiplication to obtain... Figure 4 The similarity matrix shown is a matrix where the values are the cosine similarity between the image vector and the guiding text vector calculated by the cloud server. Figure 4 The N sets of training data shown include N positive samples, i.e., text and image belonging to a pair. The indices of these N positive samples are one-to-one, and the cosine similarity of the positive samples lies on the main diagonal of the matrix. These N sets of training data also include (N*NN) negative samples. The training objective of the cloud server is to maximize the similarity of the positive samples (I1·T1, I2·T2, I3·T3, ..., I...). N ·T N The value of ) is minimized, while minimizing the similarity of negative samples (I1·T2, I1·T3, ..., I N ·T N-1 The value of ) is used to obtain the trained image encoder and text encoder. Similarly, the cloud server can also... Figure 4 The text encoder in the code is replaced with an audio encoder, and a trained audio encoder is obtained accordingly.
[0118] It should be noted that it is not limited to Figure 4 The electronic device can also subtract the image vector from the guiding text vector, as shown by the cosine similarity, and determine the similarity between the image vector and the guiding text vector based on the absolute value of the difference. The smaller the absolute value of the difference between the image vector and the guiding text vector, the higher the similarity between them; conversely, the larger the absolute value, the lower the similarity. This application does not impose any limitations on this aspect.
[0119] The electronic device can obtain image encoders, text encoders, and audio encoders from the cloud server. It should be noted that, not limited to image encoders, when the electronic device receives input containing images filtered from a specified video based on guiding parameters, it can obtain image vectors of video frames based on the video encoder. Similarly, the cloud server can... Figure 4 The video encoder is trained in the manner shown.
[0120] In one possible implementation, a cloud server can obtain the correspondence between images and aesthetic scores, and use this correspondence as training data to train an aesthetic evaluation model. The cloud server can then send the aesthetic evaluation model to an electronic device. The electronic device can then use the aesthetic evaluation model to select a result image from a set of candidate images. In this way, even when multiple images are selected based on a guiding image, the electronic device can score each image in the candidate image set based on the aesthetic evaluation model, and select the image with the highest aesthetic score as the result image, allowing the electronic device to select images that better match the user's aesthetic preferences.
[0121] The following is a schematic diagram of a training aesthetic evaluation model provided in an embodiment of this application.
[0122] For example, such as Figure 5 As shown, the cloud server can classify the acquired images according to the image scene. For example, the cloud server can classify images into scene types such as looking back, lightning, parent-child, and running. In this way, the cloud server can acquire images of each scene type and ensure that the number of images of each scene in the acquired images is relatively small, so that the aesthetic evaluation model can make a reasonable evaluation of the images of each scene.
[0123] The cloud server can perform subject detection on images, identifying the subject of the image and controlling the number of images of different subjects to minimize differences in the number of images of different subjects. Then, the cloud server can receive scores from annotators based on aesthetic evaluation metrics such as composition, color, lighting, exposure, and theme. The cloud server can then derive an aesthetic score for the image based on these scores. Note that the criteria used by annotators to score images based on composition metrics vary depending on the subject of the image. For example, using a score of 10 for each measurement indicator, when an annotator scores the composition of an image, the image can be divided into nine squares. If the main subject of the image is a person, and the area containing the person is in the top six squares, the annotator can score 6 or higher for composition. If the area containing the person is in the bottom six squares, the annotator can score 5 or lower for composition. Similarly, if the main subject of the image is an animal, and the area containing the animal is in the bottom six squares, the annotator can score 6 or higher for composition. It should be noted that the measurement indicators and corresponding scores are merely examples; other methods can also be used to determine the image's score.
[0124] The cloud server can collect ratings from multiple maintenance personnel for the same image and use the average of the multiple ratings as the image's aesthetic score.
[0125] Next, the cloud server can classify the images. Specifically, the cloud server can categorize the images according to scores and determine the number of images for each score. The cloud server can maintain a consistent number of images across each score range. In this way, the cloud server can ensure that the number of images in each score range is roughly equivalent, avoiding the problem of data class imbalance. Based on the similar number of images in each score range, the cloud server obtains an aesthetic evaluation model to reasonably evaluate the images for each score.
[0126] Cloud servers can train aesthetic evaluation models based on the correspondence between images and their aesthetic scores. For example, an aesthetic evaluation model built by a cloud server might look like this: Figure 6 As shown, the aesthetic evaluation model includes an input image module, an encoder, and a predicted distribution module. The input image module receives the set of images selected by the user. The encoder includes a backbone network, fully-connected layers, and a softmax layer. The encoder processes the input image set to obtain an image vector for each image. This image vector represents the score of various metrics for the image and can be used to calculate the aesthetic score. The predicted distribution module processes the image vectors output by the encoder to obtain the aesthetic score of the image. For example, the predicted distribution module can determine the probability of an image falling into each score category based on the image vectors. The predicted distribution module can then add the product of the score and the probability, i.e., calculate the expected value, to obtain the aesthetic score of the image.
[0127] Electronic devices can download an aesthetic evaluation model from a cloud server. After obtaining the model, the electronic device can use it to score all images in a candidate image set, obtaining an aesthetic score for each image. Based on these aesthetic scores, the electronic device can select a result image from the candidate image set; the result image is one or more images with the highest aesthetic scores in the candidate image set. Optionally, the result image is one or more images in the candidate image set whose aesthetic scores are greater than a preset aesthetic score. In this way, the aesthetic evaluation model can score each image in the candidate image set and use the one or more images with the highest aesthetic scores as the result image.
[0128] In one possible implementation, upon receiving specified input for a given video, the electronic device can display a designated interface for adding guiding parameters. This interface may include options for adding guiding parameters, which can trigger the electronic device to add these parameters. After receiving the input to add guiding parameters, the electronic device can filter and display result images from the video frames of the given video based on the added parameters. In this way, the electronic device can process the given video according to the user-selected guiding parameters and filter images from the video that meet the user's personalized and diverse needs, thereby obtaining images with a greater variety of scenes and types.
[0129] For example, such as Figure 7A As shown, the electronic device can display a desktop 700. The desktop 700 can include multiple application icons (e.g., settings application icon, calendar application icon, gallery application icon 701, etc.). The gallery application icon 701 can be used to trigger the display of the gallery application interface. Optionally, a status bar including icons such as a time indicator can be displayed above the desktop 700. Optionally, multiple tray icons (e.g., dialer application icon, messaging application icon, contacts application icon, camera application icon) can be displayed below the multiple application icons.
[0130] Electronic devices receive user requests Figure 7A After the input (e.g., clicking) of the gallery application icon 701 shown, the electronic device can display something like... Figure 7B The interface shown is 710. (As shown) Figure 7B As shown, interface 710 may include one or more video options, which can be used to trigger the electronic device to display a corresponding video playback interface. The one or more video options may include video option 711. Interface 710 may also include one or more image options, which can be used to trigger the electronic device to display an image corresponding to the image option. The one or more image options may include image option 712.
[0131] Electronic devices receive user requests Figure 7B After inputting the video option 711 shown, the electronic device can display as follows: Figure 7C The interface shown is 720. (As shown) Figure 7C As shown, interface 720 includes a video playback window 721. The video playback window 721 can be used to play the video corresponding to video option 711. Interface 720 also includes one or more functional controls, including but not limited to control 722. Control 722 can be used to trigger the electronic device to display an interface for adding guide parameters.
[0132] Electronic devices can receive user requests Figure 7C After input is received into control 722, the following display is shown in response to the input: Figure 7D The interface shown is 730.
[0133] For example, such as Figure 7D As shown, interface 730 may include one or more add options. These add options can be used to add guide parameters. The add options include add option 731, add option 732, and add option 733. Each add option includes an add control, which can trigger the electronic device to display an interface for adding the corresponding guide parameters. Specifically, add option 731 prompts the user to add guide text, and the add control included in add option 731 can trigger the electronic device to display an interface for adding guide text. Add option 732 prompts the user to add a guide image, and the add control 732A included in add option 732 can trigger the electronic device to display an interface for adding a guide image. Add option 733 prompts the user to add guide audio, and the add control included in add option 733 can trigger the electronic device to display an interface for adding guide audio. Optionally, the add options may include a modify control, which can trigger the electronic device to display an interface for modifying the corresponding guide parameters. For example, modify control 732B can trigger the electronic device to display an interface for modifying the guide image. Thus, after the electronic device receives the guide parameters added by the user, the added guide parameters can be modified through the modify control.
[0134] The interface 730 may further include a control 734, which can be used to trigger the electronic device to filter the resulting image from the video corresponding to video option 711 according to the guidance parameters added by the user. Optionally, the interface 730 may also display a specified prompt message, which can be used to prompt the user to add guidance parameters. The electronic device will then filter the resulting image, i.e., the highlight image, from the video according to the added guidance parameters. For example, the specified prompt message can be a text-based prompt message: "Please select guidance parameters. The phone will filter highlight images according to the selected guidance parameters."
[0135] Electronic devices receive user requests Figure 7D After inputting the control 732A as shown, the electronic device can display as follows: Figure 7E The interface shown is 740.
[0136] For example, such as Figure 7EAs shown, interface 740 may include one or more options that can be used to trigger an electronic device to display guide images from different sources. These options may include, but are not limited to, "Select Template Image" option 741, "Custom Input Image" option 742, "Local Input Image" option 743, and "Web Search Image" option 744.
[0137] The "Select Template Image" option 741 triggers the electronic device to display an interface containing a template image. This interface can contain one or more template images, which can serve as guide images. The electronic device can receive input for the template image and use it as the guide image. The "Custom Input Image" option 742 receives an image input by the user and uses it as the guide image. For example, the "Custom Input Image" option 742 can receive a user-drawn image input via a touchpad or input device (e.g., stylus, graphics tablet, etc.) and use it as the guide image. The "Local Input Image" option 743 triggers the electronic device to display an interface of an application storing local images, such as a gallery application interface or a file management application interface. Here, it can be used to trigger the electronic device to display... Figure 7B The interface 710 shown. The "Network Search Image" option 744 is used to trigger the electronic device to display a network search interface, which may include a search bar. The electronic device can receive keywords entered for the search bar, search for images corresponding to the keywords based on the keywords, receive input of images selected by the user, and use the selected image as a guide image.
[0138] In other examples, electronic devices can also... Figure 7C The interface shown on screen 720 displays a pop-up window with controls for adding guide parameters. This allows users to select the desired guide parameters based on the video's cover image, making the selection process convenient.
[0139] The electronic device can, in response to user input regarding the "local input image" option 743, display something like... Figure 7B The interface shown is 710. The electronic device receives user requests... Figure 7B After inputting the image option 712, the electronic device can display as follows: Figure 7F The interface shown is 750.
[0140] like Figure 7FAs shown, interface 750 may include image 751 corresponding to image option 712 and notification bar 752. Notification bar 752 may include, but is not limited to, prompt message 753, confirmation control 754, and cancellation control 755. Prompt message 753 can be used to prompt the user to confirm whether to use the image as the guide image. Confirmation control 754 is used to trigger the electronic device to filter the result image from the video corresponding to video option 711 based on image 751. Specifically, the description of how the electronic device determines the result image according to the guidance parameters can be found in [link to relevant documentation]. Figures 2 to 6 The illustrated embodiment will not be described in detail here. Optionally, after the electronic device obtains the result image through filtering, it can display the result image. The cancel control 755 is used to trigger the electronic device to cancel the display of the notification bar 752.
[0141] Optionally, after receiving user input for the "Local Input Image" option 743, the electronic device displays as follows: Figure 7B The interface shown is 710. The electronic device receives user requests... Figure 7B After image option 712 is input, image option 712 is marked with a label 713. This labeling indicates that the image corresponding to image option 712 has been selected as the guide image. The electronic device can also mark the image option selected by the user after receiving input from the user for other image options on interface 710, and use the image corresponding to the marked image option as the guide image. This facilitates the user in selecting multiple images as guide images on interface 710.
[0142] Electronic devices receive user requests Figure 7F After inputting the confirmation control 754 as shown, the electronic device can display the following: Figure 7G The notification bar 761 is shown below. Figure 7G As shown, the notification bar 761 may include, but is not limited to, the "Continue to input guide image" option 762, the "Input other guide parameters" option 763, and the "Output result image" option 764. The "Continue to input guide image" option 762 triggers the electronic device's display interface 710, allowing the electronic device to receive input for image options in the interface 710 and use the selected image as the guide image. This allows the user to select multiple images simultaneously as guide images. The "Input other guide parameters" option 763 triggers the electronic device to display, as shown in the image below. Figure 7D The interface shown is 730, or, as... Figure 7E The interface shown is 740.
[0143] Electronic devices can receive user requests Figure 7G After inputting the "Output Result Image" option 764 as shown, the following will be displayed: Figure 7H The interface shown is 760. (As shown) Figure 7HAs shown, interface 760 includes the result image. Specifically, the electronic device receives user requests for... Figure 7G After inputting the “Output Result Image” option 764, the image 751 corresponding to the image option 712 can be used as the guide image, and the video corresponding to the video option 711 can be used as the image set. The result image can be determined by the image processing method provided in the embodiment of this application.
[0144] During this process, the electronic device obtains the similarity curves between each video frame in the video corresponding to video option 711 and the guide image, which can be shown as follows: Figure 8 As shown in the figure. The horizontal axis of the curve represents the time of the video, and the vertical axis represents the similarity between each video frame and the guide image.
[0145] like Figure 8 As shown, the image subset 801 in the video corresponding to video option 711 is related to the guide image (i.e. Figure 7F Image 751 shows the highest similarity, while the image subset 802 in the video corresponding to video option 711 has the lowest similarity to the guide image. The electronic device can use image subset 801 as a candidate image set, then score the images in the candidate image set using an aesthetic evaluation algorithm, and use the image with the highest aesthetic score as the result image. Figure 7H The interface 760 shown displays the resulting image. For example, the image subset 801 includes images 801A, 801B, and 801C. The electronic device can score images 801A, 801B, and 801C using an aesthetic evaluation algorithm. Here, image 801A has the highest aesthetic score, and the electronic device can use image 801A as the resulting image and display it.
[0146] In some examples, if image 801A in the video corresponding to video option 711 has the highest similarity to the guide image, the electronic device can skip the aesthetic evaluation process and display image 801A on interface 760.
[0147] Optionally, the electronic device can use the resulting image as... Figure 7B The thumbnail of video option 711 in interface 710 is shown. In this way, the electronic device can filter the video to obtain the result image that meets the user's needs and has good aesthetics, and use the result image as the video thumbnail to give the user a better experience.
[0148] Optionally, the interface 760 may also include a sharing option for sharing the resulting image with other users. This allows the electronic device to filter images from a collection to obtain aesthetically pleasing results that meet the user's needs, and then share these results with other users, making image sharing more convenient.
[0149] Similarly, when an electronic device displays such as Figure 7D When the interface shown is 730, the electronic device can receive user requests... Figure 7D After inputting into the add control of option 731 as shown, the following will be displayed: Figure 7I The interface shown is 770.
[0150] For example, such as Figure 7I As shown, interface 770 may include one or more options. These options may be used to trigger the electronic device to display guide images from different sources. These options may include, but are not limited to, "Select Template Text" option 771, "Custom Input Text" option 772, "Local Input Text" option 773, and "Web Search Text" option 774.
[0151] The "Select Template Text" option 771 triggers the electronic device to display an interface containing template text. This interface can contain one or more template text sentences, which can serve as guiding text. The electronic device can receive input for the template text and use it as guiding text. The "Custom Input Text" option 772 triggers the electronic device to display a text box. The electronic device can receive text input for the text box and use it as guiding text. The "Local Input Text" option 773 triggers the electronic device to display an interface of an application storing local text, such as a notepad application interface. This notepad application interface can contain one or more text sentences, which can serve as guiding text. The electronic device can receive input of selected text from the user and use the selected text as guiding text. The "Web Search Text" option 774 triggers the electronic device to display a web search interface. This web search interface can include a search bar. The electronic device can receive keywords entered for the search bar, search for text corresponding to the keywords, receive input of selected text from the user, and use the selected text as guiding text. The electronic device can filter and obtain the result image from the video corresponding to video option 711 based on the user-added guidance text. Specifically, the description of how the electronic device determines the result image according to the guidance text can be found in the above embodiment, and will not be repeated here. It is understood that when the electronic device includes multiple elements such as user-added guidance text, guidance audio, and guidance image, the electronic device can filter and obtain the result image based on multiple guidance parameters.
[0152] Similarly, when an electronic device displays such as Figure 7D When the interface shown is 730, the electronic device can receive user requests... Figure 7D After inputting into the add control of option 733 as shown, the following will be displayed: Figure 7J The interface shown is 780.
[0153] For example, such as Figure 7J As shown, interface 780 may include one or more options. These options may be used to trigger the electronic device to display guide images from different sources. These options may include, but are not limited to, "Select Template Audio" option 781, "Custom Input Audio" option 782, "Local Input Audio" option 783, and "Web Search Audio" option 784.
[0154] The "Select Template Audio" option 781 triggers the electronic device to display an interface containing template audio. This interface can contain one or more template audio files, which can serve as guide audio. The electronic device can receive input for the template audio and use it as the guide audio. The "Custom Input Audio" option 782 prompts the user to input guide audio by recording audio. This option can trigger the electronic device to record audio, which can then be used as the guide audio. The "Local Input Audio" option 783 triggers the electronic device to display an interface for an application storing local audio, such as a recorder application. The "Network Search Audio" option 784 triggers the electronic device to display a network search interface, which may include a search bar. The electronic device can receive keywords entered in the search bar, search for audio corresponding to those keywords, receive user input of selected audio, and use the selected audio as the guide audio.
[0155] Electronic devices receive user requests Figure 7J After inputting the "Local Input Audio" option 783 as shown, the electronic device can display as follows: Figure 7K The interface shown is 790. (As shown) Figure 7K As shown, interface 790 may include an audio icon 791, an audio identifier 792, and a confirmation control 793. The confirmation control 793 is used to trigger the electronic device to use the audio file corresponding to the audio identifier 792 as the guiding audio. The electronic device can filter and obtain the result image from the video corresponding to video option 711 based on the user-added guiding audio. Specifically, the description of the electronic device determining the result image according to the guiding audio can be found in the above embodiment and will not be repeated here. It is understood that when the electronic device includes multiple elements such as user-added guiding audio, guiding audio, and guiding image, the electronic device can filter and obtain the result image based on multiple guiding parameters.
[0156] In some examples, if the electronic device receives a user's request... Figure 7G After entering option 763, which shows "Enter other boot parameters", the following will be displayed. Figure 7D The interface shown is 730. The electronic device can receive user requests... Figure 7D After entering option 733 as shown, the following will be displayed: Figure 7J The interface shown is 780. The electronic device can receive user requests... Figure 7J After inputting the "Local Input Audio" option 783 as shown, the electronic device can display as follows: Figure 7K The interface 790 is shown. When the electronic device receives input for the confirmation control 793, it can determine the audio corresponding to the audio identifier 792 as the guide audio. After determining the guide audio, the electronic device can, as shown... Figure 7D The interface 730 shown displays the option for the selected guide parameter. Optionally, the option for the selected guide parameter may include an identifier or thumbnail of the selected guide parameter, etc. In this way, the electronic device can prompt the user about the selected guide parameter.
[0157] Optionally, the electronic device can display a delete control among the selected guide parameters. This delete control can be used to trigger the electronic device to deselect the corresponding guide parameter. This allows users to delete unwanted guide parameters during the process of adding them.
[0158] It should be noted that, Figures 7A-7K The interface shown is merely an example. Electronic devices can implement the image processing method provided in this application embodiment through other controls, and this application embodiment does not limit this.
[0159] In this way, electronic devices can implement the image processing method provided in the embodiments of this application through the above interface.
[0160] In some application scenarios, electronic devices can use the image processing method provided in this application to mark specific actions of the subject being photographed. For example, the electronic device can receive user input and acquire a specified image of a specified subject through a camera. The electronic device can also receive user input and use the specified image as a guide image. During video recording, the electronic device can filter and obtain result images from the video stream. In this way, the electronic device uses the image of the subject as a reference to determine the result image corresponding to the subject and the guide image from the real-time recorded video stream. For example, the electronic device can receive user input to capture images of a sleeping baby, a crying baby, etc. The electronic device can receive user input and use the captured image as a guide image. Then, during the recording of a video with a baby as the subject, the electronic device can filter and obtain result images of the subject, including sleeping babies, crying babies, etc. In this way, the user can understand the specific state of the subject through the electronic device, and since the subject in the guide image is the same as the subject in the video, the electronic device can obtain a more accurate result image.
[0161] In some applications, electronic devices can capture specific actions of athletes while recording sports events, based on user-added guidance parameters. This allows the device to capture highlight moments of athletes according to the user-selected parameters. The device can also process large amounts of video data and filter the results based on user-selected guidance parameters. This allows the device to select desired footage from a large pool of video material, facilitating video editing. For example, the device can also capture images of road scenes based on user-selected guidance parameters, enabling users to record specific traffic conditions. Furthermore, the device can receive user-added guidance parameters describing specific actions of patients and, when recording videos including patients, filter the resulting images based on these parameters. This facilitates patient monitoring, and so on.
[0162] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. An image processing method, characterized in that, Applied to electronic devices, the method includes: The system receives the first video selected by the user and guidance parameters, which include one or more of guidance text, guidance audio, and guidance image. The video frames of the first video are split into multiple image subsets in chronological order, and the number of images in each image subset is less than or equal to a preset number of images. Identify the scene type to which the multiple image subsets belong; Identify a subset of target images that have the same scene type as the guidance parameters; Based on the guidance parameters, result images with a similarity greater than a preset similarity to the guidance parameters are determined from the subset of target images.
2. The method according to claim 1, characterized in that, The step of determining, based on the guidance parameters, result images from the subset of target images whose similarity to the guidance parameters is greater than a preset similarity specifically includes: Based on the guidance parameters, a candidate image set is determined from the target image subset, the candidate image set including multiple images in the target image subset whose similarity to the guidance parameters is greater than a preset similarity; An aesthetic score is determined for each image in the candidate image set. The aesthetic score is used to evaluate the image's composition, color, lighting, exposure, and theme. The result image with an aesthetic score greater than a preset aesthetic score is determined from the candidate image set.
3. The method according to claim 2, characterized in that, The step of determining a candidate image set from the target image subset based on the guidance parameters specifically includes: Based on the target image subset, a fragment vector corresponding to the target image subset is encoded, and the fragment vector includes the features of all images in the target image subset; The guidance parameters are encoded to obtain the guidance vector; Based on the fragment vector and the guiding vector, the similarity between the target image subset and the guiding parameters is determined; From the candidate image set, a set of candidate images with a similarity greater than a preset similarity is determined.
4. The method according to claim 3, characterized in that, The step of encoding the segment vector corresponding to the target image subset based on the target image subset specifically includes: The images of the target image subset are encoded to obtain image vectors; The average of the image vectors of all images in the image subset is calculated to obtain the segment vector.
5. The method according to claim 4, characterized in that, The guidance parameters include multiple items from the guidance image, guidance text, and guidance audio; the guidance vector includes multiple items from the guidance image vector, guidance text vector, and guidance audio vector; the guidance image corresponds to a first weight, the guidance text corresponds to a second weight, and the guidance audio corresponds to a third weight. The step of determining the similarity between the target image subset and the guiding parameters based on the fragment vector and the guiding vector specifically includes: Based on the similarity between the guiding image vector and the segment vector, the first weight, the similarity between the guiding text vector and the segment vector, the second weight, the similarity between the guiding audio vector and the segment vector, and the third weight, the similarity between the target image subset and the guiding parameters is determined.
6. The method according to claim 3, characterized in that, Identify a subset of images with the same scene type as the guidance parameters, specifically including: If the guidance parameters include a guidance image, based on the text description of the scene type, the scene to which the guidance image belongs is identified, and from the plurality of image subsets, the target image subset with the same scene type as the guidance image is determined; If the guidance parameters include guidance audio, based on the key information of the guidance audio, the scene to which the guidance audio belongs is determined, and from the multiple image subsets, the target image subset with the same scene type as the guidance audio is determined; If the guidance parameters include guidance text, based on the key information of the guidance text, the scene to which the guidance text belongs is determined, and from the plurality of image subsets, the target image subset with the same scene type as the guidance text is determined.
7. The method according to any one of claims 2-6, characterized in that, When there are multiple images with aesthetic scores greater than the preset aesthetic score, the resulting image has the highest aesthetic score.
8. The method according to claim 7, characterized in that, Determining the aesthetic score for each image in the candidate image set specifically includes: Based on the aesthetic evaluation model, the aesthetic scores of all images in the candidate image set are determined.
9. The method according to claim 1, characterized in that, When there are multiple images with a similarity greater than a preset similarity, the resulting image has the highest similarity.
10. The method according to claim 1, characterized in that, Before receiving the first video selected by the user and the guidance parameters, the method further includes: Display a first interface, which includes a gallery application icon; Receive the first input for the gallery application icon; In response to the first input, one or more video options are displayed, the one or more video options including the video options of the first video; Receiving the first video selected by the user specifically includes: Receive a second input for video options for the first video; In response to the second input, a second interface is displayed and the first video is selected. The second interface is used to play the first video.
11. The method according to claim 10, characterized in that, After displaying the second interface and selecting the first video, the method further includes: Received a third input for the first control on the second interface; In response to the third input, a second control, a third control, and a fourth control are displayed, wherein the second control is used to add the guide text, the third control is used to add the guide audio, and the fourth control is used to add the guide image.
12. An electronic device, characterized in that, include: A display screen, one or more processors, and one or more memories; The display screen, the one or more memories, and the one or more processors are coupled together. The one or more memories are used to store an executable program, which, when executed by the one or more processors, causes the electronic device to perform the method as described in any one of claims 1-11.
13. A readable storage medium for storing a program, characterized in that, When the program is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-11.