Method for generating emoji package and electronic equipment
By extracting and matching videos or pictures in electronic devices and generating multiple emoticons, the cumbersome problem of users needing to create emoticons in advance is solved, and the user experience is achieved conveniently generating emoticons that meet their needs is improved.
Patent Information
- Application Number
- CN202311806676.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, users need to make emoticons in advance, and the production steps are cumbersome and difficult to meet the needs of users to quickly generate emoticons that meet their needs.
By extracting and matching features of videos or pictures in electronic devices, multiple static and dynamic emoticons can be generated, and users can obtain multiple emoticons with just a simple operation.
It realizes the convenient generation of multiple emoticon packages that meet user needs by electronic devices, and improves the diversity and fun of user-created emoticon packages.
Smart Images

Figure CN120216720A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technologies, and more particularly, to a method for generating expression packs and an electronic device. Background Art
[0002] Currently, during the chatting process, users often have the need to use expression packs. Expression packs can vividly express some emotions or moods of users, are more vivid than text and voice, and can increase the fun of social interaction.
[0003] However, if users use self-made expression packs, they need to make expression packs in advance, and the current steps for making expression packs are relatively cumbersome. Summary of the Invention
[0004] This application provides a method for generating expression packs and an electronic device. This technical solution can enable the electronic device to conveniently generate expression packs that meet user needs.
[0005] In a first aspect, a method for generating expression packs is provided. The method is applied to an electronic device and includes: in response to an operation of triggering the generation of an expression pack, extracting features of a selected first target file to obtain first feature information, where the first target file is a video file or a picture; performing feature matching according to the first feature information to recall a plurality of target features; generating a plurality of expression packs corresponding to the first target file according to the plurality of target features, and the plurality of expression packs include static expression packs and / or dynamic expression packs.
[0006] Based on the embodiments of this application, when detecting an operation of triggering the generation of an expression pack, the electronic device can extract features of the first target file selected by the user, and perform feature matching according to the obtained first feature information to obtain a plurality of target features; and generate a plurality of corresponding expression packs according to the plurality of target features.
[0007] In this way, the electronic device can conveniently generate a plurality of expression packs according to simple operations of the user. In addition, the expression packs are generated from the videos or pictures selected by the user, thereby improving the diversity and fun of user-created expression packs.
[0008] In combination with the first aspect, in one implementation manner of the first aspect, the first target file is a video file, and the video file includes multiple frames of images. Among them, the extracting features of the selected first target file to obtain first feature information includes: splitting the video file into the multiple frames of images; respectively extracting features of each frame of the multiple frames of images to obtain the first feature information corresponding to each frame of the image.
[0009] Exemplarily, the first feature information may include expressions, facial features, backgrounds, actions, captions, and corresponding tag information, etc.
[0010] Based on the embodiments of the present application, if the first target file is a video file, when the electronic device performs feature extraction, it may first split the video file into multiple frames of images, and perform feature extraction on each frame of the multiple frames of images respectively, so as to obtain the feature information of each frame of image.
[0011] Combined with the first aspect, in one implementation manner of the first aspect, the generating a plurality of emoji corresponding to the first target file according to the plurality of target features includes: generating emoji corresponding to each frame of image according to the plurality of target features corresponding to each frame of image.
[0012] Based on the embodiments of the present application, the electronic device generates corresponding emoji according to the plurality of target features corresponding to each frame of image. In this way, the electronic device can generate multiple emoji from a video file, thereby improving the richness of the generated emoji and providing multiple emoji for users to choose from.
[0013] Combined with the first aspect, in one implementation manner of the first aspect, the generating emoji corresponding to each frame of image according to the plurality of target features corresponding to each frame of image includes: sorting the plurality of target features corresponding to each frame of image respectively; adding the caption corresponding to the target feature ranked first to each frame of image to generate the emoji corresponding to each frame of image.
[0014] Based on the embodiments of the present application, for a target image frame, the electronic device can sort the multiple target features that match the features, and select the caption corresponding to the target feature ranked first and add it to the target image frame, so as to generate the corresponding emoji.
[0015] In other examples, if the target feature ranked first does not correspond to a caption, the emoji corresponding to each frame of image can also be directly generated without adding a caption.
[0016] Combined with the first aspect, in one implementation manner of the first aspect, the method further includes: performing feature extraction on a preset number of consecutive images in the multiple frames of images to obtain the first feature information corresponding to the consecutive images.
[0017] Exemplarily, the preset number may be 3 or 5, etc.
[0018] Based on the embodiments of the present application, the electronic device can also perform feature extraction on multiple consecutive frames of images to obtain the corresponding first feature information.
[0019] For example, an electronic device can perform feature extraction on three consecutive frames of images, and then splice the features of each frame of image together to serve as the first feature information corresponding to the three consecutive frames of images.
[0020] In combination with the first aspect, in one implementation manner of the first aspect, the method further includes: generating a meme corresponding to the consecutive images according to multiple target features corresponding to the consecutive images, where the meme corresponding to the consecutive images is a dynamic meme.
[0021] Based on the embodiments of the present application, an electronic device can also generate dynamic memes, thereby improving the diversity of the generated memes.
[0022] In combination with the first aspect, in one implementation manner of the first aspect, the generating a meme corresponding to the consecutive images according to multiple target features corresponding to the consecutive images includes: sorting the multiple target features corresponding to the consecutive images respectively; adding the caption corresponding to the target feature ranked first to the consecutive images to generate the meme corresponding to the consecutive images.
[0023] Based on the embodiments of the present application, for consecutive image frames, an electronic device can sort multiple target features with feature matching and select the caption corresponding to the target feature ranked first to add to the consecutive image frames, thereby automatically generating memes with corresponding captions.
[0024] In combination with the first aspect, in one implementation manner of the first aspect, the performing feature matching according to the first feature information includes: performing feature matching from a cloud server according to the first feature information, where the cloud server includes feature information of online memes.
[0025] For example, the cloud server may include a feature index library, which includes feature information of online memes. It should be understood that through this feature information, the corresponding online memes can be determined.
[0026] Based on the embodiments of the present application, an electronic device can perform feature matching from the feature index library of the cloud server according to the first feature information. Since there is a large amount of feature information of online memes in the feature index library, the matching result can be made more accurate.
[0027] In combination with the first aspect, in one implementation manner of the first aspect, the method further includes: displaying the multiple memes in a first display interface, and the first display interface further includes function buttons for performing target operations on each of the multiple memes respectively.
[0028] Exemplarily, the target operation may be editing, liking, disliking, sharing, storing operation, etc.
[0029] Based on the embodiments of the present application, the electronic device can also display the generated emoticons on the display interface, so as to better display the generated emoticons to the user.
[0030] In combination with the first aspect, in one implementation of the first aspect, the method further includes: in response to an operation of the user clicking an edit function button of the dynamic emoticon, displaying a second display interface, where the second display interface includes a plurality of function buttons for editing the dynamic emoticon; in response to the user clicking a function button for editing the number of frames of the dynamic emoticon among the plurality of function buttons, displaying a third display interface, where the third display interface includes multiple frames of images included in the dynamic emoticon; and in response to an operation of the user clicking a function button for deleting the target image, deleting the target image.
[0031] Based on the embodiments of the present application, the user can also edit the generated emoticons, such as deleting the frames that the user does not like in the emoticon, so as to improve the operability of the user and provide the possibility for the user to create emoticons for the second time.
[0032] In a second aspect, a method for generating emoticons is provided. The method is applied to a cloud server and includes: obtaining network emoticons, where the network emoticons include static emoticons and dynamic emoticons; extracting feature information of each emoticon in the static emoticons to obtain second feature information of each emoticon, and extracting feature information of a plurality of consecutive frames of images included in each emoticon in the dynamic emoticons to obtain third feature information of the consecutive frames of images; and storing the second feature information and the third feature information.
[0033] Exemplarily, the cloud server can obtain network emoticons by means of web crawlers.
[0034] Based on the embodiments of the present application, the cloud server can obtain a large number of network emoticons, respectively extract the corresponding feature information from the network emoticons, and store the feature information for subsequent feature matching by the electronic device.
[0035] In a third aspect, an electronic device is provided, including: one or more processors; one or more memories; where the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the method for generating emoticons as described in the first aspect and any one of its possible implementation manners is executed.
[0036] In a fourth aspect, a device for generating emoticons is provided, including modules for implementing the method for generating emoticons as described in the first aspect and any one of its possible implementation manners.
[0037] In a fifth aspect, a cloud server is provided, including: one or more processors; one or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the method for generating an emoji as described in the second aspect is executed.
[0038] In a sixth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface is configured to receive a signal and transmit the signal to the processor. The processor processes the signal such that the method for generating an emoji as described in the first aspect and any possible implementation thereof is executed.
[0039] In a seventh aspect, a readable storage medium is provided. Instructions are stored in the readable storage medium, and when the instructions are run on an electronic device, the method for generating an emoji as described in the first aspect and any possible implementation thereof is executed.
[0040] In an eighth aspect, a program product is provided. The program product includes program code, and when the program code is run on an electronic device, the method for generating an emoji as described in the first aspect and any possible implementation thereof is executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic architecture diagram of an electronic device provided by an embodiment of the present application.
[0042] Figure 2 is a schematic diagram of a scenario applicable to an embodiment of the present application.
[0043] Figure 3 is a schematic diagram of a group of graphical user interfaces provided by an embodiment of the present application.
[0044] Figure 4 is a schematic diagram of another group of graphical user interfaces provided by an embodiment of the present application.
[0045] Figure 5 is a schematic flowchart of a method for generating an emoji provided by an embodiment of the present application.
[0046] Figure 6 is a schematic diagram of feature extraction provided by an embodiment of the present application.
[0047] Figure 7 is a schematic diagram of splitting a video into multiple frame images provided by an embodiment of the present application.
[0048] Figure 8 is a schematic diagram of feature extraction provided by an embodiment of the present application.
[0049] Figure 9It is a schematic diagram of generating emoticons by feature matching provided by an embodiment of the present application.
[0050] Figure 10 It is a schematic diagram of generating emoticons provided by an embodiment of the present application.
[0051] Figure 11 It is a schematic flowchart of another method for generating emoticons provided by an embodiment of the present application.
[0052] Figure 12 It is a schematic flowchart of a method for generating emoticons provided by an embodiment of the present application.
[0053] Figure 13 It is a schematic flowchart of another method for generating emoticons provided by an embodiment of the present application.
[0054] Figure 14 It is a schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0055] Next, the technical solutions in the present application will be described with reference to the accompanying drawings.
[0056] The method for generating emoticons in the embodiments of the present application can be applied to electronic devices such as smart phones, smart speakers, smart TVs, tablet computers, laptop computers, personal computers (PCs), ultra-mobile personal computers (UMPCs), netbooks, in-vehicle devices, wearable devices, foldable devices, and Internet of Things (IOT) devices.
[0057] Figure 1The structural schematic diagram of the electronic device 100 is shown. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0058] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0059] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0060] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.
[0061] A memory can also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0062] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus interface, etc.
[0063] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL).
[0064] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to enable communication between the processor 110 and the audio module 170.
[0065] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface.
[0066] The UART interface is a universal serial data bus for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160.
[0067] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193.
[0068] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc.
[0069] The USB interface 130 is an interface that conforms to the USB standard specification, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and peripheral devices.
[0070] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0071] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input of the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.
[0072] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110.
[0073] The wireless communication function of the electronic device 100 can be implemented through antenna 1, antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.
[0074] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100.
[0075] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.), or displays an image or video through the display screen 194. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.
[0076] The wireless communication module 160 may provide wireless communication solutions applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), Bluetooth low energy (BLE), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0077] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies.
[0078] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0079] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), or can be a display panel made of one of materials such as organic light-emitting diode (OLED), active-matrix organic light emitting diode (AMOLED), flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, or quantum dot light-emitting diodes (QLED). In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0080] The electronic device 100 can implement a shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0081] The ISP is used to process the data fed back by the camera 193. The camera 193 is used to capture still images or videos.
[0082] The digital signal processor is used to process digital signals. In addition to being able to process digital image signals, it can also process other digital signals.
[0083] The video codec is used to compress or decompress digital videos.
[0084] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 100.
[0085] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121.
[0086] The electronic device 100 can implement an audio function through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc.
[0087] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert analog audio input into digital audio signals.
[0088] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal.
[0089] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal.
[0090] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal.
[0091] The headphone jack 170D is used to connect a wired headphone.
[0092] Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, an acceleration sensor 180E, a distance sensor 180F, a fingerprint sensor 180H, a touch sensor 180K, a bone conduction sensor 180M, etc.
[0093] The keys 190 include a power-on key, volume keys, etc.
[0094] The motor 191 can generate a vibration prompt.
[0095] The indicator 192 can be an indicator light, which can be used to indicate the charging state, the change in battery level, and can also be used to indicate messages, missed calls, notifications, etc.
[0096] The SIM card interface 195 is used to connect a SIM card.
[0097] Currently, during the chatting process, users often have the need to use emoticons. Emoticons can vividly express some emotions or moods of users, are more vivid than text and voice, and can increase the fun of social interaction.
[0098] However, if users use the emoticons they made by themselves, they need to make the emoticons in advance, and the current steps of making emoticons are relatively cumbersome.
[0099] In view of this, the present application provides a method for generating emoticons and an electronic device. In this technical solution, the electronic device can conveniently generate multiple emoticons that meet the user's needs according to the user's operations.
[0100] Before introducing the embodiments of the present application, the following introduces some professional terms that the present application may involve.
[0101] Web crawler: Also known as a web spider or web robot. A web crawler can automatically collect and organize data information on the Internet on behalf of users. In the embodiments of the present application, the cloud server can obtain various web emoticons in the way of a web crawler.
[0102] Label: In machine learning, a label can be used as the target variable for training and evaluating a machine learning model.
[0103] Feature: In machine learning, a feature can be used as a symbol or marker of the characteristics of a thing.
[0104] Feature extraction: Extract the features of data through various methods.
[0105] Model training and inference: Before any machine learning model is used to solve a specific technical problem, it needs to be trained. The training of a model refers to the process of using a specified initial model to calculate the training data, and adjusting the parameters in the initial model by a certain method according to the calculation results, so that the model gradually learns certain rules and has specific functions. A model with stable functions after training can be used for inference. The inference of a model is the process of using the trained model to calculate the input data and obtain the predicted inference results.
[0106] In the training stage, first, a training set for a deep learning model needs to be constructed based on the target. The training set includes multiple training data, and each training data is set with a label. The label of the training data is the correct answer to the specific problem of the training data, and the label can represent the target of training the deep learning model using the training data.
[0107] When training the model, the training data can be input into the model after parameter initialization in batches. The model calculates the training data (i.e., "inference") to obtain the predicted results for the training data. The predicted results obtained through inference and the labels corresponding to the training data are used as the data for calculating the loss according to the loss function. The loss function is a function used to calculate the gap (i.e., "loss value") between the predicted results of the model for the training data and the labels of the training data during the model training stage. The loss function can be implemented using different mathematical functions. The expressions of common loss functions include: mean square error loss function, logarithmic loss function, least squares method, etc. The training of the model is a repeated iterative process. Each iteration infers different training data and calculates the loss value. The goal of multiple iterations is to continuously update the parameters and find the parameter configuration that minimizes or stabilizes the loss value of the loss function.
[0108] Figure 2 It is a schematic diagram of a scenario applicable to an embodiment of the present application. As Figure 2 shown, this scenario may include an electronic device and a cloud server.
[0109] The electronic device may at least include an application for generating emoticons, such as a photo album application. The electronic device may also include a key frame extraction service, a recall and sorting service, and an emoticon generation service.
[0110] Among them, the photo album application can include a meme production module and a user behavior feature collection module. The meme production module can provide the ability to produce memes, and provide the ability to display the interface and interact with the meme production interface. The user behavior feature collection module can collect data on a series of operation behaviors of the user on the generated memes, such as editing, sharing, liking, disliking, saving and other behavior data.
[0111] In other examples, the photo album application can also be replaced by an instant messaging application.
[0112] The key frame extraction service can provide the ability to perform online segmentation and feature extraction on the videos uploaded by users. Among them, the feature extraction can include feature extraction on a single image frame, and feature extraction on multiple consecutive image frames.
[0113] The recall and sorting service can provide the ability to match the features of the memes according to the features generated from the videos or pictures uploaded by users with the meme feature index library, perform feature recall, and sort the recalled features.
[0114] Exemplarily, the recall and sorting server can perform feature matching and feature recall from the meme feature storage module in the cloud server.
[0115] The meme generation service can provide the ability to convert the recalled and sorted features into meme data and generate the final memes and return them to the display interface of the electronic device.
[0116] The cloud server can at least include a meme feature index library service, a meme feature storage module, and a user behavior feature storage module.
[0117] Among them, the meme feature index library service can provide the ability to crawl popular meme pictures and tag materials through Internet crawler technology, generate a meme feature index library through algorithm training, extract captions and tags, and generate features associated with the memes.
[0118] Exemplarily, the meme feature index library service can store the extracted feature information of the network memes in the meme feature storage module.
[0119] The meme feature storage module can provide the ability to store the feature information of the memes generated by the algorithm after crawling by the meme feature index library service.
[0120] The user behavior feature storage module can provide the ability to store user behavior data.
[0121] The method for generating memes in the present application will be introduced below in combination with several graphical user interfaces (GUI).
[0122] Exemplarily, Figure 3 is a schematic diagram of a group of GUIs provided by an embodiment of the present application. Among them, from Figure 3 (a) to (h) in it show the process of an electronic device generating an emoji according to a user's operation.
[0123] Refer to Figure 3 (a) in it. This GUI is the display interface 210 of the electronic device 200, and this display interface 210 can be the display desktop of the electronic device. The display interface 210 may include a plurality of application programs installed in the electronic device 200.
[0124] After the electronic device 200 detects that the user clicks the operation of the gallery 211, it can display a GUI as shown in Figure 3 (b) in it.
[0125] Refer to Figure 3 (b) in it. This GUI can be the display interface 220 of the gallery of the electronic device 200. The display interface 220 may include a plurality of videos or pictures or animated pictures, etc. captured or saved by the electronic device, and the display interface 220 may also include a plurality of function buttons.
[0126] Exemplarily, after the electronic device 200 detects that the user long-presses the operation of the video 221, it can display a GUI as shown in Figure 3 (c) in it.
[0127] Refer to Figure 3 (c) in it. This GUI can be the display interface 230 of the electronic device 200. In this display interface 230, the video 221 is in a selected state, and selectable function boxes appear for other pictures or videos. The bottom of the display interface 230 may also include a sharing function button, a select-all function button, a creation function button, a deletion function button 231, and a more function button.
[0128] After the electronic device 200 detects that the user clicks the operation of the creation function button 231, it can display a GUI as shown in Figure 3 (d) in it.
[0129] Refer to Figure 3 (d) in it. A card 232 can be displayed above the creation function button in the display interface 230. The card 232 may include a movie option, a jigsaw option, and an emoji generation option.
[0130] When the electronic device 200 detects that the user clicks the emoji generation option in the card 232, it can display a GUI as shown in Figure 3 (e) in it.
[0131] Refer toFigure 3 In (e) of [description], the GUI is the display interface 240 of the emoticons generated for the electronic device 200. The display interface 240 may include the text information "Emoticon results generated by AI" and emoticons 241, 242, 243, and 244.
[0132] Among them, multiple function buttons may also be included below each emoticon, such as edit, like, dislike, share, or save, etc. For example, for the emoticon 244, an edit function button 2441, a like function button 2442, a dislike function button 2443, a share function button 2444, and a save function button 2445 may be included below the emoticon 244.
[0133] It can be understood that the number of emoticons generated in the display interface 240 is only illustrative. In other examples, the number of emoticons may also be 3, or 6, or 8, etc. Or, if the number of emoticons matched by the electronic device is small, such as the electronic device matches 2 emoticons, then the display interface 240 may include the 2 matched emoticons.
[0134] In this application, after the user selects the video 221 and the electronic device 200 detects the operation of the user clicking to generate an emoticon, the electronic device 200 may segment the image frames included in the video 221, extract features from each image frame, and obtain the feature information of each image frame; and perform feature matching in the feature index library in the cloud server according to the feature information of each image frame, obtain the matching result of each image frame, and generate the emoticon corresponding to each image frame according to the matching result.
[0135] For example, if the video 221 includes 5 frames of images, after the electronic device 200 detects the operation of the user clicking to generate an emoticon, the video 221 may be segmented into 5 frames of images.
[0136] It should be understood that the feature information of network emoticons may be stored in the feature index library. For example, the cloud server may obtain various network emoticons through web crawlers, extract features from the obtained network emoticons, and store them in the feature index library.
[0137] In some embodiments, the electronic device may also extract features from multiple consecutive image frames (such as three consecutive frames), obtain the feature information 2 of the multiple consecutive image frames, perform feature matching in the feature index library in the cloud server according to the feature information 2, and generate the emoticon corresponding to the multiple consecutive image frames according to the matching result. For example, the emoticon may be a graphics interchange format (GIF) image.
[0138] In some examples, for a certain image frame included in video 221, the electronic device may not match the result according to its feature information. At this time, the electronic device may not generate an emoji for this image frame.
[0139] Exemplarily, after the electronic device 200 detects the operation of the user clicking the edit function button 2441, it may display a GUI as shown in (f) of Figure 3 .
[0140] In other examples, after the user clicks the like function button 2442, the electronic device may upload the user's like behavior to the business intelligence (BI) big data platform, which is beneficial for the BI big data platform to perform model training. After the user clicks the dislike function button 2443, the electronic device may upload the user's behavior to the BI big data platform, which is beneficial for the BI big data platform to perform model training. When the user clicks the share 2443 button, the emoji 244 can be shared with other users in the pop-up page or card operation. The user can also save the emoji 244 to the electronic device or the cloud server by clicking the save function button 2445.
[0141] It should be understood that after the emoji 244 is generated, the user's operation behavior on the function button will be uploaded to the BI big data platform, which is beneficial for the BI big data platform to update or optimize the model, so that more emojis that meet the user's needs can be generated subsequently.
[0142] See Figure 3 . In (f), this GUI is the display interface 250 of the electronic device. The display interface 250 may include an emoji 244, a pause function button 251, a reverse play function button 252, an edit text function button 253, an edit size function button 254, and an edit frame number function button 255.
[0143] When the electronic device 200 detects the operation of the user clicking the pause function button 251, it can pause the dynamic playback of the emoji 244, which is convenient for the user to stay and watch the effect. When the electronic device 200 detects the operation of the user clicking the reverse play function button 252, it can play the emoji 244 in reverse order, thereby increasing the interest of the emoji. When the electronic device 200 detects the operation of the user clicking the edit text function button 253, it can edit the text in the emoji 244, which is beneficial for the user to modify the caption information of the emoji. When the electronic device 200 detects the operation of the user clicking the edit size function button 254, it can edit the size and clarity of the emoji 244 for the user to select emojis of different sizes and clarities, etc.
[0144] After the electronic device 200 detects the operation of the user clicking the edit frame number function button 255, it can display as shown in Figure 3 the GUI shown in (g) in
[0145] See Figure 3 (g) in
[0146] This GUI is the display interface 260 of the electronic device 200. The display interface 260 may include images 261, 262, 263, 264, and 265 included in the emoji package 244, and each image includes a delete function button. Figure 3 the GUI shown in (h) in
[0147] See Figure 3 (h) in
[0148] In this way, the user can delete some disliked image frames in the GIF image through convenient operations. Based on the embodiments of the present application, the electronic device can generate multiple static emojis and dynamic emojis for the user-selected video with one click. Thus, with simple operations by the user, the electronic device can conveniently generate multiple emojis, improving the user experience. The electronic device can also collect the user's operation behaviors on generating emojis, so that the big data platform can optimize the model for generating emojis, and thus the subsequent generated emojis can better meet the user's needs.
[0149] In addition, the user can also edit the emojis generated by the electronic device, thereby improving the diversity of the user's created emojis.
[0150] In some cases, when the electronic device generates emojis, the user can also provide the electronic device with corresponding tag information, so that the generated emojis can better meet the user's needs. The following will introduce this technical solution in combination with Figure 4 this
[0151] Exemplarily, Figure 4 is a schematic diagram of a group of GUIs provided by the embodiments of the present application. Among them, from Figure 4 (a) to (d) in
[0152] show the process of the electronic device generating emojis according to the user's operations. Figure 4 It should be understood that Figure 3 (a) to (b) in
[0153] Exemplarily, after the electronic device 200 detects that the user clicks the operation of generating an emoji in the card 232, it may display a GUI as shown in (c) of Figure 4 .
[0154] Referring to Figure 4 , the GUI may be the display interface 270 of the electronic device 200. The display interface 270 may include a video 221, a card 271, and a function button 273 for triggering the generation of an emoji.
[0155] Among them, the content of the card 271 can be used to provide tag information for generating an emoji by the electronic device 200. For example, the card 271 may include a prompt text "Fill in the tags you want, and the result will be more accurate", scene options: office, expression options: smile, action options: like, facial options: wearing glasses, caption options: beautiful.
[0156] It should be understood that each tag option may include multiple results for the user to choose. For example, for the scene option, the user can select other scenes by clicking the operation of the drop-down function button 282.
[0157] After the electronic device 200 detects that the user clicks the operation of the function button 273, it may display a GUI as shown in (d) of Figure 4 .
[0158] Referring to Figure 4 , the GUI may be the display interface 280 for the electronic device 200 to display the generated emoji.
[0159] Exemplarily, the display interface 280 may include emojis 281, 282, 283, and 284.
[0160] Among them, multiple function buttons may also be included below each emoji, such as edit, like, dislike, share, or save, etc.
[0161] In the embodiments of the present application, the user can select the tag information of the emoji that the electronic device wants to generate in the card 171 in the display interface 270. Thus, when the electronic device generates an emoji corresponding to the video 221, it can preferentially match the features under the above tag information in the feature index library, so that the generated emoji can better meet the user's needs. In addition, the electronic device provides the user with multiple selectable tag information, so as to improve the diversity and interest of generating emojis.
[0162] The following will introduce the technical solution for generating emojis in the present application in combination with Figures 5 - 11 .
[0163] Exemplarily, Figure 5It is a schematic flowchart of a method for generating emoticons provided by an embodiment of the present application. As Figure 5 shown, the method 300 may include steps 301 to 313.
[0164] 301, the cloud server obtains online emoticons from the Internet through a crawler.
[0165] Exemplarily, the cloud server may obtain online emoticons from emoticon websites, web pages, etc. according to the settings of developers. For example, the cloud server may obtain the emoticons released at the latest time, or the most shared emoticons by users, or all the emoticons on the website.
[0166] In this way, the cloud server can obtain a large amount of data of online emoticons, which is beneficial for subsequent electronic devices to match more suitable emoticons.
[0167] 302, the cloud server extracts features from the online emoticons to generate feature information of each online emoticon.
[0168] Exemplarily, the cloud server may extract features from each obtained online emoticon to generate feature information of each online emoticon.
[0169] Exemplarily, referring to Figure 6 , Figure 6 is a schematic diagram of feature extraction provided by an embodiment of the present application. As Figure 6 shown, the cloud server may extract features from the obtained online emoticons.
[0170] Among them, referring to Figure 6 in (a), for online emoticons in the form of pictures, that is, static emoticons, the cloud server may directly extract features. For example, for online emoticon 2, the cloud server may extract information such as the expression, action, facial features, background, caption, etc. in online emoticon 2, and generate a feature vector and label information through an algorithm. After that, the cloud server may store the feature vector, label information, and caption information in the feature index library.
[0171] It should be understood that for information such as expressions, actions, facial features, and backgrounds in pictures, corresponding feature vectors and label information may be generated. For example, for an expression, a feature vector A and label information A may be generated; for an action, a feature vector B and label information B may be generated.
[0172] In some embodiments, referring to Figure 6In (b) of , for GIF-form expression packs, that is, dynamic expression packs. For example, for online expression n, the cloud server can first perform image frame segmentation on the online expression n to obtain five frames of images; and perform feature extraction on each frame of image for information such as expression, action, facial features, and background, as well as feature extraction on text information, and generate feature vectors, label information, and caption information through algorithms. It should be understood that the feature vectors, label information, and caption information can be referred to as the feature information of the frame of image.
[0173] For GIF expression packs, in addition to performing feature extraction on each image frame, the cloud server can also perform feature extraction on multiple consecutive frames of images. For example, the online expression n includes 5 frames of images, namely frame 1, frame 2, frame 3, frame 4, and frame 5. The cloud server can perform feature extraction on consecutive frames 1, 2, and 3 to obtain their feature vectors, label information, and caption information, or can perform feature extraction on frames 2, 3, and 4. Alternatively, the cloud server can also perform feature extraction on multiple consecutive frames of images to generate multiple pieces of feature information. For example, the cloud server performs feature extraction on frames 1, 2, and 3 to obtain one piece of feature information; and performs feature extraction on frames 1, 2, 3, and 4 to obtain another piece of feature information, and can also perform feature extraction on frames 1, 2, 3, 4, and 5 to obtain another piece of feature information.
[0174] It should be understood that the present application does not limit the specific number of multiple consecutive frames of images.
[0175] It should also be understood that for all GIF expression packs obtained by the cloud server, the above method can be used for feature extraction.
[0176] 303. The cloud server stores the feature information in the feature index library.
[0177] Exemplarily, continue to refer to Figure 6 , the cloud server can store the feature information in the feature index library, so that subsequently the electronic device can perform feature matching in the feature index library according to the feature information of the selected picture or video to obtain the feature information of the matched expression pack.
[0178] 304. The cloud server stores the feature information in the BI big data platform.
[0179] It can be understood that the present application does not limit the specific order of steps 303-304.
[0180] 305. The BI big data platform performs model training and inference based on the feature information.
[0181] The BI big data platform can use the stored feature information for model training and inference, so as to improve the accuracy of the model.
[0182] It should be understood that the BI big data platform can also be a BI big data cluster.
[0183] 306. The electronic device divides the selected video into multiple frame images according to the user's operation.
[0184] Exemplarily, video 1 is stored in the electronic device. If the user hopes to generate emoticons using this video 1, the user can select video 1 and click the operation to generate emoticons. At this time, the electronic device can divide video 1 into multiple consecutive frame images. For example, see Figure 7 , Figure 7 is a schematic diagram of dividing a video into multiple frame images. If video 1 includes 5 frame images, the electronic device can divide video 1 into 5 consecutive frame images.
[0185] It should be understood that video 1 can be a video pre-stored in the electronic device or can be taken on-site by the electronic device, which is not limited in the embodiments of the present application.
[0186] 307. The electronic device extracts features from each of the multiple frame images and extracts features from a preset number of consecutive images among the multiple frame images to generate corresponding feature information.
[0187] Exemplarily, if video 1 includes 5 frame images, the electronic device can extract features from each of the 5 frame images, such as extracting information such as expressions, actions, facial features, and backgrounds in the images, and respectively generate corresponding feature information 1. Taking the preset number as 3 as an example, the electronic device can also extract features from 3 consecutive frame images to generate feature information 2.
[0188] See Figure 8 , Figure 8 is a schematic diagram of feature extraction provided by the embodiments of the present application. For each frame image, the electronic device can extract features such as expressions, actions, facial features, and backgrounds in the image to generate feature information 1.
[0189] It should be understood that the feature information 1 may include a feature vector A1 and a label A1 corresponding to the expression, a feature vector A2 and a label A2 corresponding to the action, a feature vector A3 and a label A3 corresponding to the facial feature, and a feature vector A4 and a label A4 corresponding to the background. Alternatively, the feature information 1 may further include a feature vector A5 and a label A5 corresponding to the caption information, etc.
[0190] The feature information 2 may include a feature vector B1 and a label B1 corresponding to 3 consecutive frame images. The feature B1 may include information such as expressions, actions, facial features, and backgrounds in the images, and may also include caption information.
[0191] It can be understood that the feature information may include the above-mentioned feature information 1 and feature information 2.
[0192] In some embodiments, the electronic device may also perform feature extraction only on each frame of the image, rather than on multiple frames of the image. Alternatively, when the electronic device performs feature extraction on multiple frames of the image, it may also extract features of multiple frames of images with different numbers. For example, the electronic device extracts features of 3 consecutive frames of images and extracts features of 5 consecutive frames of images.
[0193] In other examples, the image may not include a person, but may include an animal, a landscape, etc., which are not limited in the embodiments of the present application.
[0194] 308. The electronic device performs feature matching from the feature index library according to the feature information, and recalls multiple target features.
[0195] Exemplarily, for each frame of the image included in Video 1, the electronic device may perform feature matching from the feature index library according to its feature information 1, and recall multiple target features corresponding to this frame of the image.
[0196] The electronic device may perform matching according to the similarity of the vectors and the labels. For example, in the case of label matching, multiple (such as 3 or 5, etc.) target features with a similarity greater than a preset value are obtained.
[0197] For example, referring to Figure 9 , Figure 9 is a schematic diagram of generating an emoji through feature matching provided by the embodiments of the present application. For the feature information 1 corresponding to each frame of the image, the electronic device may perform feature matching from the feature index library and recall multiple target features according to the matching result. For example, the multiple target features are "feature 1, feature 2, feature 3...".
[0198] Exemplarily, for the feature information 2 corresponding to multiple consecutive frames of images, the electronic device may also perform feature matching from the feature index library according to its feature information 2, and recall multiple target features corresponding to this multiple frames of images. The electronic device may perform matching according to the similarity of the vectors and the labels. For example, in the case of label matching, multiple (such as 3 or 5, etc.) target features with a similarity greater than a preset value are obtained.
[0199] It should be understood that the emoji corresponding to the target features matched by the feature information 2 of multiple consecutive frames of images is a dynamic emoji in GIF format.
[0200] It should be understood that the embodiments of the present application do not limit the specific values of the multiple target features.
[0201] In some embodiments, if the user selects a picture instead of a video, the electronic device can directly extract information such as expressions, actions, facial features, background, etc. from the picture, and perform feature matching in the feature index library to recall multiple target features. That is, for this picture, the processing method of the electronic device is the same as that of each frame image included in the video.
[0202] 309, the electronic device sorts the multiple target features.
[0203] Exemplarily, referring further to Figure 9 , the electronic device can sort according to the similarity between the multiple target features and the feature information. For example, sort in descending order of similarity to obtain a sorting result of "Feature 1, Feature 2, Feature 3...".
[0204] Alternatively, the electronic device can also sort according to the feedback data of all users who use this function in big data statistics. For example, for similar features, most users prefer target feature 2, then target feature 2 can be ranked first to obtain a sorting result of "Feature 2, Feature 1, Feature 3...".
[0205] 310, the electronic device sends the sorted result to the BI big data platform.
[0206] The electronic device can send the sorted result to the BI big data platform, which is beneficial to the BI big data platform for model training and inference.
[0207] 311, the electronic device generates a meme according to the sorted result.
[0208] In some examples, for each frame image included in Video 1, the electronic device can generate a meme according to the feature ranked first.
[0209] For example, referring further to Figure 9 , for the first frame image, the feature ranked first is target feature 1, and its caption information is caption 1. Then the electronic device can add the caption 1 to the first frame image to generate meme 1. Refer to Figure 3 in (e), this meme 1 can be meme 241.
[0210] Exemplarily, referring to Figure 10 , Figure 10 is a schematic diagram of generating a meme provided by an embodiment of the present application. As shown in Figure 10 (a) in, for the first frame image, through feature matching, the feature ranked first that the electronic device matches is the feature corresponding to the online expression 2, and the caption information corresponding to the online expression 2 is caption 1. Then the electronic device can combine the first frame image and caption 1 to generate meme 1.
[0211] It can be understood that for the first frame image, the electronic device can also generate multiple emoji. For example, the electronic device can select the top several features and generate corresponding multiple emoji.
[0212] In other examples, the feature corresponding to the online emoji 2 may not have a caption. In this case, the electronic device can directly generate an emoji from the first frame image.
[0213] As Figure 10 shown in (b) of [], for the third frame image, through feature matching, the top-ranked feature matched by the electronic device is the feature corresponding to the online emoji 8, and the caption information corresponding to the online emoji 8 is caption 2. Then the electronic device can combine the third frame image and caption 2 to generate emoji 2.
[0214] As Figure 10 shown in (c) of [], for the fifth frame image, through feature matching, the top-ranked feature matched by the electronic device is the feature corresponding to the online emoji 3, and the caption information corresponding to the online emoji 3 is caption 3. Then the electronic device can combine the fifth frame image and caption 3 to generate emoji 3.
[0215] It should be understood that for other frame images included in the video, the electronic device can perform the above processing in the same way. However, in some cases, if the electronic device fails to match a suitable feature through feature matching, it may also not generate an emoji.
[0216] In some examples, for a series of consecutive frame images, the top-ranked feature is the target feature B, and its caption information is caption B. Then the electronic device can add the caption B to the series of consecutive frame images to generate the GIF emoji B. Refer to Figure 3 (e) of [], this emoji B can be emoji 244, and this caption B is caption 4.
[0217] It can be understood that for the series of consecutive frame images, the electronic device can also generate multiple emoji. For example, the electronic device can select the top several features and generate corresponding multiple GIF emoji.
[0218] In other examples, the target feature B may not have a caption. In this case, the electronic device can directly generate a GIF emoji from the series of frame images.
[0219] It should be understood that the embodiments of the present application do not limit the specific execution order of steps 310-311. In other examples, step 310 may not be executed either.
[0220] In some cases, the BI big data platform can also optimize the model for generating emojis according to the behavior data. Optionally, method 300 may further include steps 312-313.
[0221] 312. The electronic device sends the user's behavior data to the BI big data platform.
[0222] Exemplarily, after the electronic device generates an emoji, the electronic device can collect the user's operations on the emoji. Refer to Figure 3 In (e) of, the user can click on the edit function button, like function button, dislike function button, share function button, save function button, etc., and the electronic device can collect the user's operations.
[0223] It should be understood that for the user's operation of clicking the dislike function button, it is considered that the generated emoji is a negative sample and the user does not like the emoji.
[0224] 313. The BI big data platform optimizes the model according to the user's behavior data.
[0225] The BI big data platform can optimize the model for generating emojis according to the user's behavior data, so that the subsequent generated emojis more meet the user's needs.
[0226] Based on the embodiments of the present application, the cloud server can obtain a large amount of network emoji data from the Internet, extract its features, and store them in the feature index library, which is conducive to the subsequent electronic device matching more suitable emojis. The electronic device can extract the features of the images included in the video or a series of consecutive frames according to the user's operation, match them with the features in the feature index library, and generate static emojis and GIF emojis according to the recalled features, which can improve the diversity of the generated emojis.
[0227] Figure 11 It is a schematic flowchart of another method for generating emojis provided by the embodiments of the present application. As Figure 11 shown, the method 400 may include steps 401 to step 413.
[0228] 401. The cloud server obtains network emojis from the Internet through web crawling.
[0229] 402. The cloud server extracts the features of the network emojis and generates the feature information of each network emoji.
[0230] 403. The cloud server stores the feature information in the feature index library.
[0231] 404. The cloud server stores the feature information in the BI big data platform.
[0232] 405. The BI big data platform performs model training and inference according to the feature information.
[0233] 406. The electronic device divides the selected video into multiple frame images according to the user's operation.
[0234] 407. The electronic device extracts features from each frame image of the multiple frame images, and extracts features from a preset number of consecutive images among the multiple frame images to generate corresponding feature information.
[0235] It should be understood that for steps 401-407, reference can be made to the relevant descriptions of steps 301-307. For the sake of brevity, they will not be elaborated here.
[0236] 408. The electronic device performs feature matching in the feature index library according to the feature information and the label information provided by the user, and recalls multiple target features.
[0237] It should be understood that for the feature information, reference can be made to the relevant descriptions in the previous text.
[0238] Exemplarily, referring to Figure 4 (c) therein, the label information may be the label in card 271, such as scene, expression, action, face, coordination, etc.
[0239] Exemplarily, when performing feature matching, the electronic device can preferentially match the features of the emoji under the above labels.
[0240] In this way, the electronic device can perform feature matching in the feature index library according to the feature information and the label information provided by the user, so that the recalled target features can better meet the user's needs.
[0241] 409. The electronic device sorts the multiple target features.
[0242] 410. The electronic device sends the sorting result to the BI big data platform.
[0243] 411. The electronic device generates emojis according to the sorting result.
[0244] 412. The electronic device sends the user's behavior data to the BI big data platform.
[0245] 413. The BI big data platform optimizes the model according to the user's behavior data.
[0246] It should be understood that for steps 409-413, reference can be made to the relevant descriptions of steps 309-313 in the previous text. For the sake of brevity, they will not be elaborated here.
[0247] Based on the embodiments of the present application, the cloud server can obtain a large amount of online meme data from the Internet, extract its features, and store them in the feature index library, which is beneficial for subsequent electronic devices to match more suitable memes. The electronic device can extract the features of the images included in the video or a series of consecutive frames of images according to the user's operations and the tag information provided by the user, match the features with those in the feature index library, and generate static memes and GIF memes based on the recalled features, which can improve the diversity of the generated memes.
[0248] Figure 12 FIG. is a schematic flowchart of a method for generating memes provided by an embodiment of the present application. As Figure 12 shown, the method 500 can be applied to an electronic device, and the method 500 can include steps 510 to 530.
[0249] 510. In response to an operation of triggering meme generation, the electronic device extracts the features of the selected first target file to obtain first feature information, where the first target file includes a video file or a picture.
[0250] The first target file can be a video file, or the first target file can also be a picture.
[0251] Exemplarily, referring to Figure 3 (d) in, the first target file can be video 221. The operation of triggering meme generation can be an operation where the user clicks to generate a meme. After that, the electronic device can extract the features of the video 221 to obtain first feature information.
[0252] For example, the electronic device can extract the features of each frame of the image included in the video 221 to obtain first feature information corresponding to each frame. The first feature information can include one or more of expressions, facial features, backgrounds, actions, captions, and corresponding tag information, etc.
[0253] The electronic device can also extract the features of a series of consecutive frames of images included in the video 221 to obtain first feature information corresponding to the series of consecutive frames of images.
[0254] In some cases, the expression can also be the expression of a cartoon character or an animal, etc.
[0255] 520. The electronic device performs feature matching according to the first feature information and recalls multiple target features.
[0256] In some embodiments, the electronic device can perform feature matching from the cloud server according to the first feature information, where the cloud server includes the feature information of online memes.
[0257] For example, a cloud server can obtain a large amount of data of online emoticons through a web crawler, extract features of the online emoticons, and store the feature information of the extracted online emoticons in a feature index library. Thus, an electronic device can perform feature matching from the feature index library in the cloud server according to the first feature information and recall multiple target features.
[0258] 530, The electronic device generates multiple emoticons corresponding to the first target file according to the multiple target features, and the multiple emoticons include static emoticons and / or dynamic emoticons.
[0259] Exemplarily, if the first target file is a video file, the electronic device can generate one or more emoticons corresponding to each frame according to the multiple target features recalled for each frame of the image. The electronic device can also generate one or more emoticons corresponding to multiple consecutive frames of images according to the multiple target features recalled for the multiple consecutive frames of images, and the emoticon can be a dynamic emoticon, such as a GIF emoticon.
[0260] Exemplarily, if the first target file is a picture, the electronic device can generate one or more corresponding emoticons according to the multiple target features recalled for the picture.
[0261] Based on the embodiments of the present application, when detecting an operation that triggers the generation of an emoticon, the electronic device can extract features of the first target file selected by the user and perform feature matching according to the obtained first feature information to obtain multiple target features; and generate multiple corresponding emoticons according to the multiple target features.
[0262] In this way, the electronic device can conveniently generate multiple emoticons according to the simple operations of the user. In addition, the emoticons are generated from the videos or pictures selected by the user, thereby improving the diversity and interest of user-created emoticons.
[0263] In some embodiments, the first target file is a video file, and the video file includes multiple frames of images. Among them, when the electronic device extracts features of the selected first target file to obtain first feature information, it includes:
[0264] The electronic device divides the video file into multiple frames of images;
[0265] The electronic device respectively extracts features from each frame of the multiple frames of images to obtain first feature information corresponding to each frame of the image.
[0266] Exemplarily, the first feature information may include expressions, facial features, backgrounds, actions, captions, and corresponding tag information, etc.
[0267] Based on the embodiments of the present application, if the first target file is a video file, when the electronic device performs feature extraction, it can first split the video file into multiple frames of images, and perform feature extraction on each frame of the multiple frames of images respectively, so as to obtain the feature information of each frame of image.
[0268] In some embodiments, the electronic device generates multiple emoji corresponding to the first target file according to multiple target features, including:
[0269] The electronic device generates emoji corresponding to each frame of image respectively according to the multiple target features corresponding to each frame of image.
[0270] Based on the embodiments of the present application, the electronic device generates corresponding emoji respectively according to the multiple target features corresponding to each frame of image. In this way, the electronic device can generate multiple emoji from a video file, thereby improving the richness of the generated emoji and providing multiple emoji for users to choose from.
[0271] In some embodiments, generating emoji corresponding to each frame of image respectively according to the multiple target features corresponding to each frame of image includes:
[0272] Sort the multiple target features corresponding to each frame of image respectively;
[0273] Add the caption corresponding to the target feature ranked first to each frame of image to generate the emoji corresponding to each frame of image.
[0274] Based on the embodiments of the present application, for the target image frame, the electronic device can sort the multiple target features that match the features and select the caption corresponding to the target feature ranked first and add it to the target image frame, so as to generate the corresponding emoji.
[0275] In other examples, if the target feature ranked first does not correspond to a caption, the emoji corresponding to each frame of image can also be directly generated without adding a caption.
[0276] In some embodiments, the method 500 further includes:
[0277] The electronic device performs feature extraction on a preset number of consecutive images in the multiple frames of images to obtain the first feature information corresponding to the consecutive images.
[0278] Exemplarily, the preset number can be 3 or 5, etc.
[0279] Based on the embodiments of the present application, the electronic device can also perform feature extraction on multiple consecutive frames of images to obtain the corresponding first feature information.
[0280] For example, the electronic device may perform feature extraction on three consecutive frames of images, and then splice the features of each frame of image together to serve as the first feature information corresponding to the three consecutive frames of images.
[0281] In some embodiments, the method 500 further includes:
[0282] The electronic device generates an emoji corresponding to the consecutive images according to a plurality of target features corresponding to the consecutive images, wherein the emoji corresponding to the consecutive images is a dynamic emoji.
[0283] Exemplarily, referring to Figure 3 (e) in, the dynamic emoji may be Emoji 244.
[0284] Based on the embodiments of the present application, the electronic device can also generate a dynamic emoji, thereby improving the diversity of the generated emojis.
[0285] In some embodiments, generating an emoji corresponding to the consecutive images according to a plurality of target features corresponding to the consecutive images includes:
[0286] Sorting the plurality of target features corresponding to the consecutive images respectively;
[0287] Adding the caption corresponding to the target feature ranked first to the consecutive images to generate an emoji corresponding to the consecutive images.
[0288] Exemplarily, the electronic device may sort according to the similarity between a plurality of target features and the first feature information. For example, sorting in descending order of similarity to obtain a sorting result of "Feature 1, Feature 2, Feature 3...".
[0289] Alternatively, the electronic device may also sort according to the feedback data of all users who use this function statistically by big data. For example, for similar features, most users prefer target feature 2, then target feature 2 can be ranked first to obtain a sorting result of "Feature 2, Feature 1, Feature 3...".
[0290] Based on the embodiments of the present application, for consecutive image frames, the electronic device can sort a plurality of target features with feature matching, and select the caption corresponding to the target feature ranked first and add it to the consecutive image frame, thereby automatically generating an emoji with the corresponding caption.
[0291] In some embodiments, performing feature matching according to the first feature information includes:
[0292] The electronic device performs feature matching from the cloud server according to the first feature information, and the cloud server includes the feature information of network emojis.
[0293] For example, a cloud server may include a feature index library that includes feature information of online emoticons. It should be understood that the corresponding online emoticon can be determined according to this feature information.
[0294] Based on the embodiments of the present application, an electronic device can perform feature matching from the feature index library of the cloud server according to the first feature information. Since there is a large amount of feature information of online emoticons in the feature index library, the matching result can be made more accurate.
[0295] In some embodiments, the method 500 further includes:
[0296] Display a plurality of emoticons in a first display interface, and the first display interface further includes function buttons for performing target operations on each of the plurality of emoticons respectively.
[0297] Exemplarily, referring to Figure 3 (e) in, the first display interface may be the display interface 240, and the target operation may be editing, liking, disliking, sharing, storing operations, etc.
[0298] Exemplarily, the first display interface may be the display interface of a gallery application, or may also be the display interface of other instant messaging applications.
[0299] Based on the embodiments of the present application, the electronic device can also display the generated emoticon in the display interface, so as to better display the generated emoticon to the user.
[0300] In some embodiments, the method 500 further includes:
[0301] In response to the user's operation of clicking the edit function button of the dynamic emoticon, display a second display interface, and the second display interface includes a plurality of function buttons for editing the dynamic emoticon;
[0302] In response to the user's clicking the function button for editing the number of frames of the dynamic emoticon among the plurality of function buttons, display a third display interface, and the third display interface includes multiple frames of images included in the dynamic emoticon;
[0303] In response to the user's operation of clicking the function button for deleting the target image, delete the target image.
[0304] Exemplarily, referring to Figure 3In (e)-(h) thereof, the editing function button may be the function button 2441, the second display interface may be the display interface 250, and the multiple function buttons may be the pause function button 251, the reverse playback function button 252, the edit text function button 253, the edit size function button 254, the edit frame number function button 255, etc. The third display interface may be the display interface 260, the target image may be the image 264, and the function button for deleting the target image may be the delete function button 2641.
[0305] Based on the embodiments of the present application, the user can also edit the generated emoji, such as deleting the frames in the emoji that the user does not like, thereby improving the operability of the user and providing the possibility for the user to create emojis for secondary creation.
[0306] Figure 13 It is a schematic flowchart of a method for generating emojis provided by the embodiments of the present application. As Figure 13 shown, the method 600 can be applied to a cloud server, and the method 600 may include steps 610 to 630.
[0307] 610, the cloud server obtains network emojis, and the network emojis include static emojis and dynamic emojis.
[0308] Exemplarily, the cloud server can obtain network emojis from the Internet by means of web crawlers, and specific reference can be made to the relevant descriptions in the foregoing text.
[0309] 620, the cloud server extracts features from each emoji in the static emojis to obtain second feature information of each emoji, and extracts features from each of the continuous multi-frame images included in each emoji in the dynamic emojis to obtain third feature information of the continuous multi-frame images.
[0310] The cloud server can extract features from the obtained network emojis. For example, for static emojis, the cloud server can extract features therefrom to obtain corresponding second feature information. For dynamic emojis, the cloud server can extract features from each of the continuous multi-frame images included in each dynamic emoji to obtain corresponding third feature information.
[0311] Among them, the second feature information and the third feature information may include one or more of expressions, facial features, backgrounds, actions, captions, and corresponding tag information, etc.
[0312] Exemplarily, the second feature information and the third feature information can be used for feature matching by an electronic device to recall multiple matching target features.
[0313] 630, the cloud server stores the second feature information and the third feature information.
[0314] Exemplarily, the cloud server may store the second feature information and the third feature information in a feature index library for feature matching by the electronic device.
[0315] Based on the embodiments of the present application, the cloud server may obtain a large number of online emoticons, respectively extract features of the online emoticons to obtain corresponding feature information, and store the feature information for subsequent feature matching by the electronic device.
[0316] Figure 14 is a schematic block diagram of an electronic device provided by an embodiment of the present application. As Figure 14 shown, the electronic device 700 may include one or more processors 710; one or more memories 720; the one or more memories 720 store one or more instructions, and when the instructions are executed by the one or more processors 710, the method for generating emoticons described in any of the possible implementation manners in the foregoing is executed.
[0317] Exemplarily, the electronic device 700 may be the foregoing electronic device 100, electronic device 200, etc., or a cloud server, etc.
[0318] An embodiment of the present application further provides a device, including a processor and a communication interface. The communication interface is configured to receive a signal and transmit the signal to the processor, and the processor processes the signal so that the method for generating emoticons described in any of the possible implementation manners in the foregoing is executed.
[0319] The device may be a chip. For example, the chip may be a chip system or an independent chip, etc.
[0320] An embodiment of the present application further provides a readable storage medium, in which instructions are stored. When the instructions run on an electronic device, the electronic device executes the above-related method steps to implement the method for generating emoticons in the above embodiment.
[0321] An embodiment of the present application further provides a program product. When the program product runs on an electronic device, the electronic device executes the above-related steps to implement the method for generating emoticons in the above embodiment.
[0322] An embodiment of the present application further provides a device for generating emoticons, including modules for implementing the method for generating emoticons described in any of the foregoing embodiments.
[0323] In addition, an embodiment of the present application further provides a device, which may specifically be a chip, a component, or a module. The device may include a processor and a memory connected to each other. The memory is used to store instructions. When the device runs, the processor may execute the instructions stored in the memory so that the device executes the method for generating expression packs in the above method embodiments.
[0324] Among them, the device, readable storage medium, program product, or device provided in this embodiment is all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0325] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0326] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0327] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other may be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0328] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0329] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0330] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs that can store program codes.
[0331] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for generating emoticons, characterized in that, The method is applied to an electronic device, and the method includes: In response to an operation of generating an emoji, perform feature extraction on a selected first target file to obtain first feature information, where the first target file includes a video file or a picture; Perform feature matching according to the first feature information to recall multiple target features; Generate multiple emojis corresponding to the first target file according to the multiple target features, where the multiple emojis include static emojis and / or dynamic emojis.
2. The method according to claim 1, wherein The first target file is a video file, and the video file includes multiple frames of images. Among them, the performing feature extraction on the selected first target file to obtain first feature information includes: Split the video file into the multiple frames of images; Perform feature extraction on each frame of the multiple frames of images respectively to obtain the first feature information corresponding to each frame of image.
3. The method according to claim 2, wherein The generating multiple emojis corresponding to the first target file according to the multiple target features includes: Generate emojis corresponding to each frame of image respectively according to the multiple target features corresponding to each frame of image.
4. The method according to claim 3, characterized in that, The generating emojis corresponding to each frame of image respectively according to the multiple target features corresponding to each frame of image includes: Sort the multiple target features corresponding to each frame of image respectively; Add the caption corresponding to the target feature ranked first to each frame of image to generate an emoji corresponding to each frame of image.
5. The method according to any one of claims 2-4, characterized in that, The method further includes: Perform feature extraction on a preset number of consecutive images in the multiple frames of images to obtain the first feature information corresponding to the consecutive images.
6. The method according to claim 5, characterized in that, The method further includes: Generate an emoji corresponding to the consecutive images according to the multiple target features corresponding to the consecutive images, where the emoji corresponding to the consecutive images is a dynamic emoji.
7. The method according to claim 6, wherein The generating an emoji corresponding to the consecutive images according to the multiple target features corresponding to the consecutive images includes: Sort the multiple target features corresponding to the consecutive images respectively; Add the caption corresponding to the target feature ranked first to the consecutive images to generate an emoji corresponding to the consecutive images.
8. The method according to any one of claims 1-7, characterized in that, The performing feature matching according to the first feature information includes: Perform feature matching from a cloud server according to the first feature information, where the cloud server includes feature information of online emojis.
9. The method according to any one of claims 1-7, characterized in that, The performing feature matching according to the first feature information to recall multiple target features includes: Perform feature matching from a cloud server according to the first feature information and label information provided by a user, where the cloud server includes feature information of online emojis.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: Display the multiple emojis on a first display interface, and the first display interface further includes function buttons for performing target operations on each of the multiple emojis respectively.
11. The method according to claim 10, wherein The method further includes: In response to an operation of the user clicking an edit function button of the dynamic emoji, display a second display interface, and the second display interface includes multiple function buttons for editing the dynamic emoji; In response to the user clicking on the function button for editing the number of frames of the dynamic emoji among the multiple function buttons, a third display interface is displayed, and the third display interface includes multiple frames of images included in the dynamic emoji. In response to the user's operation of clicking on the function button for deleting the target image, the target image is deleted.
12. A method for generating an emoji, characterized in that, The method is applied to a cloud server, and the method includes: Obtaining network emojis, where the network emojis include static emojis and dynamic emojis; Performing feature extraction on each emoji in the static emojis to obtain second feature information of each emoji, and performing feature extraction on consecutive multiple frames of images included in each emoji in the dynamic emojis to obtain third feature information of the consecutive multiple frames of images; Storing the second feature information and the third feature information.
13. An electronic device, characterized in that, including: One or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the method for generating emojis according to any one of claims 1-11 is executed.
14. A cloud server, characterized in that, including: One or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the method for generating emojis according to claim 12 is executed.
15. A chip, characterized in that, The chip includes a processor and a communication interface. The communication interface is used to receive a signal and transmit the signal to the processor, and the processor processes the signal so that the method for generating emojis according to any one of claims 1-12 is executed.
16. A readable storage medium, characterized in that, Instructions are stored in the readable storage medium, and when the instructions are run on an electronic device, the method for generating emojis according to any one of claims 1-12 is executed.