Emoji generating method and electronic device
Through the feature extraction and matching technology of electronic devices and cloud servers, the problem of cumbersome emoticons for users is solved, and the user experience is achieved conveniently generating diverse emoticons and improving user experience.
Patent Information
- Application Number
- PCT/CN2024/140926
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the steps for making user-made emoticon packages are cumbersome, making it difficult to easily generate emoticon packages that meet their needs.
Feature extraction and matching of videos or images selected by users through electronic devices, recall multiple target features, generate static and dynamic emoticon packages, and use the cloud server's feature index library to match and sort features, providing a variety of emoticon package generation methods.
It realizes convenient generation of multiple emoticon packages, which improves the diversity and fun of user-created emoticon packages, simplifies the production process, and enhances the user experience.
Smart Images

Figure CN2024140926_03072025_PF_FP_ABST
Abstract
Description
Method and electronic device for generating emoticon package
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 25, 2023, with application number 202311806676.5 and application name “Method and electronic device for generating emoticon packages”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of electronic technology, and more specifically, to a method for generating an emoticon package and an electronic device. Background Art
[0003] Currently, users often need to use emoticons during chats. Emoticons can vividly express some of the user's emotions or moods, are more vivid than text and voice, and can increase the fun of social interaction.
[0004] However, if users use their own emoticon packages, they need to prepare the emoticon packages in advance, and the current steps for preparing the emoticon packages are relatively complicated. Summary of the Invention
[0005] The present application provides a method and electronic device for generating an emoticon package. The technical solution enables the electronic device to conveniently generate an emoticon package that meets user needs.
[0006] In a first aspect, a method for generating an emoticon package is provided, which is applied to an electronic device, and the method includes: in response to an operation that triggers the generation of an emoticon package, performing feature extraction on a selected first target file to obtain first feature information, wherein the first target file is a video file or a picture; performing feature matching based on the first feature information to recall multiple target features; and generating multiple emoticon packages corresponding to the first target file based on the multiple target features, wherein the multiple emoticon packages include static emoticon packages and / or dynamic emoticon packages.
[0007] Based on the embodiment of the present application, when an operation that triggers the generation of an emoticon package is detected, the electronic device can extract features of the first target file selected by the user, and perform feature matching based on the obtained first feature information to obtain multiple target features; and generate corresponding multiple emoticon packages based on the multiple target features.
[0008] In this way, the electronic device can generate multiple emoticon packages conveniently according to the user's simple operation. In addition, the emoticon package is generated by the video or picture selected by the user, thereby improving the diversity and fun of the user's creation of the emoticon package.
[0009] In combination with the first aspect, in an implementation of the first aspect, the first target file is a video file, and the video file includes multiple frames of images, wherein the feature extraction of the selected first target file to obtain the first feature information includes: dividing the video file into the multiple frames of images; and performing feature extraction on each frame of the multiple frames to obtain the first feature information corresponding to each frame of the images.
[0010] Exemplarily, the first feature information may include expressions, facial features, background, actions, accompanying text and corresponding tag information, etc.
[0011] Based on the embodiment of the present application, if the first target file is a video file, the electronic device can first divide the video file into multiple frames of images when performing feature extraction, and perform feature extraction on each frame of the multiple frames respectively, so as to obtain feature information of each frame of the image.
[0012] In combination with the first aspect, in an implementation method of the first aspect, generating multiple emoticon packages corresponding to the first target file based on the multiple target features includes: generating emoticon packages corresponding to each frame image according to the multiple target features corresponding to each frame image.
[0013] Based on the embodiments of the present application, the electronic device generates corresponding emoticon packages based on multiple target features corresponding to each frame of the image. In this way, the electronic device can generate multiple emoticon packages based on a video file, thereby increasing the richness of the generated emoticon packages and providing users with a variety of emoticon packages for users to choose from.
[0014] In combination with the first aspect, in an implementation method of the first aspect, generating an emoticon package corresponding to each frame image according to the multiple target features corresponding to each frame image includes: sorting the multiple target features corresponding to each frame image; adding the text corresponding to the target feature ranked first to each frame image, and generating an emoticon package corresponding to each frame image.
[0015] Based on the embodiment of the present application, for the target image frame, the electronic device can sort the multiple target features of feature matching, and select the text corresponding to the target feature ranked first and add it to the target image frame, so as to generate a corresponding emoticon package.
[0016] In other examples, if the target feature ranked first does not correspond to a text, an emoticon package corresponding to each frame of the image can be directly generated without adding a text.
[0017] In combination with the first aspect, in an implementation manner of the first aspect, the method further includes: performing feature extraction on a preset number of continuous images in the multiple frames of images to obtain the first feature information corresponding to the continuous images.
[0018] For example, the preset number may be 3 or 5, etc.
[0019] Based on the embodiment of the present application, the electronic device can also perform feature extraction on multiple consecutive frames of images to obtain corresponding first feature information.
[0020] For example, the electronic device may extract features from three consecutive image frames and then stitch together the features of each image frame to serve as the first feature information corresponding to the three consecutive image frames.
[0021] In combination with the first aspect, in an implementation of the first aspect, the method further includes: generating an emoticon package corresponding to the continuous images based on multiple target features corresponding to the continuous images, wherein the emoticon package corresponding to the continuous images is a dynamic emoticon package.
[0022] Based on the embodiments of the present application, the electronic device can also generate dynamic emoticon packages, thereby improving the diversity of the generated emoticon packages.
[0023] In combination with the first aspect, in an implementation method of the first aspect, generating an emoticon package corresponding to the continuous image based on multiple target features corresponding to the continuous image includes: sorting the multiple target features corresponding to the continuous image respectively; adding the text corresponding to the target feature ranked first to the continuous image, and generating an emoticon package corresponding to the continuous image.
[0024] Based on the embodiment of the present application, for continuous image frames, the electronic device can sort multiple target features of feature matching, and select the text corresponding to the target feature ranked first and add it to the continuous image frame, so as to automatically generate an emoticon package with corresponding text.
[0025] In combination with the first aspect, in an implementation of the first aspect, the performing feature matching based on the first feature information includes: performing feature matching from a cloud server based on the first feature information, the cloud server including feature information of a network emoticon package.
[0026] For example, the cloud server may include a feature index library, which includes feature information of network emoticon packages. It should be understood that the corresponding network emoticon package can be determined based on the feature information.
[0027] Based on the embodiment of the present application, the electronic device can perform feature matching from the feature index library of the cloud server according to the first feature information. Since the feature index library contains a large amount of feature information of online emoticon packages, the matching results can be made more accurate.
[0028] In combination with the first aspect, in an implementation of the first aspect, the method further includes: displaying the multiple emoticon packages in a first display interface, and the first display interface also includes function buttons for performing target operations on each of the multiple emoticon packages.
[0029] Exemplarily, the target operation may be editing, liking, disliking, sharing, saving, etc.
[0030] Based on the embodiment of the present application, the electronic device can also display the generated emoticon package in the display interface, so as to better show the generated emoticon package to the user.
[0031] In combination with the first aspect, in an implementation of the first aspect, the method further includes: in response to the user clicking the editing function button of the dynamic emoticon package, displaying a second display interface, the second display interface including multiple function buttons for editing the dynamic emoticon package; in response to the user clicking the function button for editing the number of frames of the dynamic emoticon package among the multiple function buttons, displaying a third display interface, the third display interface including multiple frame images included in the dynamic emoticon package; in response to the user clicking the function button for deleting the target image, deleting the target image.
[0032] Based on the embodiments of the present application, users can also edit the generated emoticon package, such as deleting frames that the user does not like in the emoticon package, thereby improving the user's operability and providing the possibility for users to re-create emoticon packages.
[0033] In a second aspect, a method for generating an emoticon package is provided, which is applied to a cloud server and includes: obtaining an online emoticon package, which includes a static emoticon package and a dynamic emoticon package; performing feature extraction on each emoticon package in the static emoticon package to obtain second feature information of each emoticon package, and performing feature extraction on continuous multi-frame images included in each emoticon package in the dynamic emoticon package to obtain third feature information of the continuous multi-frame images; and storing the second feature information and the third feature information.
[0034] For example, the cloud server may obtain the online emoticon package by means of a web crawler.
[0035] Based on the embodiment of the present application, the cloud server can obtain a large number of online emoticon packages, extract features from the online emoticon packages respectively, obtain corresponding feature information, and store the feature information to facilitate subsequent feature matching by electronic devices.
[0036] In a third aspect, an electronic device is provided, comprising: one or more processors; one or more memories; the one or more memories storing one or more programs, wherein when the one or more programs are executed by one or more processors, the method for generating an emoticon package as described in the first aspect and any possible implementation thereof is executed.
[0037] In a fourth aspect, a device for generating an emoticon package is provided, comprising a module for implementing the method for generating an emoticon package as described in the first aspect and any possible implementation thereof.
[0038] In a fifth aspect, a cloud server is provided, comprising: one or more processors; one or more memories; the one or more memories storing one or more programs, and when the one or more programs are executed by one or more processors, the method for generating an emoticon package as described in the second aspect is executed.
[0039] In the sixth aspect, a chip is provided, comprising a processor and a communication interface, wherein the communication interface is used to receive a signal and transmit the signal to the processor, and the processor processes the signal so that the method for generating an emoticon package as described in the first aspect and any possible implementation thereof is executed.
[0040] In a seventh aspect, a readable storage medium is provided, wherein instructions are stored in the readable storage medium. When the instructions are run on an electronic device, the method for generating an emoticon package as described in the first aspect and any possible implementation thereof is executed.
[0041] In an eighth aspect, a program product is provided, comprising a program code. When the program code is run on an electronic device, the method for generating an emoticon package as described in the first aspect and any possible implementation thereof is executed. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] FIG1 is a schematic architecture diagram of an electronic device provided in an embodiment of the present application.
[0043] FIG2 is a schematic diagram of a scenario to which the embodiments of the present application may be applied.
[0044] FIG3 is a schematic diagram of a set of graphical user interfaces provided in an embodiment of the present application.
[0045] FIG4 is a schematic diagram of another set of graphical user interfaces provided in an embodiment of the present application.
[0046] FIG5 is a schematic flowchart of a method for generating an emoticon package provided in an embodiment of the present application.
[0047] FIG6 is a schematic diagram of a feature extraction method provided in an embodiment of the present application.
[0048] FIG7 is a schematic diagram of dividing a video into multiple frames of images provided in an embodiment of the present application.
[0049] FIG8 is a schematic diagram of a feature extraction method provided in an embodiment of the present application.
[0050] FIG9 is a schematic diagram of generating an emoticon package through feature matching provided in an embodiment of the present application.
[0051] FIG10 is a schematic diagram of generating an emoticon package provided in an embodiment of the present application.
[0052] FIG11 is a schematic flowchart of another method for generating an emoticon package provided in an embodiment of the present application.
[0053] FIG12 is a schematic flowchart of a method for generating an emoticon package provided in an embodiment of the present application.
[0054] FIG13 is a schematic flowchart of another method for generating an emoticon package provided in an embodiment of the present application.
[0055] FIG14 is a schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0056] The technical solution in this application will be described below with reference to the accompanying drawings.
[0057] The method for generating emoticons in the embodiments of the present application can be applied to electronic devices such as smart phones, smart speakers, smart TVs, tablet computers, laptops, personal computers (PCs), ultra-mobile personal computers (UMPCs), netbooks, vehicle-mounted devices, wearable devices, foldable devices, and Internet of Things (IOT) devices.
[0058] 1 shows a schematic structural diagram of an electronic device 100. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0059] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0060] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0061] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0062] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0063] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus interface.
[0064] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL).
[0065] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170.
[0066] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via a PCM bus interface.
[0067] The UART interface is a universal serial data bus used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is generally used to connect the processor 110 and the wireless communication module 160.
[0068] The MIPI interface can be used to connect the processor 110 with peripheral devices such as the display screen 194 and the camera 193.
[0069] The GPIO interface can be configured via software. The GPIO interface can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc.
[0070] The USB interface 130 is an interface that complies with USB standards, and may be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 may be used to connect a charger to charge the electronic device 100, and may also be used to transfer data between the electronic device 100 and peripheral devices.
[0071] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0072] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device via the power management module 141.
[0073] The power management module 141 is used to connect the battery 142 , the charging management module 140 and the processor 110 .
[0074] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0075] The mobile communication module 150 can provide wireless communication solutions including 2G / 3G / 4G / 5G applied on the electronic device 100.
[0076] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0077] The wireless communication module 160 can provide wireless communication solutions for application on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), Bluetooth low energy (BLE), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.
[0078] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150 , and antenna 2 is coupled to wireless communication module 160 , so that electronic device 100 can communicate with the network and other devices through wireless communication technology.
[0079] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0080] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), or a display panel made of one of the following materials: organic light-emitting diode (OLED), active-matrix organic light-emitting diode (AMOLED), flexible light-emitting diode (FLED), MiniLED, MicroLed, Micro-oLed, or quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1.
[0081] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0082] The ISP is used to process data fed back by the camera 193. The camera 193 is used to capture still images or videos.
[0083] Digital signal processors are used to process digital signals. In addition to processing digital image signals, they can also process other digital signals.
[0084] Video codecs are used to compress or decompress digital video.
[0085] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100.
[0086] The internal memory 121 may be used to store computer executable program codes, which include instructions. The processor 110 executes the instructions stored in the internal memory 121 to execute various functional applications and data processing of the electronic device 100.
[0087] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0088] The audio module 170 is used to convert digital audio information into analog audio signals for output, and is also used to convert analog audio input into digital audio signals.
[0089] The speaker 170A, also called a "horn", is used to convert audio electrical signals into sound signals.
[0090] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals.
[0091] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals.
[0092] The headphone jack 170D is used to connect a wired headphone.
[0093] Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, an acceleration sensor 180E, a distance sensor 180F, a fingerprint sensor 180H, a touch sensor 180K, a bone conduction sensor 180M, etc.
[0094] The buttons 190 include a power button, a volume button, and the like.
[0095] Motor 191 can generate vibration prompts.
[0096] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.
[0097] The SIM card interface 195 is used to connect a SIM card.
[0098] Currently, users often need to use emoticons during chats. Emoticons can vividly express some of the user's emotions or moods, are more vivid than text and voice, and can increase the fun of social interaction.
[0099] However, if users use their own emoticon packages, they need to prepare the emoticon packages in advance, and the current steps for preparing the emoticon packages are relatively complicated.
[0100] In view of this, the present application provides a method and electronic device for generating emoticon packages. In this technical solution, the electronic device can conveniently generate multiple emoticon packages that meet user needs based on user operations.
[0101] Before introducing the embodiments of the present application, the following is an introduction to some professional terms that may be involved in the present application.
[0102] Crawler: Also known as a web crawler or web robot. A crawler can automatically collect and organize data and information on the Internet on behalf of a user. In the embodiment of the present application, the cloud server can obtain various online emoticon packages through a crawler.
[0103] Labels: In machine learning, labels can be used to train and evaluate target variables of machine learning models.
[0104] Features: In machine learning, features can be used as symbols, signs, etc. of the characteristics of things.
[0105] Feature extraction: Extract the features of data through various methods.
[0106] Model training and inference: Any machine learning model must be trained before it can be used to solve specific technical problems. Model training involves using a specified initial model to perform calculations on training data. Based on the calculation results, the parameters of the initial model are adjusted using a specific method, allowing the model to gradually learn certain patterns and acquire specific functions. Once trained and stable, the model can be used for inference. Model inference involves using the trained model to perform calculations on input data to obtain predicted inference results.
[0107] During the training phase, it is first necessary to build a training set for the deep learning model based on the goal. The training set includes multiple training data, each of which is set with a label. The label of the training data is the correct answer to the specific question of the training data. The label can represent the goal of training the deep learning model using the training data.
[0108] When training a model, training data can be input in batches into the model after parameter initialization. The model performs calculations (i.e., "inference") on the training data to obtain prediction results for the training data. The prediction results obtained through inference and the labels corresponding to the training data are used as data for calculating the loss according to the loss function. The loss function is a function used to calculate the gap (i.e., "loss value") between the model's prediction results for the training data and the labels of the training data during the model training phase. The loss function can be implemented using different mathematical functions. Commonly used loss function expressions include: mean square error loss function, logarithmic loss function, least squares method, etc. Model training is an iterative process. Each iteration infers different training data and calculates the loss value. The goal of multiple iterations is to continuously update the parameters and find the parameter configuration that minimizes or stabilizes the loss value of the loss function.
[0109] Figure 2 is a schematic diagram of a scenario to which the embodiments of the present application may be applied. As shown in Figure 2 , the scenario may include an electronic device and a cloud server.
[0110] The electronic device may include at least an application for generating an emoticon package, such as a photo album application, and may also include a key frame extraction service, a recall sorting service, and an emoticon package generation service.
[0111] The photo album application can include an emoticon creation module and a user behavior feature collection module. The emoticon creation module provides emoticon creation capabilities, along with user interface display and interaction capabilities. The user behavior feature collection module can collect data on a series of user actions on the generated emoticons, such as editing, sharing, liking, disliking, and saving.
[0112] In other examples, the photo album application may also be replaced by an instant messaging application.
[0113] The keyframe extraction service can provide online segmentation and feature extraction capabilities for user-uploaded videos. Feature extraction can include feature extraction for a single image frame or for multiple consecutive frames.
[0114] The recall and sorting service can provide a feature index library for matching emoticons generated based on features in videos or pictures uploaded by users, perform feature recall, and sort the recalled features.
[0115] Exemplarily, the recall sorting server may perform feature matching and feature recall from the emoticon feature storage module in the cloud server.
[0116] The emoticon package generation service can provide the ability to convert the recalled and sorted features into emoticon package data and generate the final emoticon package and return it to the display interface of the electronic device.
[0117] The cloud server may include at least an emoticon package feature index library service, an emoticon package feature storage module, and a user behavior feature storage module.
[0118] Among them, the emoticon feature index library service can provide the ability to crawl popular emoticon package images and label materials through Internet crawler technology, generate an emoticon package feature index library through algorithm training, extract accompanying text and labels, and generate features associated with emoticons.
[0119] Exemplarily, the emoticon package feature index library service may store the extracted feature information of the online emoticon package in the emoticon package feature storage module.
[0120] The emoticon package feature storage module can provide an emoticon package feature index library service to store the feature information of the emoticon package generated by the post-crawler algorithm.
[0121] The user behavior feature storage module can provide storage capabilities for user behavior data.
[0122] The following will introduce the method of generating emoticon packages of this application in combination with several sets of graphical user interfaces (GUIs).
[0123] For example, Figure 3 is a schematic diagram of a set of GUIs provided in an embodiment of the present application, wherein (a) to (h) in Figure 3 illustrate a process in which an electronic device generates an emoticon package based on a user's operation.
[0124] 3 ( a ), the GUI is a display interface 210 of the electronic device 200 , which may be a display desktop of the electronic device and may include multiple applications installed in the electronic device 200 .
[0125] When the electronic device 200 detects that the user clicks on the gallery 211 , it may display a GUI as shown in FIG. 3( b ).
[0126] 3( b ), the GUI may be a display interface 220 of a gallery of the electronic device 200. The display interface 220 may include multiple videos, pictures, or dynamic pictures captured or saved by the electronic device, and may also include multiple function buttons.
[0127] For example, after the electronic device 200 detects that the user has long-pressed the video 221 , it may display a GUI as shown in (c) in FIG. 3 .
[0128] 3 (c), the GUI may be a display interface 230 of the electronic device 200. In the display interface 230, the video 221 is selected, and selectable function boxes appear for other pictures or videos. The bottom of the display interface 230 may also include a share function button, a select all function button, a create function button, a delete function button 231, and more function buttons.
[0129] When the electronic device 200 detects that the user clicks the creation function button 231, it may display a GUI as shown in (d) in FIG. 3 .
[0130] 3 (d), a card 232 may be displayed above the creation function button in the display interface 230. The card 232 may include a movie option, a puzzle option, and an emoticon package generation option.
[0131] When the electronic device 200 detects that the user clicks the "Generate Emoticon Package" option in the card 232, the GUI shown in (e) of FIG. 3 may be displayed.
[0132] 3 (e), the GUI is a display interface 240 of the emoticon package generated by the electronic device 200. The display interface 240 may include text information "AI generated emoticon package results for you" and emoticon package 241, emoticon package 242, emoticon package 243 and emoticon package 244.
[0133] Each emoticon package may further include multiple function buttons below, such as edit, like, dislike, share, or save, etc. For example, for emoticon package 244, the emoticon package 244 may include an edit function button 2441, a like function button 2442, a dislike function button 2443, a share function button 2444, and a save function button 2445 below.
[0134] It is understood that the number of emoticon packages generated in the display interface 240 is merely illustrative. In other examples, the number of emoticon packages may also be 3, 6, or 8, etc. Alternatively, if the electronic device matches a small number of emoticon packages, such as 2 emoticon packages, the display interface 240 may include the 2 matched emoticon packages.
[0135] In the present application, after the user selects video 221, when the electronic device 200 detects that the user clicks on the operation of generating an emoticon package, the electronic device 200 can segment the image frames included in the video 221, and perform feature extraction on each image frame to obtain feature information of each image frame; and perform feature matching in the feature index library in the cloud server based on the feature information of each image frame to obtain the matching results for each image frame, and generate emoticon packages corresponding to each image frame based on the matching results.
[0136] For example, the video 221 includes 5 frames of images. After the electronic device 200 detects that the user clicks on the operation of generating an emoticon package, the video 221 can be divided into 5 frames of images.
[0137] It should be understood that the feature index library may store feature information of network emoticon packages. For example, the cloud server may obtain various network emoticon packages through a web crawler, extract features from the obtained network emoticon packages, and store them in the feature index library.
[0138] In some embodiments, the electronic device can also perform feature extraction on multiple continuous image frames (such as three consecutive frames) to obtain feature information 2 of the multiple continuous image frames, and perform feature matching in the feature index library in the cloud server based on the feature information 2, and generate emoticon packages corresponding to the multiple continuous image frames based on the matching results. For example, the emoticon package can be a graphics interchange format (GIF) image.
[0139] In some examples, the electronic device may not find a matching result based on the feature information of a certain image frame included in the video 221. In this case, the electronic device may not generate an emoticon package for the image frame.
[0140] For example, after the electronic device 200 detects that the user clicks the edit function button 2441 , it may display a GUI as shown in (f) of FIG. 3 .
[0141] In other examples, after the user clicks the like function button 2442, the electronic device can upload the user's like behavior to the business intelligence (BI) big data platform, which is conducive to the BI big data platform for model training. After the user clicks the step on function button 2443, the electronic device can upload the user's behavior to the BI big data platform, which is conducive to the BI big data platform for model training. When the user clicks the share 2443 button, the emoticon package 244 can be shared with other users on the pop-up page or card operation. The user can also save the emoticon package 244 to an electronic device or cloud server by clicking the save function button 2445.
[0142] It should be understood that after generating the emoticon package 244, the user's operation of the function button will be uploaded to the BI big data platform, which will help the BI big data platform to update or optimize the model, so that an emoticon package that better meets user needs can be generated in the future.
[0143] 3( f ), the GUI is a display interface 250 of the electronic device. The display interface 250 may include an emoticon package 244, a pause function button 251, a rewind function button 252, an edit text function button 253, an edit size function button 254, and an edit frame number function button 255.
[0144] When the electronic device 200 detects that the user clicks the pause function button 251, the dynamic playback of the emoticon package 244 can be paused, making it easier for the user to stop and watch the effect. When the electronic device 200 detects that the user clicks the reverse function button 252, the emoticon package 244 can be played in reverse order, thereby increasing the fun of the emoticon package. When the electronic device 200 detects that the user clicks the edit text function button 253, the text in the emoticon package 244 can be edited, which is beneficial for the user to modify the text information of the emoticon package. When the electronic device 200 detects that the user clicks the edit size function button 254, the size and clarity of the emoticon package 244 can be edited, allowing the user to select emoticon packages of different sizes and clarity, etc.
[0145] When the electronic device 200 detects that the user clicks the edit frame number function button 255 , it may display a GUI as shown in FIG. 3( g ).
[0146] 3( g ), the GUI is a display interface 260 of the electronic device 200. The display interface 260 may include images 261, 262, 263, 264, and 265 included in the emoticon package 244, and each image includes a delete function button.
[0147] For example, after the electronic device 200 detects that the user clicks the delete function button 2641 in the image 264 , it may display a GUI as shown in (h) in FIG. 3 .
[0148] 3(h), the display interface 260 includes images 261, 262, 263, and 265. In this way, the user can delete some undesirable image frames in the GIF image through a convenient operation.
[0149] Based on the embodiments of the present application, the electronic device can generate multiple static and dynamic emoticon packages for a user-selected video with one click. This allows the user to conveniently generate multiple emoticon packages with a simple operation, thereby improving the user experience. The electronic device can also collect the user's operational behavior in generating emoticon packages, enabling the big data platform to optimize the emoticon package generation model, thereby ensuring that subsequently generated emoticon packages better meet user needs.
[0150] In addition, users can also edit the emoticon packages generated by electronic devices, thereby increasing the diversity of user-created emoticon packages.
[0151] In some cases, when an electronic device generates an emoticon package, the user can also provide the electronic device with corresponding tag information, so that the emoticon package generated by the electronic device can better meet the user's needs. The following will introduce this technical solution in conjunction with Figure 4.
[0152] For example, Figure 4 is a schematic diagram of a set of GUIs provided in an embodiment of the present application, wherein (a) to (d) in Figure 4 illustrate a process in which an electronic device generates an emoticon package based on a user's operation.
[0153] It should be understood that (a) to (b) in FIG. 4 can refer to the relevant description of (a) to (b) in FIG. 3 in the previous text.
[0154] For example, after the electronic device 200 detects that the user clicks on the operation of generating an emoticon package in the card 232 , the GUI shown in (c) of FIG. 4 may be displayed.
[0155] 4( c ), the GUI may be a display interface 270 of the electronic device 200. The display interface 270 may include a video 221, a card 271, and a function button 273 for triggering the generation of an emoticon package.
[0156] The content of card 271 can be used to provide tag information for generating an emoticon package for electronic device 200. For example, card 271 can include a prompt text "Fill in the tag you want, the result will be more accurate", scene option: office, expression option: smile, action option: like, face option: wearing glasses, and text option: beautiful.
[0157] It should be understood that each label option can include multiple results for the user to select. For example, for the scene option, the user can select other scenes by clicking the drop-down function button 282.
[0158] When the electronic device 200 detects that the user clicks the function button 273 , it may display a GUI as shown in FIG. 4( d ).
[0159] Referring to (d) in FIG. 4 , the GUI may be a display interface 280 for displaying the generated emoticon package on the electronic device 200 .
[0160] Exemplarily, the display interface 280 may include an emoticon package 281 , an emoticon package 282 , an emoticon package 283 , and an emoticon package 284 .
[0161] Among them, each emoticon package can also include multiple function buttons under it, such as edit, like, dislike, share or save, etc.
[0162] In the embodiment of the present application, the user can select the tag information of the emoticon package that they want the electronic device to generate in the card 171 in the display interface 270. Therefore, when the electronic device generates the emoticon package corresponding to the video 221, it can prioritize matching the features under the above tag information in the feature index library, so that the generated emoticon package can better meet the user's needs. In addition, the electronic device provides the user with multiple selectable tag information, thereby increasing the diversity and fun of the generated emoticon package.
[0163] The following will introduce the technical solution for generating emoticon packages in this application in conjunction with Figures 5-11.
[0164] For example, Figure 5 is a schematic flow chart of a method for generating an emoticon package provided in an embodiment of the present application. As shown in Figure 5 , the method 300 may include steps 301 to 313 .
[0165] 301, the cloud server obtains the network emoticon package from the Internet through crawlers.
[0166] For example, the cloud server can obtain network emoticon packages from emoticon package websites, web pages, etc. according to the developer's settings. For example, the cloud server can obtain the emoticon packages released at the latest time, or the emoticon packages shared most by users, or all emoticon packages on the website.
[0167] In this way, the cloud server can obtain a large amount of online emoticon package data, which is conducive to subsequent electronic devices matching more suitable emoticon packages.
[0168] 302 , the cloud server extracts features from the online emoticon package and generates feature information of each online emoticon package.
[0169] Exemplarily, the cloud server may perform feature extraction on each acquired network emoticon package to generate feature information of each network emoticon package.
[0170] For example, referring to Figure 6, which is a schematic diagram of a feature extraction method provided by an embodiment of the present application, as shown in Figure 6, the cloud server can perform feature extraction on the acquired network emoticon package.
[0171] For example, as shown in Figure 6 (a), for image-based emoticons (i.e., static emoticons), the cloud server can directly perform feature extraction. For example, for emoticon 2, the cloud server can extract information such as the expression, action, facial features, background, and accompanying text, and generate a feature vector and label information through an algorithm. The cloud server can then store this feature vector, label information, and accompanying text information in a feature index library.
[0172] It should be understood that corresponding feature vectors and label information can be generated for expressions, actions, facial features, background, and other information in the image. For example, for expressions, feature vector A and label information A can be generated; for actions, feature vector B and label information B can be generated.
[0173] In some embodiments, referring to FIG6(b), for a GIF-formatted emoji package, i.e., a dynamic emoji package, for example, for an online emoji n, the cloud server may first segment the emoji n into five frames; then, for each frame, perform feature extraction of information such as expression, action, facial features, background, and text information, and algorithmically generate a feature vector, label information, and accompanying text information. It should be understood that the feature vector, label information, and accompanying text information may be referred to as the feature information of the frame image.
[0174] For GIF emoticon packages, in addition to extracting features from each image frame, the cloud server can also extract features from consecutive multi-frame images. For example, an online emoticon n includes five frames of images, namely frame 1, frame 2, frame 3, frame 4, and frame 5. The cloud server can extract features from consecutive frames 1, frame 2, and frame 3 to obtain their feature vectors, label information, and accompanying text information, and can also extract features from frames 2, frame 3, and frame 4. Alternatively, the cloud server can also extract features from multiple consecutive multi-frame images to generate multiple feature information. For example, the cloud server extracts features from frames 1, frame 2, and frame 3 to obtain one feature information; and extracts features from frames 1, frame 2, frame 3, and frame 4 to obtain another feature information; and can also extract features from frames 1, frame 2, frame 3, frame 4, and frame 5 to obtain another feature information.
[0175] It should be understood that the present application does not limit the specific number of consecutive multi-frame images.
[0176] It should also be understood that the above method can be used to extract features for all GIF emoticon packages obtained by the cloud server.
[0177] 303. The cloud server stores the feature information in a feature index library.
[0178] For example, referring to FIG6 , the cloud server may store the feature information in a feature index library, so that subsequent electronic devices may perform feature matching in the feature index library based on the feature information of the picture or video selected by the user to obtain the feature information of the matched emoticon package.
[0179] 304. The cloud server stores the feature information in the BI big data platform.
[0180] It is understandable that the present application does not limit the specific order of steps 303-304.
[0181] 305, the BI big data platform performs model training and inference based on feature information.
[0182] The BI big data platform can use stored feature information for model training and inference, thereby improving the accuracy of the model.
[0183] It should be understood that the BI big data platform can also be called a BI big data cluster.
[0184] 306 , the electronic device divides the selected video into multiple frames of images according to the user's operation.
[0185] For example, if a video 1 is stored in an electronic device and a user wishes to generate an emoticon package using the video 1, the user can select the video 1 and click on the operation of generating an emoticon package. At this time, the electronic device can segment the video 1 into multiple consecutive frames. For example, referring to FIG7 , FIG7 is a schematic diagram of segmenting a video into multiple frames. If the video 1 includes five frames, the electronic device can segment the video 1 into five consecutive frames.
[0186] It should be understood that the video 1 may be a video pre-stored in the electronic device, or may be a video shot on-site by the electronic device, which is not limited in the embodiment of the present application.
[0187] 307 , the electronic device performs feature extraction on each frame of the multiple frames of images, and performs feature extraction on a preset number of consecutive images in the multiple frames of images, to generate corresponding feature information.
[0188] Exemplarily, video 1 includes 5 frames of images, and the electronic device can perform feature extraction on each of the 5 frames, such as extracting expressions, actions, facial features, background and other information in the image, and generate corresponding feature information 1 respectively; taking the preset number of 3 as an example, the electronic device can also perform feature extraction on 3 consecutive frames of images to generate feature information 2.
[0189] 8 is a schematic diagram of a feature extraction method provided by an embodiment of the present application. For each frame of image, the electronic device can extract features of information such as expressions, actions, facial features, and background in the image to generate feature information 1.
[0190] It should be understood that the feature information 1 may include a feature vector A1 and label A1 corresponding to the expression, a feature vector A2 and label A2 corresponding to the action, a feature vector A3 and label A3 corresponding to the facial features, and a feature vector A4 and label A4 corresponding to the background. Alternatively, the feature information 1 may also include a feature vector A5 and label A5 corresponding to the accompanying text information.
[0191] The feature information 2 may include feature vectors B1 and labels B1 corresponding to three consecutive frames of images. The feature B1 may include information such as expressions, actions, facial features, background, etc. in the image, and may also include text information.
[0192] It can be understood that the characteristic information may include the above-mentioned characteristic information 1 and characteristic information 2.
[0193] In some embodiments, the electronic device may perform feature extraction on only each frame of an image, rather than on multiple frames of images. Alternatively, when performing feature extraction on multiple frames of images, the electronic device may extract features from multiple frames of images of different numbers, for example, the electronic device may extract features from three consecutive frames of images and then extract features from five consecutive frames of images.
[0194] In other examples, the image may not include people, but may include animals or scenery, etc., which is not limited in the embodiments of the present application.
[0195] 308 , the electronic device performs feature matching from a feature index library according to the feature information, and recalls multiple target features.
[0196] Exemplarily, for each frame image included in the video 1, the electronic device may perform feature matching from a feature index library based on its feature information 1, and recall multiple target features corresponding to the frame image.
[0197] The electronic device can match the vectors with the tags based on their similarity. For example, in the case of tag matching, multiple (eg, 3 or 5) target features with similarities greater than a preset value are obtained.
[0198] For example, see Figure 9, which is a schematic diagram of an embodiment of the present application, illustrating an example of generating an emoticon package through feature matching. For feature information 1 corresponding to each frame of image, the electronic device can perform feature matching from a feature index library and, based on the matching results, recall multiple target features. For example, the multiple target features may be "feature 1, feature 2, feature 3..."
[0199] For example, for feature information 2 corresponding to multiple consecutive frames of images, the electronic device can also perform feature matching from the feature index library based on the feature information 2, and recall multiple target features corresponding to the multiple frames of images. The electronic device can also perform matching based on vector similarity and labels. For example, in the case of label matching, multiple (e.g., 3 or 5, etc.) target features with a similarity greater than a preset value are obtained.
[0200] It should be understood that the emoticon package corresponding to the target feature matched by the feature information 2 corresponding to the continuous multiple frames of images is a dynamic emoticon package in GIF format.
[0201] It should be understood that the embodiments of the present application do not limit the specific values of the multiple target features.
[0202] In some embodiments, if the user selects an image instead of a video, the electronic device can directly extract information such as expressions, actions, facial features, and background information from the image, perform feature matching within a feature index library, and retrieve multiple target features. In other words, the electronic device processes the image in the same manner as it would each frame of a video.
[0203] At 309 , the electronic device sorts the multiple target features.
[0204] For example, referring to FIG9 , the electronic device may sort multiple target features according to their similarity to the feature information, for example, sorting them in descending order of similarity, and obtaining a sorting result of “feature 1, feature 2, feature 3…”.
[0205] Alternatively, the electronic device can also sort the feedback data of all users who use the function based on big data statistics. For example, for similar features, most users prefer target feature 2, so target feature 2 can be ranked first, and the sorting result is "feature 2, feature 1, feature 3..."
[0206] 310. The electronic device sends the sorting result to the BI big data platform.
[0207] Electronic devices can send the sorting results to the BI big data platform, which is beneficial for the BI big data platform to perform model training and reasoning.
[0208] 311. The electronic device generates an emoticon package according to the sorting result.
[0209] In some examples, for each frame of image included in video 1, the electronic device can generate an emoticon package based on the top-ranked feature.
[0210] For example, referring to FIG9 , for the first frame image, the top feature is target feature 1, and its accompanying text information is accompanying text 1. The electronic device can add accompanying text 1 to the first frame image to generate emoticon package 1. Referring to FIG3 (e), emoticon package 1 can be emoticon package 241.
[0211] For example, see Figure 10, which is a schematic diagram of an emoticon package generation method provided by an embodiment of the present application. As shown in Figure 10 (a), for the first frame of image, the electronic device performs feature matching. If the first matched feature is the feature corresponding to network emoticon 2, and the accompanying text information corresponding to network emoticon 2 is accompanying text 1, the electronic device can combine the first frame of image and accompanying text 1 to generate emoticon package 1.
[0212] It is understandable that the electronic device may also generate multiple emoticon packages for the first frame image. For example, the electronic device may take the features that rank in the first few positions and generate corresponding multiple emoticon packages.
[0213] In other examples, the feature corresponding to the online emoticon 2 may not have a caption. In this case, the electronic device may directly generate an emoticon package from the first frame image.
[0214] As shown in (b) in Figure 10, for the third frame image, the electronic device performs feature matching, and the first matched feature is the feature corresponding to the network emoticon 8, and the accompanying text information corresponding to the network emoticon 8 is accompanying text 2. The electronic device can then combine the third frame image and accompanying text 2 to generate emoticon package 2.
[0215] As shown in (c) in Figure 10, for the fifth frame image, the electronic device performs feature matching, and the first matched feature is the feature corresponding to the network emoticon 3, and the accompanying text information corresponding to the network emoticon 3 is accompanying text 3. The electronic device can then combine the fifth frame image and accompanying text 3 to generate emoticon package 3.
[0216] It should be understood that the electronic device can also perform the above processing on other frame images included in the video. However, in some cases, if the electronic device does not match the appropriate features through feature matching, it may not generate an emoticon package.
[0217] In some examples, for a plurality of consecutive image frames, if the top-ranked feature is target feature B and its accompanying text is accompanying text B, the electronic device may add accompanying text B to the plurality of image frames to generate a GIF emoticon package B. Referring to (e) in FIG3 , the emoticon package B may be emoticon package 244, and the accompanying text B may be accompanying text 4.
[0218] It is understandable that for the continuous multi-frame image, the electronic device can also generate multiple emoticon packages. For example, the electronic device can take the features that rank in the top few positions and generate corresponding multiple GIF emoticon packages.
[0219] In other examples, the target feature B may not have accompanying text. In this case, the electronic device can directly generate a GIF emoticon package from the multiple frames of images.
[0220] It should be understood that the embodiment of the present application does not limit the specific execution order of steps 310-311. In other examples, step 310 may not be performed.
[0221] In some cases, the BI big data platform can also optimize the model for generating emoticons based on the behavioral data. Optionally, the method 300 can also include steps 312-313.
[0222] 312. The electronic device sends the user's behavior data to the BI big data platform.
[0223] For example, after the electronic device generates an emoticon package, the electronic device can collect the user's operations on the emoticon package. Referring to (e) in Figure 3, the user can click the edit function button, the like function button, the dislike function button, the share function button, and the save function button, and the electronic device can collect the user's operations.
[0224] It should be understood that when a user clicks the step-on function button, the generated emoticon package is considered a negative sample, and the user does not like the emoticon package.
[0225] 313, BI big data platform optimizes the model based on user behavior data.
[0226] The BI big data platform can optimize the model for generating emoticon packages based on user behavior data, so that the subsequently generated emoticon packages are more in line with user needs.
[0227] Based on the embodiments of the present application, a cloud server can obtain a large amount of online emoticon package data from the internet, extract features from it, and store it in a feature index library, thereby facilitating subsequent electronic devices to match more suitable emoticon packages. Based on user operations, electronic devices can extract features from images or consecutive multi-frame images included in a video and match them with features in the feature index library. Based on the retrieved features, static emoticon packages and GIF emoticon packages can be generated, thereby increasing the diversity of the generated emoticon packages.
[0228] FIG11 is a schematic flow chart of another method for generating an emoticon package provided in an embodiment of the present application. As shown in FIG11 , the method 400 may include steps 401 to 413 .
[0229] 401, the cloud server obtains the emoticon package from the Internet through a crawler.
[0230] 402 , the cloud server extracts features from the online emoticon package and generates feature information for each online emoticon package.
[0231] 403. The cloud server stores the feature information in the feature index library.
[0232] 404. The cloud server stores the feature information in the BI big data platform.
[0233] 405, the BI big data platform performs model training and inference based on feature information.
[0234] 406 , the electronic device divides the selected video into multiple frames of images according to the user's operation.
[0235] 407 , the electronic device performs feature extraction on each frame of the multiple frames of images, and performs feature extraction on a preset number of consecutive images in the multiple frames of images, to generate corresponding feature information.
[0236] It should be understood that steps 401-407 can refer to the relevant description of steps 301-307, and for the sake of brevity, they are not repeated here.
[0237] 408 , the electronic device performs feature matching from a feature index library based on the feature information and the tag information provided by the user, and recalls multiple target features.
[0238] It should be understood that the feature information can be found in the relevant description above.
[0239] For example, referring to (c) in FIG. 4 , the tag information may be a tag in the card 271 , such as scene, expression, action, face, coordination, and the like.
[0240] Exemplarily, when performing feature matching, the electronic device may give priority to matching the features of the emoticon packages under the above tags.
[0241] In this way, the electronic device can perform feature matching in a feature index library based on the feature information and the tag information provided by the user, so that the recalled target features can better meet the needs of the user.
[0242] At 409 , the electronic device sorts the multiple target features.
[0243] 410. The electronic device sends the sorting result to the BI big data platform.
[0244] 411. The electronic device generates an emoticon package according to the sorting result.
[0245] 412. The electronic device sends the user's behavior data to the BI big data platform.
[0246] 413, the BI big data platform optimizes the model based on user behavior data.
[0247] It should be understood that steps 409-413 can refer to the relevant description of steps 309-313 in the previous text, and for the sake of brevity, they are not repeated here.
[0248] Based on the embodiments of the present application, a cloud server can obtain a large amount of online emoticon package data from the Internet, extract features from it, and store it in a feature index library, thereby facilitating subsequent electronic devices to match more suitable emoticon packages. The electronic device can extract features from images or continuous multi-frame images included in the video based on user operations and user-provided tag information, and match them with features in the feature index library. Based on the recalled features, static emoticon packages and GIF emoticon packages can be generated, which can improve the diversity of the generated emoticon packages.
[0249] FIG12 is a schematic flow chart of a method for generating an emoticon package provided by an embodiment of the present application. As shown in FIG12 , the method 500 can be applied to an electronic device, and the method 500 can include steps 510 to 530.
[0250] 510. In response to the operation of triggering the generation of an emoticon package, the electronic device extracts features of a selected first target file to obtain first feature information. The first target file includes a video file or a picture.
[0251] The first target file may be a video file, or the first target file may be a picture.
[0252] 3 (d), the first target file may be video 221. The triggering operation of generating an emoticon package may be a user clicking on the "generate emoticon package" operation, after which the electronic device may perform feature extraction on the video 221 to obtain first feature information.
[0253] For example, the electronic device may extract features from each frame of the video 221 to obtain first feature information corresponding to each frame. The first feature information may include one or more of expression, facial features, background, action, text, and corresponding tag information.
[0254] The electronic device may also perform feature extraction on the continuous multiple-frame images included in the video 221 to obtain first feature information corresponding to the continuous multiple-frame images.
[0255] In some cases, the expression may also be an expression of a cartoon character or an animal.
[0256] 520. The electronic device performs feature matching based on the first feature information and recalls multiple target features.
[0257] In some embodiments, the electronic device may perform feature matching from a cloud server based on the first feature information, where the cloud server includes feature information of an online emoticon package.
[0258] For example, the cloud server can obtain a large amount of network emoticon package data through a web crawler, perform feature extraction on the network emoticon package, and store the feature information of the extracted network emoticon package in a feature index library, so that the electronic device can perform feature matching from the feature index library in the cloud server based on the first feature information and recall multiple target features.
[0259] 530. The electronic device generates a plurality of emoticon packages corresponding to the first target file according to the plurality of target features, wherein the plurality of emoticon packages include static emoticon packages and / or dynamic emoticon packages.
[0260] For example, if the first target file is a video file, the electronic device can generate one or more emoticon packages corresponding to each frame based on the multiple target features recalled from each frame. The electronic device can also generate one or more emoticon packages corresponding to multiple consecutive frames based on the multiple target features recalled from multiple consecutive frames. The emoticon packages can be dynamic emoticon packages, such as GIF emoticon packages.
[0261] Exemplarily, if the first target file is a picture, the electronic device may generate one or more corresponding emoticon packages based on multiple target features recalled from the picture.
[0262] Based on the embodiment of the present application, when an operation that triggers the generation of an emoticon package is detected, the electronic device can extract features of the first target file selected by the user, and perform feature matching based on the obtained first feature information to obtain multiple target features; and generate corresponding multiple emoticon packages based on the multiple target features.
[0263] In this way, the electronic device can generate multiple emoticon packages conveniently according to the user's simple operation. In addition, the emoticon package is generated by the video or picture selected by the user, thereby improving the diversity and fun of the user's creation of the emoticon package.
[0264] In some embodiments, the first target file is a video file including multiple frames of images, wherein the electronic device performs feature extraction on the selected first target file to obtain first feature information, including:
[0265] The electronic device divides the video file into multiple frames of images;
[0266] The electronic device performs feature extraction on each frame of the multiple frames of images to obtain first feature information corresponding to each frame of the images.
[0267] Exemplarily, the first feature information may include expressions, facial features, background, actions, accompanying text and corresponding tag information, etc.
[0268] Based on the embodiment of the present application, if the first target file is a video file, the electronic device can first divide the video file into multiple frames of images when performing feature extraction, and perform feature extraction on each frame of the multiple frames respectively, so as to obtain feature information of each frame of the image.
[0269] In some embodiments, the electronic device generates multiple emoticon packages corresponding to the first target file according to multiple target features, including:
[0270] The electronic device generates an emoticon package corresponding to each frame of image according to multiple target features corresponding to each frame of image.
[0271] Based on the embodiments of the present application, the electronic device generates corresponding emoticon packages based on multiple target features corresponding to each frame of the image. In this way, the electronic device can generate multiple emoticon packages based on a video file, thereby increasing the richness of the generated emoticon packages and providing users with a variety of emoticon packages for users to choose from.
[0272] In some embodiments, generating an expression package corresponding to each frame of image according to multiple target features corresponding to each frame of image includes:
[0273] Sort the multiple target features corresponding to each frame image separately;
[0274] The caption corresponding to the first-ranked target feature is added to each frame of the image to generate an emoticon package corresponding to each frame of the image.
[0275] Based on the embodiment of the present application, for the target image frame, the electronic device can sort the multiple target features of feature matching, and select the text corresponding to the target feature ranked first and add it to the target image frame, so as to generate a corresponding emoticon package.
[0276] In other examples, if the target feature ranked first does not correspond to a text, an emoticon package corresponding to each frame of the image can be directly generated without adding a text.
[0277] In some embodiments, the method 500 further includes:
[0278] The electronic device extracts features from a preset number of continuous images in the multiple frames of images to obtain first feature information corresponding to the continuous images.
[0279] For example, the preset number may be 3 or 5, etc.
[0280] Based on the embodiment of the present application, the electronic device can also perform feature extraction on multiple consecutive frames of images to obtain corresponding first feature information.
[0281] For example, the electronic device may extract features from three consecutive image frames and then stitch together the features of each image frame to serve as the first feature information corresponding to the three consecutive image frames.
[0282] In some embodiments, the method 500 further includes:
[0283] The electronic device generates an emoticon package corresponding to the continuous images according to a plurality of target features corresponding to the continuous images, wherein the emoticon package corresponding to the continuous images is a dynamic emoticon package.
[0284] Exemplarily, referring to (e) in FIG. 3 , the dynamic emoticon package may be the emoticon package 244 .
[0285] Based on the embodiments of the present application, the electronic device can also generate dynamic emoticon packages, thereby improving the diversity of the generated emoticon packages.
[0286] In some embodiments, generating expression packages corresponding to the consecutive images based on multiple target features corresponding to the consecutive images includes:
[0287] Sort multiple target features corresponding to consecutive images respectively;
[0288] The caption corresponding to the first-ranked target feature is added to the continuous images to generate the emoticon package corresponding to the continuous images.
[0289] Exemplarily, the electronic device may sort the multiple target features according to their similarity to the first feature information, for example, sorting them in descending order of similarity, and obtaining a sorting result of “feature 1, feature 2, feature 3…”.
[0290] Alternatively, the electronic device can also sort the feedback data of all users who use the function based on big data statistics. For example, for similar features, most users prefer target feature 2, so target feature 2 can be ranked first, and the sorting result is "feature 2, feature 1, feature 3..."
[0291] Based on the embodiment of the present application, for continuous image frames, the electronic device can sort multiple target features of feature matching, and select the text corresponding to the target feature ranked first and add it to the continuous image frame, so as to automatically generate an emoticon package with corresponding text.
[0292] In some embodiments, performing feature matching based on the first feature information includes:
[0293] The electronic device performs feature matching from a cloud server based on the first feature information, where the cloud server includes feature information of the network emoticon package.
[0294] For example, the cloud server may include a feature index library, which includes feature information of network emoticon packages. It should be understood that the corresponding network emoticon package can be determined based on the feature information.
[0295] Based on the embodiment of the present application, the electronic device can perform feature matching from the feature index library of the cloud server according to the first feature information. Since the feature index library contains a large amount of feature information of online emoticon packages, the matching results can be made more accurate.
[0296] In some embodiments, the method 500 further includes:
[0297] A plurality of emoticon packages are displayed in the first display interface, and the first display interface also includes function buttons for performing target operations on each of the plurality of emoticon packages.
[0298] For example, referring to (e) in FIG. 3 , the first display interface may be the display interface 240 , and the target operation may be editing, liking, disliking, sharing, storing, or the like.
[0299] Exemplarily, the first display interface may be a display interface of a gallery application, or may be a display interface of other instant messaging applications.
[0300] Based on the embodiment of the present application, the electronic device can also display the generated emoticon package in the display interface, so as to better show the generated emoticon package to the user.
[0301] In some embodiments, the method 500 further includes:
[0302] In response to a user clicking an edit function button of the dynamic emoticon package, a second display interface is displayed, where the second display interface includes a plurality of function buttons for editing the dynamic emoticon package;
[0303] In response to a user clicking a function button for editing the number of frames of the dynamic emoticon package among the plurality of function buttons, a third display interface is displayed, the third display interface including the plurality of frame images included in the dynamic emoticon package;
[0304] In response to a user's operation of clicking a function button for deleting a target image, the target image is deleted.
[0305] For example, referring to (e)-(h) in FIG3 , the edit function button may be function button 2441, the second display interface may be display interface 250, the multiple function buttons may be pause function button 251, rewind function button 252, edit text function button 253, edit size function button 254, and edit frame number function button 255, etc. The third display interface may be display interface 260, the target image may be image 264, and the function button for deleting the target image may be delete function button 2641.
[0306] Based on the embodiments of the present application, users can also edit the generated emoticon package, such as deleting frames that the user does not like in the emoticon package, thereby improving the user's operability and providing the possibility for users to re-create emoticon packages.
[0307] FIG13 is a schematic flow chart of a method for generating an emoticon package provided by an embodiment of the present application. As shown in FIG13 , the method 600 can be applied to a cloud server, and the method 600 can include steps 610 to 630.
[0308] At 610 , the cloud server obtains an online emoticon package, which includes a static emoticon package and a dynamic emoticon package.
[0309] For example, the cloud server may obtain the emoticon package from the Internet by means of a web crawler. For details, please refer to the relevant description in the above text.
[0310] 620. The cloud server performs feature extraction on each emoticon package in the static emoticon package to obtain second feature information of each emoticon package, and performs feature extraction on the continuous multi-frame images included in each emoticon package in the dynamic emoticon package to obtain third feature information of the continuous multi-frame images.
[0311] The cloud server can perform feature extraction on the acquired online emoticon packages. For example, for static emoticon packages, the cloud server can perform feature extraction on them to obtain corresponding second feature information. For dynamic emoticon packages, the cloud server can perform feature extraction on the consecutive multiple frames of images included in each emoticon package to obtain corresponding third feature information.
[0312] The second characteristic information and the third characteristic information may include one or more of expressions, facial features, background, actions, accompanying text and corresponding tag information.
[0313] Exemplarily, the second feature information and the third feature information may be used by the electronic device to perform feature matching to retrieve multiple matching target features.
[0314] 630. The cloud server stores the second feature information and the third feature information.
[0315] Exemplarily, the cloud server may store the second feature information and the third feature information in a feature index library for feature matching by the electronic device.
[0316] Based on the embodiment of the present application, the cloud server can obtain a large number of online emoticon packages, extract features from the online emoticon packages respectively, obtain corresponding feature information, and store the feature information to facilitate subsequent feature matching by electronic devices.
[0317] Figure 14 is a schematic block diagram of an electronic device provided in an embodiment of the present application. As shown in Figure 14, the electronic device 700 may include one or more processors 710; one or more memories 720; the one or more memories 720 store one or more instructions, which, when executed by the one or more processors 710, enable the method for generating an emoticon package as described in any of the possible implementations described above to be executed.
[0318] Exemplarily, the electronic device 700 may be the electronic device 100, the electronic device 200, the cloud server, etc. mentioned above.
[0319] An embodiment of the present application also provides a device including a processor and a communication interface, the communication interface being used to receive a signal and transmit the signal to the processor, the processor processing the signal so that the method for generating an emoticon package as described in any possible implementation method described above is executed.
[0320] The device may be a chip. For example, the chip may be a chip system or an independent chip.
[0321] An embodiment of the present application also provides a readable storage medium, which stores instructions. When the instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the method for generating an emoticon package in the above-mentioned embodiment.
[0322] An embodiment of the present application also provides a program product. When the program product is run on an electronic device, the electronic device executes the above-mentioned related steps to implement the method for generating an emoticon package in the above-mentioned embodiment.
[0323] An embodiment of the present application also provides a device for generating an emoticon package, including a module for implementing the method for generating an emoticon package as described in any of the embodiments above.
[0324] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store instructions, and when the device is running, the processor can execute the instructions stored in the memory to enable the device to execute the method of generating emoticon packages in the above-mentioned method embodiments.
[0325] Among them, the equipment, readable storage medium, program product or device provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.
[0326] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0327] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0328] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0329] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0330] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0331] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0332] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for generating expression packs, characterized in that, The method is applied to an electronic device, and the method includes: In response to an operation of triggering the generation of an emoji, perform feature extraction on a selected first target file to obtain first feature information, where the first target file includes a video file or a picture; Perform feature matching according to the first feature information to recall multiple target features; Generate multiple emojis corresponding to the first target file according to the multiple target features, where the multiple emojis include static emojis and / or dynamic emojis.
2. The method according to claim 1, wherein The first target file is a video file, and the video file includes multiple frames of images. Among them, the performing feature extraction on the selected first target file to obtain first feature information includes: Split the video file into the multiple frames of images; Perform feature extraction on each frame of the multiple frames of images respectively to obtain the first feature information corresponding to each frame of image.
3. The method according to claim 2, wherein The generating multiple emojis corresponding to the first target file according to the multiple target features includes: Generate emojis corresponding to each frame of image respectively according to the multiple target features corresponding to each frame of image.
4. The method according to claim 3, wherein The generating emojis corresponding to each frame of image respectively according to the multiple target features corresponding to each frame of image includes: Sort the multiple target features corresponding to each frame of image respectively; Add the caption corresponding to the target feature ranked first to each frame of image to generate the emoji corresponding to each frame of image.
5. The method according to any one of claims 2 to 4, characterized in that, The method further includes: Perform feature extraction on a preset number of consecutive images in the multiple frames of images to obtain the first feature information corresponding to the consecutive images.
6. The method according to claim 5, characterized in that The method further includes: Generate an emoji corresponding to the consecutive images according to the multiple target features corresponding to the consecutive images, where the emoji corresponding to the consecutive images is a dynamic emoji.
7. The method according to claim 6, wherein The generating an emoji corresponding to the consecutive images according to the multiple target features corresponding to the consecutive images includes: Sort the multiple target features corresponding to the consecutive images respectively; Add the caption corresponding to the target feature ranked first to the consecutive images to generate the emoji corresponding to the consecutive images.
8. The method according to any one of claims 1-7, characterized in that, The performing feature matching according to the first feature information includes: Perform feature matching from a cloud server according to the first feature information, where the cloud server includes feature information of online emojis.
9. The method according to any one of claims 1-7, characterized in that, The performing feature matching according to the first feature information to recall multiple target features includes: Perform feature matching from a cloud server according to the first feature information and label information provided by a user, where the cloud server includes feature information of online emojis.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: Display the multiple emojis in a first display interface, and the first display interface further includes function buttons for performing target operations on each of the multiple emojis respectively.
11. The method according to claim 10, wherein The method further includes: In response to an operation of a user clicking an edit function button of the dynamic emoji, display a second display interface, where the second display interface includes multiple function buttons for editing the dynamic emoji; In response to the user clicking on the function button for editing the number of frames of the dynamic emoji among the multiple function buttons, a third display interface is displayed, and the third display interface includes multiple frames of images included in the dynamic emoji. In response to the user clicking on the function button for deleting the target image, the target image is deleted.
12. A method for generating emoticons, characterized in that, The method is applied to a cloud server, and the method includes: Obtaining network emojis, where the network emojis include static emojis and dynamic emojis; Performing feature extraction on each emoji in the static emojis to obtain second feature information of each emoji, and performing feature extraction on consecutive multiple frames of images included in each emoji in the dynamic emojis to obtain third feature information of the consecutive multiple frames of images; Storing the second feature information and the third feature information.
13. An electronic device, characterized in that, Including: One or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the method for generating emojis according to any one of claims 1-11 is executed.
14. A cloud server, characterized in that, Including: One or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the method for generating emojis according to claim 12 is executed.
15. A chip, characterized in that, The chip includes a processor and a communication interface. The communication interface is used to receive a signal and transmit the signal to the processor, and the processor processes the signal so that the method for generating emojis according to any one of claims 1-12 is executed.
16. A readable storage medium, characterized in that, Instructions are stored in the readable storage medium, and when the instructions are run on an electronic device, the method for generating emojis according to any one of claims 1-12 is executed.
Citation Information
Patent Citations
Facial expression generation method and device, terminal and storage medium
CN107977928A
Expression production method, device, terminal and computer readable storage medium
CN108280166A
Emoji package picture processing method and device and server
CN114880512A
Image searching method and electronic equipment
CN117076702A