Emoji Generation Method, Device, Electronic Device and Computer Readable Storage Medium

Through the automated emoticon package generation method, users can quickly generate high-quality emoticon packages, solving the problems of difficulty and low efficiency of user-made emoticon packages, and improving user participation and production efficiency.

CN113538628BActive Publication Date: 2025-06-03GUANGZHOU KUGOU COMP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110744504.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-30
Publication Date
2025-06-03
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

It is difficult and inefficient for users to make emoticons, and usually requires professionals to make them, and the user participation is not high.

Method used

Provides a method for generating emoticons, which automatically generates emoticons by obtaining the original material and determining the emoticons category and generation parameters. The method includes obtaining materials, determining categories and parameters, and using this information to generate emoticons.

Benefits of technology

It reduces the difficulty and time for users to make emoticons, improves the efficiency and user participation of emoticons, and realizes the flexibility and diversity of emoticons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113538628B_ABST
    Figure CN113538628B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of information processing, and discloses a method, device, electronic device and computer-readable storage medium for generating emoticons. The method for generating emoticons includes: obtaining original materials for generating emoticons; determining the category of the emoticon to be generated and the emoticon generation parameters corresponding to the category of the emoticon; the category of the emoticon includes one or more of the following categories: audio emoticons, static picture emoticons, dynamic picture emoticons, video emoticons; generating an emoticon by using the original materials and the emoticon generation parameters. The present invention realizes the automatic generation of emoticons, reduces the difficulty for users to make emoticons, and also improves the production efficiency and flexibility of emoticons.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and particularly to a method, device, electronic device and computer-readable storage medium for generating expression packs. Background Art

[0002] In the mobile Internet era, relying on the continuous development of social networking and the Internet, the way people communicate with each other has also changed accordingly. It has evolved from the earliest text communication to the use of some simple symbols and emoji (visual emotion symbols), and gradually to the increasingly diversified expression pack culture. By using some self-made pictures of popular elements and matching them with a series of corresponding texts, specific emotions are expressed.

[0003] Generally, the production of expression packs involves a lot of image processing technology, which makes it difficult for users to produce them by themselves, resulting in low efficiency of expression pack production. Therefore, they are usually produced by professionals and then provided for users to use, with low user participation. Even when users make expression packs by themselves, they usually need to complete it manually, which reduces the efficiency of expression pack production. Summary of the Invention

[0004] The main purpose of the present invention is to provide a method, device, electronic device and computer-readable storage medium for generating expression packs, aiming to solve the technical problem of the great difficulty and low efficiency of users in making expression packs.

[0005] The first aspect of the present invention provides a method for generating an expression pack, the method comprising:

[0006] Obtaining original materials for generating an expression pack;

[0007] Determining the category of the expression pack to be generated and the expression pack generation parameters corresponding to the category of the expression pack; the category of the expression pack includes one or more of the following categories: audio expression packs, static picture expression packs, dynamic picture expression packs, video expression packs;

[0008] Generating an expression pack by using the original materials and the expression pack generation parameters.

[0009] Optionally, in the first implementation manner of the first aspect of the present invention, when the determined category of the expression pack to be generated includes an audio expression pack, the generating an expression pack by using the original materials and the expression pack generation parameters includes:

[0010] Obtaining the text data corresponding to the original materials;

[0011] Generating audio data by using the musical score data and the text data; the musical score data is the expression pack generation parameter corresponding to the audio expression pack;

[0012] Generate an audio meme corresponding to the original material by using the audio data.

[0013] Optionally, in the second implementation manner of the first aspect of the present invention, the obtaining of the text data corresponding to the original material includes:

[0014] When the original material is an audio material, convert the audio material into text to obtain the text data;

[0015] When the original material is a picture material, recognize the text information or semantic information in the picture material to obtain the text data.

[0016] Optionally, in the third implementation manner of the first aspect of the present invention, when the original material is an audio material, the generating of the audio data by using the sheet music data and the text data includes:

[0017] Recognize the user pronunciation characteristics corresponding to the audio material;

[0018] Generate audio data by using the sheet music data, the text data, and the user pronunciation characteristics.

[0019] Optionally, in the fourth implementation manner of the first aspect of the present invention, the generating of the audio meme corresponding to the original material by using the audio data includes:

[0020] Generate an audio meme corresponding to the original material by using the audio data and the picture selected by the user; the audio meme is a meme carrying the audio data in the picture.

[0021] Optionally, in the fifth implementation manner of the first aspect of the present invention, when the determined meme category to be generated includes static picture memes, dynamic picture memes, or video memes, the generating of the meme by using the original material and the meme generation parameters includes:

[0022] Perform preprocessing on the original material;

[0023] Process the preprocessed original material according to the meme generation parameters to obtain a meme;

[0024] Among them, the performing of the preprocessing on the original material includes:

[0025] When the original material includes a picture material, recognize the foreground image and the background image in the picture material, and set the channel of the background image to a transparent channel;

[0026] Perform content feature recognition on the foreground image, and based on the content feature recognition result, determine whether the foreground image contains a person or an animal with facial features;

[0027] If so, perform facial feature recognition on the person or animal to obtain the facial feature points of the person or animal in the foreground image.

[0028] Optionally, in the seventh implementation manner of the first aspect of the present invention, when the determined category of the meme to be generated includes static picture memes, the processing of the preprocessed original material according to the meme generation parameters to obtain the meme includes:

[0029] Transform the facial feature points of the person or animal in the foreground image according to the facial feature point transformation matrix to obtain an expression image;

[0030] Add a text watermark to the expression image according to the picture text description input by the user to obtain a static picture meme; the facial feature point transformation matrix and the picture text description are the meme generation parameters corresponding to the static picture meme.

[0031] Optionally, in the eighth implementation manner of the first aspect of the present invention, when the determined category of the meme to be generated includes dynamic picture memes, the processing of the preprocessed original material according to the meme generation parameters to obtain the meme includes:

[0032] Successively transform the facial feature points of the person or animal in the foreground image according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp to obtain multiple frames of expression images with a time sequence;

[0033] Perform dynamic picture synthesis on each frame of the expression image to obtain a dynamic picture meme; the timestamp and the facial feature point transformation matrix are the meme generation parameters corresponding to the dynamic picture meme.

[0034] Optionally, in the ninth implementation manner of the first aspect of the present invention, when the determined category of the meme to be generated includes video memes, the processing of the preprocessed original material according to the meme generation parameters to obtain the meme includes:

[0035] Generate the expression audio of the person or animal according to the facial features recognized in the foreground image;

[0036] Calculate the volume corresponding to each frame of the expression audio according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to the timestamps, and sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple expression images with a time sequence;

[0037] Perform audio-video synthesis on the expression audio and each frame of the expression images to obtain a video file;

[0038] Associate and save the video file with a preset video expression icon and a preset video display icon respectively to obtain a video emoji package.

[0039] Optionally, in the tenth implementation manner of the first aspect of the present invention, when the determined emoji category to be generated includes video emojis, the processing of the preprocessed original material according to the emoji generation parameters to obtain emojis includes:

[0040] When the original material further includes text material, perform audio synthesis on the text material to obtain the expression audio of the person or animal;

[0041] Calculate the volume corresponding to each frame of the expression audio according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to the timestamps, and sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple expression images with a time sequence;

[0042] Perform audio-video synthesis on the expression audio and each frame of the expression images to obtain a video file;

[0043] Associate and save the video file with a preset video expression icon and a preset video display icon respectively to obtain a video emoji package.

[0044] Optionally, in the eleventh implementation manner of the first aspect of the present invention, after processing the preprocessed original material according to the emoji generation parameters to obtain emojis, it further includes:

[0045] Detect whether there is a touch operation on the video expression icon;

[0046] If so, obtain the video file associated with the video expression icon and obtain the video display icon associated with the video file;

[0047] Send the video file and the video display icon.

[0048] Optionally, in the twelfth implementation manner of the first aspect of the present invention, after processing the preprocessed original material according to the emoji generation parameters to obtain emojis, it further includes:

[0049] Receive the video file and the video display icon;

[0050] Display the video display icon and the playing duration of the video file, and play the video file.

[0051] The second aspect of the present invention provides a meme generation device, and the device includes:

[0052] An acquisition module, configured to acquire original materials for generating memes;

[0053] A determination module, configured to determine the type of meme to be generated and the meme generation parameters corresponding to the type of meme; the type of meme includes one or more of the following types: audio memes, static picture memes, dynamic picture memes, video memes;

[0054] A generation module, configured to generate memes by using the original materials and the meme generation parameters.

[0055] Optionally, in the first implementation manner of the second aspect of the present invention, when the determined type of meme to be generated includes audio memes, the generation module is specifically configured to:

[0056] Obtain the text data corresponding to the original materials;

[0057] Generate audio data by using the musical score data and the text data; the musical score data is the meme generation parameter corresponding to the audio meme;

[0058] Generate the audio meme corresponding to the original materials by using the audio data.

[0059] Optionally, in the second implementation manner of the second aspect of the present invention, the generation module is further configured to:

[0060] When the original materials are audio materials, convert the audio materials into text to obtain the text data;

[0061] When the original materials are picture materials, recognize the text information or semantic information in the picture materials to obtain the text data.

[0062] Optionally, in the third implementation manner of the second aspect of the present invention, when the original materials are audio materials, the generation module is further configured to:

[0063] Recognize the user pronunciation characteristics corresponding to the audio materials;

[0064] Generate audio data by using the musical score data, the text data and the user pronunciation characteristics.

[0065] Optionally, in the fourth implementation manner of the second aspect of the present invention, the generating module is further configured to:

[0066] Generate an audio meme corresponding to the original material by using the audio data and the picture selected by the user; the audio meme is a meme carrying the audio data in the picture.

[0067] Optionally, in the fifth implementation manner of the second aspect of the present invention, when the determined meme category to be generated includes static picture memes, dynamic picture memes or video memes, the generating module includes:

[0068] A preprocessing unit, configured to preprocess the original material;

[0069] A processing unit, configured to process the preprocessed original material according to the meme generation parameters to obtain a meme;

[0070] Wherein, the preprocessing unit is specifically configured to:

[0071] When the original material includes picture material, identify the foreground image and the background image in the picture material, and set the channel of the background image as a transparent channel;

[0072] Perform content feature recognition on the foreground image, and determine whether the foreground image contains a person or an animal with facial features according to the content feature recognition result;

[0073] If so, identify the facial feature points of the person or animal in the foreground image to obtain the facial feature points of the person or animal in the foreground image.

[0074] Optionally, in the seventh implementation manner of the second aspect of the present invention, when the determined meme category to be generated includes static picture memes, the processing unit is specifically configured to:

[0075] Transform the facial feature points of the person or animal in the foreground image according to the facial feature point transformation matrix to obtain an expression picture;

[0076] Add a text watermark to the expression picture according to the picture text description input by the user to obtain a static picture meme; the facial feature point transformation matrix and the picture text description are the meme generation parameters corresponding to the static picture memes.

[0077] Optionally, in the eighth implementation manner of the second aspect of the present invention, when the determined meme category to be generated includes dynamic picture memes, the processing unit is specifically configured to:

[0078] According to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each of the timestamps, the facial feature points of the person or animal in the foreground image are sequentially transformed to obtain multiple frames of expression images with a time sequence.

[0079] Perform dynamic image synthesis on each frame of the expression images to obtain a dynamic image emoji; the timestamp and the facial feature point transformation matrix are the emoji generation parameters corresponding to the dynamic image type emoji.

[0080] Optionally, in the ninth implementation manner of the second aspect of the present invention, when the determined emoji category to be generated includes video type emojis, the processing unit is specifically configured to:

[0081] Generate the expression audio of the person or animal according to the facial features recognized in the foreground image.

[0082] According to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, calculate the volume corresponding to each frame of audio in the expression audio, and sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression images with a time sequence.

[0083] Perform audio-video synthesis on the expression audio and each frame of the expression images to obtain a video file.

[0084] Associate and save the video file with a preset video expression icon and a preset video display icon respectively to obtain a video emoji.

[0085] Optionally, in the tenth implementation manner of the second aspect of the present invention, when the determined emoji category to be generated includes video type emojis, the processing unit is specifically configured to:

[0086] When the original material further includes text material, perform audio synthesis on the text material to obtain the expression audio of the person or animal.

[0087] According to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, calculate the volume corresponding to each frame of audio in the expression audio, and sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression images with a time sequence.

[0088] Perform audio-video synthesis on the expression audio and each frame of the expression images to obtain a video file.

[0089] Associate and save the video file with a preset video expression icon and a preset video display icon respectively to obtain a video emoji.

[0090] Optionally, in the eleventh implementation manner of the second aspect of the present invention, the emoji generation device further includes:

[0091] A detection module, configured to detect whether there is a touch operation on the video emoji icon;

[0092] An obtaining module, configured to, if there is a touch operation on the video emoji icon, obtain a video file associated with the video emoji icon and obtain a video display icon associated with the video file;

[0093] A sending module, configured to send the video file and the video display icon.

[0094] Optionally, in the twelfth implementation manner of the second aspect of the present invention, the emoji generation device further includes:

[0095] A receiving module, configured to receive the video file and the video display icon;

[0096] A playing module, configured to display the video display icon and the playing duration of the video file, and play the video file.

[0097] The third aspect of the present invention provides an electronic device, including: a memory and at least one processor, where instructions are stored in the memory; the at least one processor invokes the instructions in the memory to enable the electronic device to execute the above-mentioned emoji generation method.

[0098] The fourth aspect of the present invention provides a computer-readable storage medium, where instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is enabled to execute the above-mentioned emoji generation method.

[0099] In the technical solution provided by the present invention, the user only needs to provide the materials for making emojis, and the system will automatically determine the category of the emojis to be generated and the emoji generation parameters corresponding to the emoji category; finally, the original materials and the emoji generation parameters are used to generate emojis. The process of making emojis is automatically completed without user operation, thereby reducing the difficulty of making emojis by the user and improving the efficiency of making emojis at the same time. In addition, since the materials are provided by the user himself / herself and the emoji generation parameters can also be set by the user himself / herself, the flexibility and diversity of making emojis are realized, and the enthusiasm and experience of user participation are improved. Description of the Drawings

[0100] Figure 1 It is a schematic diagram of the first embodiment of the emoji generation method in the embodiment of the present invention;

[0101] Figure 2 It is a schematic diagram of the second embodiment of the emoji generation method in the embodiment of the present invention;

[0102] Figure 3 Schematic diagram of the third embodiment of the emoji generation method in the embodiments of the present invention;

[0103] Figure 4 Schematic diagram of the fourth embodiment of the emoji generation method in the embodiments of the present invention;

[0104] Figure 5 Schematic diagram of the fifth embodiment of the emoji generation method in the embodiments of the present invention;

[0105] Figure 6 Schematic diagram of the sixth embodiment of the emoji generation method in the embodiments of the present invention;

[0106] Figure 7 Schematic diagram of an embodiment of the emoji generation device in the embodiments of the present invention;

[0107] Figure 8 Schematic diagram of an embodiment of the electronic device in the embodiments of the present invention. Detailed implementation manners

[0108] The embodiments of the present invention provide an emoji generation method, device, electronic device and computer-readable storage medium. The present invention reduces the difficulty for users to make emojis and also improves the efficiency of emoji making. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the term "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0109] To facilitate the understanding of the present invention, relevant terms involved in the embodiments of the present application are first introduced.

[0110] (1) Emoji refers to a form of expression that uses pictures to represent personal emotions, generally in the form of static pictures or dynamic pictures. The present invention further expands the form of expression of emojis and creatively proposes audio emojis and video emojis with audio.

[0111] (2) Original materials refer to the basic materials used to make emoticons, and generally pictures are used to make emoticons. The present invention further expands the scope of materials for emoticons. In addition to the commonly used picture materials, audio materials, text materials, and video materials are creatively proposed.

[0112] (3) Emotion types refer to the emotions of the made emoticons. Generally, emotion-based expressions such as crying, laughing, sad, happy, excited, nervous, and scared are used. The present invention further expands the scope of expressions of emoticons. In addition to the conventional emotion-based expressions, scene expressions (such as scenes of working overtime, shopping, dining, etc.) and action expressions (such as actions of goodbye, shaking hands, making a fist, raising a leg, etc.) are creatively proposed.

[0113] (4) Emoticon categories refer to the classification of emoticons. Emoticon categories include one or more of the following categories: audio-based emoticons, static picture-based emoticons, dynamic picture-based emoticons, video-based emoticons.

[0114] (5) Emoticon generation parameters are used to make emoticons. Different emoticon types correspond to different emoticon generation parameters. Preferably, each emoticon type corresponds to one emoticon generation parameter, and each emoticon generation parameter includes configuration parameters for making one or more categories of emoticons. Emoticon generation parameters can either be personalized settings by users themselves or default settings by developers.

[0115] For the convenience of further understanding, the specific process of the embodiment of the present invention is described below. For the technical solution provided by the embodiment of the present invention, the execution subject of each step can be an electronic device. In one possible implementation, the electronic device can be a terminal device such as a smart phone, a tablet computer, or a personal computer. In another possible implementation, the electronic device can also be a smart TV.

[0116] Please refer to Figure 1 , the first embodiment of the emoticon generation method in the embodiment of the present invention includes:

[0117] 101. Obtain the original materials for generating emoticons;

[0118] In this embodiment, the production of emoticons requires corresponding original materials. For example, picture emoticons require pictures, animated emoticons require one or more pictures, and voice emoticons require voices, etc. The materials in this embodiment need to be provided in advance, which can be provided by users themselves or be system-built-in. It should be noted that the materials in this embodiment are not emoticon template materials and can be any pictures, texts, audios, videos, etc.

[0119] In one embodiment, preferably, the materials for making emoticons include: picture materials, audio materials, text materials, and video materials.

[0120] Optionally, the expression type corresponding to the emoji can be further obtained. Emojis of different expression types have different audio-visual feelings and convey different information.

[0121] In one embodiment, the expression types preferably expressed by the emojis include:

[0122] A. Emotional expressions, such as crying, laughing, sadness, happiness, excitement, nervousness, fear, etc.;

[0123] B. Scene expressions, such as scenes of overtime work, shopping, dining together, etc.

[0124] C. Action expressions, such as actions of goodbye, shaking hands, making a fist, raising a leg, etc.

[0125] 102. Determine the emoji category to be generated and the emoji generation parameters corresponding to the emoji category; the emoji category includes one or more of the following categories: audio emojis, static picture emojis, dynamic picture emojis, video emojis;

[0126] In this embodiment, after obtaining the materials and expression types, it is also necessary to further determine the emoji category to be generated and the emoji generation parameters corresponding to the emoji category. Preferably, each emoji type corresponds to an emoji generation parameter. In this embodiment, the determination method of the emoji category is not limited, and it can be selected by the user or randomly selected by the machine.

[0127] In this embodiment, the emoji generation parameters are specifically determined according to the material type, emoji category, etc. For example, the picture size, dimensions, format requirements, video size, duration, whether to have audio, etc., and the font, color, etc. of the text can all be used as emoji generation parameters.

[0128] In addition, to improve the flexibility of emoji production, in one embodiment, the emoji generation parameters can be set by the user according to the actual production needs of the emojis.

[0129] 103. Generate emojis using the original materials and emoji generation parameters.

[0130] In this embodiment, the processing method of the original materials is not limited and is specifically determined according to the emoji production needs. For example, perform size scaling processing, format conversion processing, picture compression processing, etc. on the picture materials, and perform style layout processing on the text materials. During the processing of the original materials, the materials are further processed according to the emoji generation parameters to obtain the corresponding emojis.

[0131] Since the present invention supports a large number of material types, various types of emoticons can be made. In one embodiment, preferably, the emoticons that can be made include: static picture emoticons, dynamic picture emoticons, video emoticons, and audio emoticons.

[0132] The present invention also supports users to customize emoticons. Optionally, in one embodiment, before the above step 101, it further includes:

[0133] Obtain the emoticon generation mode selected by the user. The emoticon generation mode includes one-key generation mode and custom mode;

[0134] If the emoticon generation mode is the one-key generation mode, then execute steps 101-103. Otherwise, perform UI (User Interface) human-computer interaction operations. Through the UI human-computer interaction operations, obtain the custom materials for the emoticons to be generated, and generate the custom emoticon generation parameters;

[0135] After the UI human-computer interaction operations are completed, process the custom materials according to the custom emoticon generation parameters to obtain custom emoticons.

[0136] In this embodiment, for emotion-based emoticons with simple meaning expressions, the one-key generation mode is preferably used for production. For scene-based emoticons and action-based emoticons with rich meaning expressions, the custom mode is preferably used for production. In the custom mode, the custom materials can be not only pictures, texts, and audios, but also videos. Users can edit the custom materials or associate and combine different custom materials to form scene-based emoticons or action-based emoticons with richer meaning expressions.

[0137] In this embodiment, users only need to provide the materials for making emoticons, and the system will automatically determine the category of the emoticons to be generated and the emoticon generation parameters corresponding to the emoticon category; finally, use the original materials and the emoticon generation parameters to generate emoticons. The emoticon production process is automatically completed without user operation, thereby reducing the difficulty of users making emoticons and improving the emoticon production efficiency at the same time. In addition, since the materials are provided by the users themselves and the emoticon generation parameters can also be set by the users themselves, the flexibility and diversity of emoticon production are realized, and the user participation enthusiasm and experience are improved.

[0138] Next, taking emotion-based emoticons as an example, the specific generation methods of various emoticons will be illustrated by using the one-key generation mode.

[0139] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the generation process of an embodiment of the audio emoticon of the present invention. The second embodiment of the emoticon generation method in the embodiment of the present invention includes:

[0140] 201. Obtain the original materials for generating emoticons.

[0141] 202. Determine the category of the emoticon to be generated and the emoticon generation parameters corresponding to the emoticon category; the emoticon category includes one or more of the following categories: audio emoticons, static picture emoticons, dynamic picture emoticons, video emoticons.

[0142] In this embodiment, the descriptions of the above steps 201 and 202 refer to the first embodiment, and will not be elaborated here. The implementation process of generating audio emoticons will be specifically described below.

[0143] 203. When the determined category of the emoticon to be generated includes audio emoticons, obtain the text data corresponding to the original materials.

[0144] In this embodiment, there is no limit to the method of obtaining the text data corresponding to the original materials, and it is preferably obtained by processing the original materials.

[0145] Optionally, in one embodiment, the above step 203 includes:

[0146] When the original material is an audio material, convert the audio material into text to obtain the text data.

[0147] When the original material is a picture material, identify the text information or semantic information in the picture material to obtain the text data.

[0148] In this optional embodiment, if the original material is an audio material, perform text recognition on the audio material to obtain the corresponding text data. If the original material is a picture material, perform recognition on the text or semantics contained in the picture material to obtain the corresponding text data.

[0149] 204. Use the musical score data and the text data to generate audio data; the musical score data is the emoticon generation parameter corresponding to the audio emoticon.

[0150] In this embodiment, to further perform personalized processing on the audio data, through the musical score data set in the emoticon generation parameters, the obtained text data is converted into audio data with personalized pronunciation characteristics. Among them, the musical score data contains multiple notes and the audio corresponding to each note.

[0151] Optionally, in one embodiment, the above step 204 includes:

[0152] When the original material is an audio material, identify the user pronunciation characteristics corresponding to the audio material.

[0153] Use the musical score data, the text data and the user pronunciation characteristics to generate audio data.

[0154] In this alternative embodiment, when the original material is audio material, the user's pronunciation features of the audio material are further identified, and then according to the sheet music data and the user's voice features, the text data is converted into corresponding audio data with the user's pronunciation features, so as to realize the personalized characteristics of the audio data. For example, it can be to use the user's own pronunciation features or the pronunciation features of a certain person to generate audio emoticons.

[0155] For example, a section of the user's voice is pre-recorded as the audio material. When making an emoticon, first extract the timbre and pitch of this section of the voice to obtain the user's pronunciation features. One-to-one correspondence is established between the seven musical notes and seven audio, that is: C, D, E, F, G, A, B, and the Chinese expression forms are duo, lai, mi, fa, suo, la, xi, which are respectively in one-to-one correspondence with the 7 audio with the singing names of do, re, mi, fa, sol, la, si, so as to obtain 7 letter audios and use them as sheet music data. At the same time, these seven letter audios are respectively stored in the sheet music data in one-to-one correspondence with seven pitches (from pitch 1 to pitch 7).

[0156] Then, pitch extraction is performed on the recognized text data to obtain the pitch numbers of each character in the text data, and then according to the correspondence between the seven pitches and the seven letter audios in the sheet music data, the note pronunciations of each character in the text data are obtained.

[0157] Finally, according to the user's pronunciation features, the audio data corresponding to the text data is generated, and the generation of the audio emoticon is completed. The audio data uses the above-mentioned note pronunciations, so as to obtain an audio emoticon that has both the user's pronunciation features and a sense of rhythm.

[0158] 205. Generate an audio emoticon corresponding to the original material using the audio data.

[0159] In this embodiment, after generating the personalized audio data, the generated personalized audio data can be used to generate an audio emoticon corresponding to the original material. For example, after associating the audio data with the emoticon icon, it can be used as an audio emoticon.

[0160] Optionally, in an embodiment, the above step 205 includes:

[0161] Generate an audio emoticon corresponding to the original material using the audio data and the picture selected by the user; the audio emoticon is an emoticon with audio data carried in the picture.

[0162] In this embodiment, only the original materials need to be provided, and the system will automatically recognize the literal meaning of the original materials and convert them into audio, thereby improving the production efficiency of the emoji. In addition, this embodiment further associates the picture with the audio, so that the audio emoji not only has the expression way and meaning of the picture, but also has the expression way and meaning of the sound, thus further enriching the content expressed by the emoji.

[0163] In this embodiment, the audio emoji preferably consists of audio data, an audio emoji icon, and an audio display icon. Among them, the audio emoji icon can be generated from a picture associated with the audio data.

[0164] Optionally, in one embodiment, the audio emoji is sent in the following manner:

[0165] Detect whether there is a touch operation on the audio emoji icon;

[0166] If so, obtain the audio data associated with the audio emoji icon and obtain the audio display icon associated with the audio data;

[0167] Send the audio data and the audio display icon.

[0168] Optionally, in one embodiment, the audio emoji is received and played in the following manner:

[0169] Receive the audio data and the audio display icon;

[0170] Display the audio display icon and the playing duration of the audio data, and play the audio data.

[0171] This embodiment further expands the expression form of the emoji. In addition to pictures and animations, audio can also be used as an emoji. At the same time, personalized settings for voice pronunciation are added to the audio emoji, including personalized settings for audio features and emotional expressions, which not only improves the entertainment of the emoji, but also further provides more expression forms of the emoji, thereby enhancing the user experience.

[0172] Please refer to Figure 3 , Figure 3 This is the third embodiment of the emoji generation method in the embodiments of the present invention. In this embodiment, the emoji generation method includes:

[0173] 301. Obtain the original materials for generating the emoji;

[0174] 302. Determine the category of the emoji to be generated and the emoji generation parameters corresponding to the emoji category; the emoji category includes one or more of the following categories: audio emoji, static picture emoji, dynamic picture emoji, video emoji;

[0175] In this embodiment, the descriptions of steps 301 and 302 above refer to the first embodiment and will not be elaborated herein.

[0176] 303. When the determined categories of the emoticons to be generated include static picture emoticons, dynamic picture emoticons, or video emoticons, preprocess the original materials.

[0177] 304. Process the preprocessed original materials according to the emoticon generation parameters to obtain emoticons.

[0178] In this embodiment, there is no limit to the preprocessing method of the original materials, which is specifically determined according to the needs of emoticon production. For example, perform size scaling processing, format conversion processing, picture compression processing, etc. on picture materials, and perform style typesetting processing on text materials. After the preprocessing is completed, the actual processing and production of emoticons can be carried out, and the materials are processed specifically according to the emoticon generation parameters to obtain the corresponding emoticons.

[0179] Optionally, in an embodiment, when the original materials include picture materials, step 303 above includes:

[0180] Identify the foreground image and the background image in the picture materials, and set the channel of the background image to a transparent channel.

[0181] Perform content feature recognition on the foreground image, and judge whether the foreground image contains a person or an animal with facial features according to the content feature recognition result.

[0182] If so, recognize the facial features of the person or the animal to obtain the facial feature points of the person or the animal in the foreground image.

[0183] In this alternative embodiment, before making emoticons, it is necessary to preprocess the picture materials first. Specifically, first identify the foreground image and the background image in the picture materials. The foreground image and the background image are classifications of the content in the picture. Generally, the foreground image is the focus of the whole picture, while the background image exists as the background of the foreground image.

[0184] In this embodiment, it is assumed that the threshold range of the pixel gray value of the foreground image is 0 to 255. By comparing the pixel gray value in the picture materials with this threshold, the pixels falling within this range are called the foreground image, and the pixels not falling within this range are called the background image.

[0185] To avoid the appearance of the background image in the emoticons or interference with the emotional meaning expressed by the emoticons, it is also necessary to further set the channel of the background image to a transparent channel.

[0186] In this alternative embodiment, after setting the background of the picture material to be transparent, the foreground is further preprocessed. The specific processing method is to first perform content recognition on the foreground image to determine what content is in the foreground image. Only when the content in the foreground image is suitable for making emoticons can the subsequent processing continue. The content in the foreground image can be artificially selected in advance when setting the picture material, such as only selecting pictures with human or animal images as the picture material. Of course, it can also be that the user uses random pictures as the picture material. In this case, it is necessary to perform content feature recognition on the foreground image in the picture material. This embodiment does not limit the method of content feature recognition. For example, a pre-trained content feature recognition model can be used for recognition, such as a human feature recognition model, an animal feature recognition model, etc. Through the model, the content features in the image can be automatically recognized, such as having human face features or having the face features of a certain animal.

[0187] In this embodiment, by preprocessing the original material, the content expression form of the emoticon can be further improved and the content of the emoticon can be enriched, thereby enhancing the user participation and experience in making emoticons.

[0188] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the generation process of an embodiment of the static picture emoticon of the present invention. The fourth embodiment of the emoticon generation method in the embodiment of the present invention includes:

[0189] 401. Obtain the original material for generating emoticons;

[0190] 402. Determine the category of the emoticon to be generated and the emoticon generation parameters corresponding to the emoticon category; the emoticon category includes one or more of the following categories: audio emoticons, static picture emoticons, dynamic picture emoticons, video emoticons;

[0191] In this embodiment, the descriptions of the above steps 401 and 402 refer to the first embodiment and will not be elaborated in this embodiment. The following specifically describes the implementation process of generating a static picture emoticon with the obtained material being a picture material and the expression type being an emotion-based expression.

[0192] 403. When the determined category of the emoticon to be generated includes static picture emoticons, according to the facial feature point transformation matrix, transform the facial feature points of the person or animal in the foreground image to obtain an expression picture;

[0193] 404. Add a text watermark to the expression picture according to the picture text description input by the user to obtain a static picture emoticon; the facial feature point transformation matrix and the picture text description are the emoticon generation parameters corresponding to the static picture emoticons.

[0194] In this embodiment, the picture caption of the static emoji needs to be input by the user and used as an emoji generation parameter. Different static emojis can use the same picture caption or different picture captions, which is specifically determined according to actual application requirements.

[0195] In this embodiment, the facial feature point transformation matrix specifically refers to the matrix formed by the pixel displacement amounts of each facial feature point in the foreground image, which contains the adjustment methods of each facial feature point. By transforming the facial feature points of a person or an animal in the foreground image, an expression map with specific expression meanings can be formed. For example, opening the eyes wider, closing the eyes, opening the mouth wider, etc.

[0196] Since the emojis in this embodiment are for emotional expressions, pictures with the facial features of a person or an animal need to be selected to make emojis. If the content feature recognition result of the foreground image is that there is a person or an animal with facial features in the foreground image, then this picture material meets the prerequisite for making emotional emojis.

[0197] In this embodiment, emotional expressions are usually reflected by changes in facial features. For example, the mouth will open when happy, and the mouth will close and the eyebrows will frown when frustrated. Therefore, emoji generation parameters are preset, such as configuring a facial feature point transformation matrix for adjusting the positions of facial feature points in the foreground image, so as to make the person or animal in the foreground image have emotional expression features through feature point transformation.

[0198] In this embodiment, if the content feature recognition result is that there is no person or animal with facial features in the foreground image, then the emoji production is exited or a prompt message is popped up. If the content feature recognition result is that there is a person or an animal with facial features in the foreground image, then the facial feature points of the person or animal are continuously recognized, such as the facial feature points of the facial features, and then according to the facial feature point transformation matrix in the emoji generation parameters, the facial feature points of the person or animal in the current foreground image are transformed, so as to obtain an expression map with emotional features. Among them, the facial feature point transformation matrix contains the adjustment methods of each facial feature point, such as widening the eyes, closing the eyes, opening the mouth, frowning the eyebrows, etc. By transforming the facial feature points of a person or an animal in the foreground image, an emotional expression map with specific expression meanings is formed.

[0199] In this embodiment, to further enrich the meaning expressed by the emoji, a caption is added to the obtained expression map, specifically using the picture caption in the emoji generation parameter, and the caption is added to the expression map in the form of a watermark, so as to obtain a static picture emoji with emotional expression features and a caption. Users can use this emoji to express their emotions or convey relevant information, etc.

[0200] In this embodiment, any picture is specifically used as the material for the emoticon pack, which reduces the requirements for pictures in creating emoticon packs and increases the autonomy and flexibility of users in creating emoticon packs. By generating emoticon packs with one click, the efficiency of creating emoticon packs is greatly improved and the difficulty of creating emoticon packs is reduced.

[0201] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the generation process of an embodiment of the dynamic picture emoticon pack of the present invention. The fifth embodiment of the emoticon pack generation method in the embodiments of the present invention includes:

[0202] 501. Obtain the original material for generating the emoticon pack;

[0203] 502. Determine the category of the emoticon pack to be generated and the emoticon pack generation parameters corresponding to the category of the emoticon pack; the category of the emoticon pack includes one or more of the following categories: audio emoticon packs, static picture emoticon packs, dynamic picture emoticon packs, video emoticon packs;

[0204] In this embodiment, the descriptions of the above steps 501 and 502 refer to the first embodiment, and will not be elaborated in this embodiment. The following specifically describes the implementation process of generating a dynamic picture emoticon pack with the obtained material being a picture material and the expression type being an emotion-based expression.

[0205] 503. When the determined category of the emoticon pack to be generated includes a dynamic picture emoticon pack, sequentially transform the facial feature points of the person or animal in the foreground image according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, and obtain multiple frames of expression pictures with a time sequence;

[0206] 504. Synthesize the dynamic pictures of each frame of the expression pictures to obtain a dynamic picture emoticon pack; the timestamp and the facial feature point transformation matrix are the emoticon pack generation parameters corresponding to the dynamic picture emoticon pack.

[0207] The difference between this embodiment and the above fourth embodiment is that the emoticon pack generation parameters of this embodiment are configured with multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, and each timestamp corresponds to one frame of the dynamic picture.

[0208] In this embodiment, based on the time sequence corresponding to the timestamp, sequentially transform the facial feature points of the person or animal in the recognized foreground image according to the facial feature point transformation matrices respectively corresponding to each timestamp, so as to obtain multiple frames of emotion-based expression pictures with a time sequence. The expression features of the person or animal in each frame of the expression picture can be the same or different, and the parameter values in the configuration file are specifically adjusted according to actual needs.

[0209] After obtaining multiple frames of expression images, the dynamic image synthesis technology can be used to synthesize each expression image into a dynamic image expression pack, or it can also be synthesized into a video animation expression pack.

[0210] In one embodiment, the synthesized dynamic image expression pack can be further associated with audio for playback, so that while the dynamic image expression pack is being displayed, the associated audio is also played, thereby further enriching the expression content of the dynamic image expression pack.

[0211] In one embodiment, the audio frames and picture frames can be further synthesized into an audio-video to obtain a video expression pack with audio, so that while the video animation expression pack is being played, the corresponding audio content is also played, thereby further enriching the expression content of the video animation expression pack.

[0212] In this embodiment, multiple timestamps and multiple face feature point transformation matrices are introduced into the configuration file. Through multiple rounds of face feature transformation, one picture is transformed into multiple pictures, and finally, through picture synthesis or video synthesis, a dynamic image expression pack or a video animation expression pack is finally generated, further enriching the types of expression packs that can be made and enhancing the fun of users making expression packs.

[0213] Please refer to Figure 6 , Figure 6 which is a schematic diagram of the generation process of an embodiment of the video expression pack of the present invention. The sixth embodiment of the expression pack generation method in the embodiments of the present invention includes:

[0214] 601. Obtain the original materials for generating the expression pack;

[0215] 602. Determine the type of expression pack to be generated and the expression pack generation parameters corresponding to the type of expression pack; the type of expression pack includes one or more of the following types: audio expression pack, static picture expression pack, dynamic picture expression pack, video expression pack;

[0216] In this embodiment, the descriptions of the above steps 601 and 602 refer to the first embodiment, and will not be elaborated in this embodiment. The following specifically describes the implementation process of generating a video expression pack with the obtained material being a picture material and the expression type being an emotion expression.

[0217] 603. When the determined type of expression pack to be generated includes a video expression pack and the material includes a picture material, generate the expression audio of the person or animal according to the facial features of the person or animal recognized in the foreground image;

[0218] 604. Calculate the volume corresponding to each frame of the audio in the expression audio according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to the timestamps, and sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression images in chronological order;

[0219] The description of the above steps 603-604 refers to the second embodiment. The difference between this embodiment and the above fourth embodiment is that this embodiment can automatically generate the expression audio with emotions corresponding to the person or animal according to the recognized features of the person or animal in the foreground image. For example, if the recognized person feature is a young man in his twenties, then automatically generate the expression audio with the features of this person and emotions at the same time, such as the crying audio or the hearty laughter audio of the young man.

[0220] There is no limit to the way of generating the expression audio in this embodiment. It is preferred to pre-save the audio features of various people and animals in various emotions, and then generate the corresponding expression audio based on the audio features.

[0221] Optionally, in an embodiment, when the original material includes picture material and text material, perform audio synthesis on the text material to obtain the expression audio of the person or animal.

[0222] This optional embodiment can automatically generate the expression audio with corresponding emotions and expressing the content of the text material corresponding to the person or animal according to the text material. For example, if the recognized person feature is a young man in his twenties, the text material is "It's really delicious", and the emotion is happy, then the automatically generated expression audio is: The young man says happily "It's really delicious".

[0223] In this embodiment, when transforming the facial feature points of the person or animal in the foreground image, further calculate the difference between the uppermost point and the lowermost point of the lips of the person or animal in the facial feature point transformation matrix, and calculate the volume corresponding to each frame of the audio in the expression audio according to the calculated difference size of the upper and lower lip points (preferably taking the maximum value), where there is a positive correlation between the difference of the upper and lower lip points and the volume corresponding to each frame of the audio in the expression audio.

[0224] 605. Perform audio-visual synthesis on the expression audio and each frame of expression image to obtain a video file;

[0225] 606. Associate and save the video file with the preset video expression icon and the preset video display icon respectively to obtain a video emoji.

[0226] In this embodiment, it is preferred to match audio with animation to form a video emoji. The difference between the video emoji and the existing video emojis is that the former belongs to the vocal animation emoji, while the latter belongs to the silent animation emoji. The video emoji of this embodiment not only expresses expressions in the form of animation, but also conveys more information of the emoji in the form of audio, enhancing the richness of the information expressed by the emoji. At the same time, generating the video emoji with one key greatly improves the production efficiency of the emoji.

[0227] In this embodiment, to facilitate the user to intuitively operate the video emoji, a video emoji icon and a video display icon are introduced. Among them, the video emoji icon is an icon applied to represent the video emoji on the emoji sender side. The video emoji icon is associated with the video file of the video emoji. The user can send the video file associated with the video emoji icon by touching the video emoji icon, thus realizing the sending of the video emoji. The video display icon is an icon applied to represent the video emoji on the emoji receiver side. The video display icon is also associated with the video file of the video emoji. After the emoji receiver receives the video file and the video display icon, they are displayed and played on the receiver side, thus realizing the reception and playback of the video emoji.

[0228] The form of the video emoji icon and the video display icon in this embodiment is not limited. It can be a specific configured icon or a frame of the video file used as the icon.

[0229] Optionally, in one embodiment, the video emoji is sent in the following manner:

[0230] Detect whether there is a touch operation on the video emoji icon;

[0231] If so, obtain the video file associated with the video emoji icon and obtain the video display icon associated with the video file;

[0232] Send the video file and the video display icon.

[0233] Optionally, in one embodiment, the video emoji is received and played in the following manner:

[0234] Receive the video file and the video display icon;

[0235] Display the video display icon and the playback duration of the video file, and play the video file.

[0236] In this alternative embodiment, the video meme consists of three parts: a video file, a video emoticon icon, and a video display icon. The three are saved in association. Among them, the content of the video emoticon icon and the video display icon is not limited. It can be a configured icon or a frame of the video file as an icon. The video emoticon icon can distinguish different video files. Each video emoticon icon corresponds to an id, and each id corresponds to a video file. When the user clicks on the video emoticon icon, it triggers the sending of the video file and the video display icon. The receiving party receives the video display icon and the video file, displays the video display icon on the interface, and automatically plays the video file.

[0237] The method for generating a meme in the embodiment of the present invention has been described above. Next, the meme generation device in the embodiment of the present invention will be described. Please refer to Figure 7 In an embodiment, the meme generation device in the embodiment of the present invention includes:

[0238] An acquisition module 701, configured to acquire the original materials for generating the meme;

[0239] A determination module 702, configured to determine the type of the meme to be generated and the meme generation parameters corresponding to the type of the meme; the type of the meme includes one or more of the following types: audio memes, static picture memes, dynamic picture memes, video memes;

[0240] A generation module 703, configured to generate a meme by using the original materials and the meme generation parameters.

[0241] Optionally, in an embodiment, when the determined type of the meme to be generated includes audio memes, the generation module 703 is specifically configured to:

[0242] Obtain the text data corresponding to the original materials;

[0243] Generate audio data by using the musical score data and the text data; the musical score data is the meme generation parameter corresponding to the audio meme;

[0244] Generate an audio meme corresponding to the original materials by using the audio data.

[0245] Optionally, in an embodiment, the generation module 703 is further configured to:

[0246] When the original materials are audio materials, convert the audio materials into text to obtain text data;

[0247] When the original materials are picture materials, recognize the text information or semantic information in the picture materials to obtain text data.

[0248] Optionally, in an embodiment, when the original materials are audio materials, the generation module 703 is further configured to:

[0249] Identify the user pronunciation features corresponding to the audio material;

[0250] Generate audio data by using the music score data, text data, and user pronunciation features.

[0251] Optionally, in one embodiment, the generation module 703 is further configured to:

[0252] Generate an audio meme corresponding to the original material by using the audio data and the picture selected by the user; the audio meme is a meme carrying audio data in the picture.

[0253] Optionally, in one embodiment, when the determined meme category to be generated includes static picture memes, dynamic picture memes, or video memes, the generation module 703 includes:

[0254] A preprocessing unit, configured to preprocess the original material;

[0255] A processing unit, configured to process the preprocessed original material according to the meme generation parameters to obtain a meme.

[0256] Optionally, in one embodiment, when the original material includes picture material, the preprocessing unit is specifically configured to:

[0257] Identify the foreground image and the background image in the picture material, and set the channels of the background image to transparent channels;

[0258] Perform content feature recognition on the foreground image, and determine whether the foreground image contains a person or an animal with facial features according to the content feature recognition result;

[0259] If so, identify the facial features of the person or animal to obtain the facial feature points of the person or animal in the foreground image.

[0260] Optionally, in one embodiment, when the determined meme category to be generated includes static picture memes, the processing unit is specifically configured to:

[0261] Transform the facial feature points of the person or animal in the foreground image according to the facial feature point transformation matrix to obtain an expression picture;

[0262] Add a text watermark to the expression picture according to the picture text description to obtain a static picture meme; the facial feature point transformation matrix and the picture text description are the meme generation parameters corresponding to the static picture memes.

[0263] Optionally, in one embodiment, when the determined meme category to be generated includes dynamic picture memes, the processing unit is specifically configured to:

[0264] According to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression images with a time sequence;

[0265] Perform dynamic image synthesis on each frame of the expression image to obtain a dynamic image emoji; The timestamp and the facial feature point transformation matrix are the emoji generation parameters corresponding to the dynamic image type emoji.

[0266] Optionally, in an embodiment, when the determined emoji category to be generated includes video type emojis, the processing unit is specifically configured to:

[0267] Generate the expression audio of the person or animal according to the facial features recognized in the foreground image;

[0268] According to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, calculate the volume corresponding to each frame of audio in the expression audio, and sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression images with a time sequence;

[0269] Perform audio-visual synthesis on the expression audio and each frame of the expression image to obtain a video file;

[0270] Associate the video file with a preset video expression icon and a preset video display icon respectively to obtain a video emoji.

[0271] Optionally, in an embodiment, when the determined emoji category to be generated includes video type emojis, the processing unit is specifically configured to:

[0272] When the original material further includes text material, perform audio synthesis on the text material to obtain the expression audio of the person or animal;

[0273] According to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, calculate the volume corresponding to each frame of audio in the expression audio, and sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression images with a time sequence;

[0274] Perform audio-visual synthesis on the expression audio and each frame of the expression image to obtain a video file;

[0275] Associate the video file with a preset video expression icon and a preset video display icon respectively to obtain a video emoji.

[0276] Optionally, in an embodiment, the emoji generation device further includes:

[0277] A detection module, configured to detect whether there is a touch operation on the video expression icon;

[0278] An acquisition module, configured to obtain a video file associated with a video emoticon icon and a video display icon associated with the video file if there is a touch operation on the video emoticon icon;

[0279] A sending module, configured to send the video file and the video display icon.

[0280] Optionally, in an embodiment, the emoticon generating device further includes:

[0281] A receiving module, configured to receive the video file and the video display icon;

[0282] A playing module, configured to display the video display icon and the playing duration of the video file, and play the video file.

[0283] In this embodiment, the user only needs to provide the materials for making emoticons, and the system will automatically determine the category of the emoticons to be generated and the emoticon generation parameters corresponding to the emoticon category; finally, the original materials and the emoticon generation parameters are used to generate emoticons. The process of making emoticons is automatically completed without the need for user operation, thereby reducing the difficulty of making emoticons for users and improving the efficiency of making emoticons. In addition, since the materials are provided by the users themselves and the emoticon generation parameters can also be set by the users themselves, the flexibility and diversity of making emoticons are realized, and the enthusiasm and experience of users are improved.

[0284] Above Figure 7 The emoticon generating device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Next, the electronic device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0285] Figure 8 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 500 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 510 (for example, one or more processors) and a memory 520, and one or more storage media 530 (for example, one or more mass storage devices) for storing application programs 533 or data 532. Among them, the memory 520 and the storage media 530 may be transient storage or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the electronic device 500. Further, the processor 510 may be configured to communicate with the storage media 530 and execute a series of instruction operations in the storage media 530 on the electronic device 500.

[0286] The electronic device 500 may further include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 8 The illustrated structure of the electronic device does not constitute a limitation on the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0287] The present invention also provides an electronic device, which includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute the steps of the emoji generation method in the above-mentioned various embodiments.

[0288] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or it may also be a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the emoji generation method.

[0289] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0290] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0291] Above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for generating expression packs, characterized in that, the method includes: Obtaining the original materials for generating expression packs; Determining the category of the expression pack to be generated and the expression pack generation parameters corresponding to the category of the expression pack; the category of the expression pack includes one or more of the following categories: audio expression packs, static picture expression packs, dynamic picture expression packs, video expression packs; Using the original materials and the expression pack generation parameters to generate an expression pack; Among them, when the determined category of the expression pack to be generated includes a static picture expression pack, a dynamic picture expression pack or a video expression pack, the using the original materials and the expression pack generation parameters to generate an expression pack includes: Preprocessing the original materials; Processing the preprocessed original materials according to the expression pack generation parameters to obtain an expression pack; Among them, the preprocessing the original materials includes: When the original materials contain picture materials, identifying the foreground image and the background image in the picture materials, and setting the channels of the background image to transparent channels; Performing content feature recognition on the foreground image, and judging whether the foreground image contains a person or an animal with facial features according to the content feature recognition result; If so, identifying the facial features of the person or animal to obtain the facial feature points of the person or animal in the foreground image; Among them, when the determined category of the expression pack to be generated includes a video expression pack, the processing the preprocessed original materials according to the expression pack generation parameters to obtain an expression pack includes: Generating the expression audio of the person or animal according to the facial features of the person or animal identified in the foreground image; Calculating the volume corresponding to each frame of audio in the expression audio according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, and sequentially transforming the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression pictures with a time sequence; Performing audio-video synthesis on the expression audio and each frame of the expression pictures to obtain a video file; Associating and saving the video file with a preset video expression icon and a preset video display icon respectively to obtain a video expression pack.

2. The method for generating expression packs according to claim 1, characterized in that, when the determined category of the expression pack to be generated includes an audio expression pack, the using the original materials and the expression pack generation parameters to generate an expression pack includes: Obtaining the text data corresponding to the original materials; Using the music score data and the text data to generate audio data; the music score data is the expression pack generation parameter corresponding to the audio expression pack; Using the audio data to generate the audio expression pack corresponding to the original materials.

3. The method for generating expression packs according to claim 2, characterized in that, the obtaining the text data corresponding to the original materials includes: When the original materials are audio materials, converting the audio materials into text to obtain the text data; When the original materials are picture materials, identifying the text information or semantic information in the picture materials to obtain the text data.

4. The meme generation method according to claim 2, wherein, when the original material is audio material, the generating of audio data by using the sheet music data and the text data includes: identifying the user pronunciation characteristics corresponding to the audio material; generating audio data by using the sheet music data, the text data and the user pronunciation characteristics.

5. The meme generation method according to claim 2, wherein, the generating of the audio meme corresponding to the original material by using the audio data includes: generating the audio meme corresponding to the original material by using the audio data and the picture selected by the user; the audio meme is the meme with the audio data carried in the picture.

6. The meme generation method according to claim 1, wherein, when the determined meme category to be generated includes static picture memes, the processing of the preprocessed original material according to the meme generation parameters to obtain a meme includes: transforming the facial feature points of the person or animal in the foreground image according to the facial feature point transformation matrix to obtain an expression image; adding a text watermark to the expression image according to the picture text description input by the user to obtain a static picture meme; the facial feature point transformation matrix and the picture text description are the meme generation parameters corresponding to the static picture memes.

7. The meme generation method according to claim 1, wherein, when the determined meme category to be generated includes dynamic picture memes, the processing of the preprocessed original material according to the meme generation parameters to obtain a meme includes: successively transforming the facial feature points of the person or animal in the foreground image according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to the timestamps to obtain multiple frames of expression images with a time sequence; performing dynamic picture synthesis on each frame of the expression images to obtain a dynamic picture meme; the timestamps and the facial feature point transformation matrices are the meme generation parameters corresponding to the dynamic picture memes.

8. The meme generation method according to claim 1, wherein, when the determined meme category to be generated includes video memes, the processing of the preprocessed original material according to the meme generation parameters to obtain a meme includes: when the original material further includes text material, performing audio synthesis on the text material to obtain the expression audio of the person or animal; calculating the volume corresponding to each frame of audio in the expression audio according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to the timestamps, and successively transforming the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression images with a time sequence; performing audio-video synthesis on the expression audio and each frame of the expression images to obtain a video file; associating and saving the video file with a preset video expression icon and a preset video display icon respectively to obtain a video meme.

9. The meme generation method according to claim 1 or 8, wherein, After processing the preprocessed original material according to the meme generation parameters to obtain a meme, the following steps are further included: Detect whether there is a touch operation on the video emoticon icon; If so, obtain the video file associated with the video emoticon icon and obtain the video display icon associated with the video file; Send the video file and the video display icon.

10. The meme generation method according to claim 9, wherein, After processing the preprocessed original material according to the meme generation parameters to obtain a meme, the following steps are further included: Receive the video file and the video display icon; Display the video display icon and the playing duration of the video file, and play the video file.

11. A meme generation device, wherein, The device includes: An acquisition module for acquiring the original material for generating a meme; A determination module for determining the category of the meme to be generated and the meme generation parameters corresponding to the meme category; the meme category includes one or more of the following categories: audio memes, static picture memes, dynamic picture memes, video memes; A generation module for generating a meme by using the original material and the meme generation parameters; Wherein, when the determined meme category to be generated includes static picture memes, dynamic picture memes or video memes, the generation module includes: A preprocessing unit for preprocessing the original material; A processing unit for processing the preprocessed original material according to the meme generation parameters to obtain a meme; Wherein, the preprocessing unit is specifically used for: When the original material includes picture material, identify the foreground image and the background image in the picture material, and set the channel of the background image to a transparent channel; Perform content feature recognition on the foreground image, and judge whether the foreground image contains a person or an animal with facial features according to the content feature recognition result; If so, perform facial feature recognition on the person or animal to obtain the facial feature points of the person or animal in the foreground image; Wherein, when the determined meme category to be generated includes video memes, the processing unit is specifically used for: Generate the expression audio of the person or animal according to the facial features of the person or animal recognized in the foreground image; Calculate the volume corresponding to each frame of audio in the expression audio according to multiple timestamps and the facial feature point transformation matrices respectively corresponding to each timestamp, and sequentially transform the facial feature points of the person or animal in the foreground image to obtain multiple frames of expression pictures with a time sequence; Perform audio-video synthesis on the expression audio and each frame of the expression pictures to obtain a video file; Associate and save the video file with a preset video emoticon icon and a preset video display icon respectively to obtain a video meme.

12. An electronic device, wherein, The electronic device includes: a memory and at least one processor, and instructions are stored in the memory; The at least one processor invokes the instructions in the memory to cause the electronic device to execute the emoji generation method according to any one of claims 1-10.

13. A computer-readable storage medium, on which instructions are stored, wherein, when the instructions are executed by a processor, the emoji generation method according to any one of claims 1-10 is implemented.

Citation Information

Patent Citations

  • Dynamic expression generation method, device, computer readable storage medium and computer device

    CN109120866A

  • Voice expression display method and device and voice expression generation method and device

    CN112910752A