Information processing method and device, computer equipment and storage medium
By displaying emoji templates and generating emojis in real time during video recording, the problem of low emoji generation efficiency is solved, enabling the rapid generation of multiple emojis, simplifying user operations, and improving both fun and generation efficiency.
Patent Information
- Application Number
- CN202410558824.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-07
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies for generating emojis are inefficient, unable to quickly generate multiple emojis, and require complex, time-consuming, and labor-intensive user operations.
By displaying a set of emoji templates at fixed points in response to preset conditions during video recording, and displaying partial content of the video frame in the filled area, multiple emoji sets are automatically generated, reducing manual operation by the user.
It improves the efficiency of emoji generation, allowing users to generate multiple emojis without manually selecting materials, reducing the difficulty of operation, and increasing the fun and user participation in video recording.
Smart Images

Figure CN120915899A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and in particular, to an information processing method and device, computer equipment and storage medium. BACKGROUND
[0002] With the development of the Internet, various online communication applications have emerged. In the process of chatting using these applications, in order to express the content that the user wants to say more vividly and accurately express the emotions that the user needs to express, such as happiness, sadness, surprise, etc., the user usually uses different expressions to increase the interest and flexibility in the chatting process, so that the chat is more interesting and the user's feelings can be accurately conveyed.
[0003] At present, in order to meet the individual needs of users, users can customize the processing of materials they like through some additional tools, thereby generating a single expression. However, this method can only generate a single expression each time, and the efficiency of generating expressions is low. SUMMARY
[0004] Therefore, it is necessary to provide an information processing method and device, computer equipment and storage medium capable of quickly and automatically generating multiple expressions in view of the above technical problems.
[0005] In a first aspect, the present disclosure provides an information processing method. The method comprises:
[0006] starting video recording in response to a video recording trigger event;
[0007] displaying at least one expression template in an expression template group in response to reaching a freeze point in the video recording process that meets a preset condition, the expression template having a filling area;
[0008] displaying partial content of a video picture captured during video recording in the filling area;
[0009] obtaining a recorded video and an expression group containing multiple expressions in response to the end of video recording, each expression in the expression group containing partial content displayed in the filling area of the displayed expression template.
[0010] In a second aspect, the present disclosure also provides an information processing device. The device comprises:
[0011] a video recording module configured to start video recording in response to a video recording trigger event;
[0012] The template display module is configured to display at least one expression template in the expression template group in response to reaching a freeze frame point in the video recording process that meets a preset condition, the expression template having a fill area; and display partial content of a video picture captured during the video recording in the fill area.
[0013] The data acquisition module is configured to obtain a recorded video and an expression group containing a plurality of expressions in response to the end of the video recording, each expression in the expression group containing the partial content displayed in the fill area of the displayed expression template.
[0014] In a third aspect, the present disclosure also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps in any of the method embodiments described above when executing the computer program.
[0015] In a fourth aspect, the present disclosure also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in any of the method embodiments described above.
[0016] In a fifth aspect, the present disclosure also provides a computer program product. The computer program product includes a computer program, and the computer program is executed by a processor to implement the steps in any of the method embodiments described above.
[0017] The above information processing method, device, computer device, storage medium, and computer program product can display at least one expression template in the expression template group in response to reaching a freeze frame point in the video recording process that meets a preset condition during the video recording process, can display the expression template in real time during the video recording process, increase the interest, make the recorded video more lively and interesting, and can generate an expression according to the displayed expression template in real time during the video recording process. The partial content of the video picture captured during the video recording is displayed in the fill area, so that the user can accurately see the content in the fill area of the current expression template, thereby the user can adjust the partial content of the video picture captured, the generated expression can be more in line with the user's needs, and the user only needs to adjust the partial content of the video picture captured, without other operations, saving the user's time and effort. Since the expression group containing a plurality of expressions is generated during the video recording process, the user does not need to manually select materials and manually process, reducing the user operation process and the user operation difficulty, the expression group containing a plurality of expressions can be generated at one time, and the expression generation efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to make the technical solutions in the specific embodiments of the present disclosure or the prior art clearer, the accompanying drawings needed in the specific embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are some embodiments of the present disclosure, and other accompanying drawings can be obtained by those of ordinary skill in the art without any creative work on the premise of the accompanying drawings.
[0019] Figure 1 An application environment diagram of the information processing method in one embodiment;
[0020] Figure 2 A flow diagram of the information processing method in one embodiment;
[0021] Figure 3 A diagram of the expression template in one embodiment;
[0022] Figure 4 A diagram of displaying the expression template in the interface of video recording in one embodiment;
[0023] Figure 5 A diagram of adding the expression group to the expression library of the chat session in one embodiment;
[0024] Figure 6 A diagram of displaying at least one expression template in the expression template group when the video recording time length meets the time constraint condition in one embodiment;
[0025] Figure 7 A diagram of the relationship between the time point and the video in one embodiment;
[0026] Figure 8 A diagram of the action prompt information and the expression template displayed when reaching two key points in one embodiment;
[0027] Figure 9 A diagram of the rhythm of the background music in one embodiment;
[0028] Figure 10 A diagram of the relationship between the rhythm key point and the expression template in one embodiment;
[0029] Figure 11 A diagram of the relationship between the time point, the rhythm key point and the video in one embodiment;
[0030] Figure 12 A diagram of the relationship between the recording node corresponding to the recording content and the video in one embodiment;
[0031] Figure 13 A diagram of the generated expression in one embodiment;
[0032] Figure 14 a flowchart of a process for displaying an expression template in one embodiment;
[0033] Figure 15 a schematic diagram of a template generation interface in one embodiment;
[0034] Figure 16 a schematic diagram of testing an expression in one embodiment;
[0035] Figure 17 a flowchart of a process for information processing in another embodiment;
[0036] Figure 18 a schematic block diagram of an information processing apparatus in one embodiment;
[0037] Figure 19 a schematic diagram of an internal structure of a computer device in one embodiment;
[0038] Figure 20 a schematic diagram of an internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0039] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, further detailed description will be made to the present disclosure in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, and are not used to limit the present disclosure.
[0040] It should be noted that the terms "first", "second", and the like in the description of the specification and claims and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, device, product or equipment including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0041] In this paper, the term "and / or" is only a description of the relationship between the associated objects, which means that there can be three relationships. For example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this paper generally represents a "or" relationship between the front and rear associated objects.
[0042] The embodiments of the present disclosure provide an information processing method, which can be applied to, for example Figure 1The application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The data storage system can store recorded videos, expression template groups, and generated expression groups. In response to a video recording trigger event, the terminal 102 starts video recording. During video recording, in response to reaching a preset condition that meets the still point in the video recording process, the terminal 102 displays at least one expression template in the expression template group, and the expression template has a fill area. In the fill area, the terminal 102 displays the local content of the video screen captured during video recording. In response to the end of video recording, the terminal 102 obtains a recorded video and an expression group containing multiple expressions. Each expression in the expression group contains at least one expression template displayed during video recording, and the local content displayed in the fill area of the displayed expression template. The expression group is used to select an expression from the expression group in response to a user operation in a conversation and send it in the conversation. The recorded video is used for playback browsing. Usually, the recorded video and the expression group are generated by the server 104. Among them, the terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The server 104 can be implemented by a standalone server or a server cluster composed of multiple servers. It should be noted that the above application environment is only an example. In some embodiments, this scheme can also be applied to the server or the terminal alone, and can also be applied to a system including the terminal and the server, and realized through the interaction of the terminal and the server.
[0043] In one embodiment, as Figure 2 shown, an information processing method is provided. This method is applied to the terminal 102 in Figure 1 for example, including the following steps:
[0044] S202, in response to a video recording trigger event, start video recording.
[0045] Among them, the video recording trigger event refers to an event that triggers the start or stop of video recording. These events can be triggered by manually operating the recording device, or by an automated program or sensor. For example, time trigger, action trigger, sound trigger, environment trigger, etc.
[0046] Specifically, the recording of a video can be started by manually operating a button or a touch screen of the terminal. The user can control the start and stop of the video recording by pressing a button on the button or the touch screen. A timing time can also be set in the terminal, and the terminal starts the video recording at a specific time point. In addition, some video recording terminals can also trigger recording by sound. When the terminal detects a sound of a certain intensity or frequency, it will start video recording. Some video recording terminals can also trigger recording by detecting changes in the temperature or light of the environment. When the temperature or light of the environment reaches a set threshold, the terminal will start video recording.
[0047] S204, in response to reaching a freeze frame point in the video recording process that meets a preset condition, displaying at least one expression template in an expression template group, the expression template having a filling area.
[0048] The preset condition can be determined according to different application scenarios, and the freeze frame point can be determined according to the preset condition. For example, the preset condition can be a condition for restricting the recording duration in the video recording process, for example, when the recording duration in the video recording process reaches 5s, the 5s time point can be determined as the freeze frame point. When there is background music in the video recording process, the preset condition can also be a condition for restricting the background music, for example, the rhythm of a certain rhythm point of the background music is relatively strong, and the rhythm point corresponding to the relatively strong rhythm can be the freeze frame point. The preset condition can also restrict the video shooting content, and the video content can be, for example, facial expressions, body movements, etc. When the expression or body movement in the video content in the video recording process matches the preset expression and body movement, the node of the recorded video can be the freeze frame point.
[0049] The expression template group can include a plurality of different expression templates. The expression template is a template or model for creating and editing expressions. In some embodiments of the present disclosure, each expression template has a filling area. The filling area can generally be an area that needs to fill in the local content in the video recording process. The local content can include facial expressions, bodies, backgrounds, and the like. As shown in Figure 3 The filling area can fill in the facial expression.
[0050] Specifically, the terminal displays at least one expression template in the expression template group in response to reaching a freeze frame point that meets a preset condition during the video recording process.
[0051] In some exemplary embodiments, as Figure 4As shown, during the terminal video recording process, the terminal usually displays a video recording interface. Therefore, at least one expression template in the expression template group can be displayed in the video recording interface, can be displayed in the center of the video recording interface, or can be displayed in other positions in the video recording interface. In some embodiments of the present disclosure, the display position of at least one expression template in the expression template group is not limited as long as the at least one expression template in the expression template group can be displayed completely.
[0052] S206, in the filling area, display the partial content of the video picture captured by the video recording.
[0053] The partial content can be a certain specific area or part in the video picture.
[0054] Specifically, after displaying at least one expression template in the expression template group, in order to enable the user to see the style of the currently generated expression, the partial content of the video picture captured by the video recording can be displayed in the filling area, so that the user can adjust the partial content of the video picture captured by the video recording in a targeted manner, thereby generating an expression more in line with the user's needs. The terminal can generate an expression based on the partial content displayed in the filling area and the expression template.
[0055] In some exemplary embodiments, the partial content displayed in the filling area can be partial content of the material type specified by the filling area. For example, the filling area needs to display a human face, and the partial content can be a human face. The filling area needs to display a body part, such as a hand, and the partial content can be a hand. In this way, when the partial content is displayed in the filling area, the user can determine the general style of the generated expression, thereby making targeted adjustments.
[0056] S208, in response to the end of the video recording, obtain a recorded video and an expression group containing a plurality of expressions, each expression in the expression group containing at least one expression template displayed during the video recording process and the partial content displayed in the filling area of the displayed expression template; the expression group is used to select an expression from the expression group in response to a user operation in a conversation and send in the conversation, and the recorded video is used for playback browsing.
[0057] In particular, in response to the end of the terminal video recording, since the terminal has displayed at least one expression template in the video recording process and displayed the partial content in the filling area in the expression template, the terminal can generate an expression according to the displayed expression template and the partial content displayed in the filling area in the video recording process. In addition, a plurality of different expression templates are displayed in the video recording process, so the server or the terminal can generate a plurality of different expressions according to the plurality of different expression templates and the partial content in the video recording, and also generate a recording video according to the video pictures captured in the recording process. The generated expression group can select an expression from the expression group in the conversation in response to a user operation and send it in the conversation. The recording video is used for playback browsing or published to a short video client.
[0058] For example, a user starts a video recording in a certain short video client using a terminal in response to a video recording trigger event. During the video recording, in response to the server detecting a freeze point that meets a preset condition in the video recording process, the server displays at least one expression template in the expression template group in the short video client of the terminal. The expression template has a filling area. The terminal displays the partial content of the video pictures captured during the video recording in the filling area. During the video recording, the server generates an expression according to the content in the filling area and the displayed expression template. In response to the video recording result, the terminal and the server obtain a recording video and an expression group containing a plurality of expressions. As shown in Figure 5 When the short video client also has a chat function, the expression group can be added to the expression library of the chat conversation, and the user can select an expression from the expression group in the expression library in the conversation in response to a user operation and send it in the conversation when using the chat function of the short video client in the future. The user can also publish the recording video on the short video client. Subsequently, the recording video published in the short video client is played back for browsing. In addition, when the short video client does not have a chat function, the generated expression group can be stored locally. Subsequently, when the user chats through other chat clients, the user can select an expression from the expression group stored locally in the conversation in response to a user operation and send it in the conversation.
[0059] In the above information processing method, at least one expression template in the expression template group is displayed in response to reaching a freeze frame point in the video recording process that meets a preset condition. The expression template can be displayed in real time during the video recording process, increasing the interest and making the recorded video more lively and interesting. The expression can also be generated in real time during the video recording process according to the displayed expression template. The partial content of the video picture captured during the video recording is displayed in the filling area, so that the user can accurately see the content in the filling area of the current expression template, thereby enabling the user to adjust the partial content of the video picture captured in a targeted manner, so that the generated expression is more in line with the user's needs. The user only needs to adjust the partial content of the video picture captured in a targeted manner, without the need for other operations, thereby saving the user's time and effort. Since the expression group containing multiple expressions is generated during the video recording process, the user does not need to manually select the material and manually process, thereby reducing the user's operation and processing process and reducing the user's operation difficulty. The expression group containing multiple expressions can be generated at one time, thereby improving the efficiency of expression generation.
[0060] In one embodiment, the preset condition includes a time constraint condition for a recording duration of the video recording, and the freeze frame point is a time point at which the recording duration meets the time constraint condition. The displaying at least one expression template in the expression template group in response to reaching the freeze frame point in the video recording process that meets the preset condition includes:
[0061] In response to reaching the time point at which the recording duration meets the time constraint condition, at least one expression template in the expression template group is displayed in the interface of the video recording.
[0062] The time constraint condition can be one or more preset time durations or a preset time interval. For example, the preset time durations are 5s, 10s, and 15s, and the freeze frame points can be the time points of 5s, 10s, and 15s after the start of the video recording.
[0063] When the video recording process reaches the time point at which the video recording duration meets the preset one or more time durations, at least one expression template in the expression template group is displayed in the interface of the video recording each time a time point at which the video recording duration meets the preset time point is reached.
[0064] For example, as Figure 6As shown, taking an example of displaying an expression template, when the time constraint condition is a preset time length, for example, 3s, 6s and 9s, during the video recording process, when the recording time length reaches 3s, it can be determined that a key frame point is reached, and the bounce expression template is displayed in the video recording interface at 3s. When the recording time length reaches 6s, it can be determined that another key frame point is reached, and the cheer expression template is displayed in the video recording interface at 6s. When the recording time length reaches 9s, it can be determined that the last key frame point is reached, and the NO.1 expression template is displayed in the video recording interface at 9s. When the time constraint condition is a preset time interval, for example, the time interval is 3s, when the first 3s time interval is reached, it can be determined that the first key frame point is reached. When the first key frame point is reached, the bounce expression template is displayed in the video recording interface. When the second 3s time interval is reached, it can be determined that the second key frame point is reached. When the second key frame point is reached, the cheer expression template is displayed in the video recording interface. When the third 3s time interval is reached, it can be determined that the third key frame point is reached. When the third key frame point is reached, the NO.1 expression template is displayed in the video recording interface.
[0065] Before displaying the next set of expression templates, the at least one expression template displayed in the video recording interface is hidden, wherein the expression template displayed in the video recording interface is different each time.
[0066] Specifically, in general, the expression template displayed in the video recording interface is different each time, which can ensure that the generated expression is different. In addition, before displaying the next set of expression templates, the at least one expression template displayed in the video recording interface is generally hidden, which can ensure that the at least one expression template displayed in each set does not interfere with each other, i.e., only one set of expression templates is displayed in the video recording interface at the same time.
[0067] For example, the time for hiding the displayed expression template can be set, and the displayed expression template can be hidden after 1s or 2s. The time for displaying the next set of expression templates can also be set, and if the time for displaying the next set of expression templates is 5s, the at least one displayed expression template can be hidden between 2s and 5s. It should be noted that this is only an example of how to hide the displayed expression template, and those skilled in the art can set more ways to hide the expression template. For example, when it is determined that the expression indicated by a certain displayed expression template is successfully generated, the expression template of the successfully generated expression can be hidden or switched to the next set of expression templates to be displayed.
[0068] In an exemplary embodiment, before video recording, the user can specify the display timing of the expression template, and can choose to display the expression template and generate the expression based on the time point of the recording time length meeting the time constraint condition. For example,Figure 7 As shown, the maximum playing time range can be specified first, and then a plurality of time points can be selected and moved, and the plurality of time points obtained by the selection and movement can be the time constraint condition. During the video recording process, the end user displays the expression template according to the plurality of time points obtained by the selection and movement within the maximum playing time range, so as to generate the expression according to the displayed expression template. In addition, the number of generated expressions is not more than the number of time points within the range. If the selected time points are a1, a2, a3 and a4, then the A expression template can be displayed when the recording duration of the video recording process reaches a1, the B expression template can be displayed when the recording duration reaches a2, the C expression template can be displayed when the recording duration reaches a3, and the D expression template can be displayed when the recording duration reaches a4. In addition, after the expression template is displayed, the local content of the video picture during the video recording process can not conform to the material type specified by the filling area, for example, the facial expression of the local content is an angry expression, and the facial expression of the material type specified by the filling area is a happy expression. In this case, although the expression template is displayed, the expression generated according to the expression template can not be successful. Therefore, although four expression templates are displayed, the number of generated expressions will not be more than four.
[0069] In this embodiment, the expression template is displayed in response to reaching the time point that makes the recording duration conform to the time constraint condition. By displaying different expression templates during the recording process, the user can have a more rich and interesting experience, and the user's participation and fun can be increased. Before displaying the next group of expression templates, at least one expression template displayed in the interface of the video recording is hidden, and the expression templates displayed in the interface of the video recording are different each time. By constantly switching different expression templates, the video recording can be more interesting and creative, the user can have more choices and possibilities, and the attractiveness of the video recording can be improved. In addition, the time points are determined by using the time constraint condition, the timing of the expression template display can be accurately determined, and the timing of the expression template display can be flexibly adjusted by using the time constraint condition, thereby improving the operability.
[0070] In one embodiment, the display of at least one expression template in the expression template group in the interface of the video recording includes:
[0071] The action prompt information is displayed in the interface of the video recording, and at least one expression template in the expression template group indicated by the action prompt information is displayed. The action prompt information is used to prompt the limb action during the video recording process. For example, facial expression action, such as making a happy expression, making a cheering expression, making an angry expression, etc.; body action, such as hand action, arm action, etc.
[0072] Specifically, when a time point is reached such that the recording duration meets the time constraint condition, action prompt information is displayed in the interface of the video recording, and at least one expression template in the expression template group indicated by the action prompt information is displayed, the filling area in the expression template is usually filled with the partial content of the local video picture associated with the action prompt information. For example, the action prompt information is to make a happy face, and the filling area in at least one expression template in the expression template group indicated by the action prompt information is usually also displayed with the partial content of the happy face part in the video picture.
[0073] In some exemplary embodiments, as shown in Figure 8 For example, as shown in the embodiment, two key points are reached, and an expression template is displayed at each key point. When the first key point is reached, the interface of the video recording can display action prompt information, which can be to make a surprised face. The interface of the video recording can also display a bounce expression template indicated by the action prompt information of making a surprised face, and the filling area in the bounce expression template can display the content of the video picture containing the surprised facial expression. When the second key point is reached, another action prompt information can be displayed in the interface of the video recording, which can be to make a happy face. The interface of the video recording can also display a cheer expression template indicated by the action prompt information of making a happy face, and the filling area in the cheer expression template can display the content of the video picture containing the happy facial expression.
[0074] In this embodiment, by displaying the action prompt information and displaying at least one expression template in the expression template group indicated by the action prompt information, the user can accurately make the corresponding body action according to the action prompt information, and the success rate of expression generation can be improved.
[0075] In one embodiment, the displaying, in the filling area, the partial content of the video picture captured by the video recording, comprises:
[0076] In response to the body action contained in the partial content of the video picture captured by the video recording matching the body action prompted by the action prompt information, the partial content of the video picture captured by the video recording is displayed in the filling area.
[0077] The matching can be matching of types of the body action, for example, the body action is an A action of hand swinging, and the body action contained in the partial content of the video picture is a B action of hand swinging, although the A action and the B action are different, but both are actions of hand swinging, and at this time, the matching can be considered. The matching can also be matching of the body action, for example, the body action is a happy expression of a face, and the expression of the face contained in the partial content of the video picture is a happy expression, and at this time, the matching can be considered. According to different application scenarios, a person skilled in the art can flexibly select the matching condition.
[0078] Specifically, in response to the body action contained in the partial content of the picture photographed by the video recording matching the body action of the action prompt information, the partial content of the video picture photographed by the video recording matching the body action of the action prompt information is displayed in the filling area.
[0079] In some exemplary embodiments, in some application scenarios, the filling area only needs to display a certain type of body action. For example, the filling area only needs to display the expression of the face, and does not pay attention to the expression type of the expression of the face. If the body action prompted by the action prompt information is to swing a happy expression, the expression of the face contained in the partial content of the video picture photographed by the video recording is a sad expression, although the expressions are different, but the partial content contains the expression of the face, and it can also be considered as matching, and at this time, the sad expression can be displayed in the filling area. In other application scenarios, the filling area needs to display the body action that meets the action prompt information, for example, the action prompt information is to prompt the hand to swing a “v” action, and the hand action contained in the partial content of the video picture photographed by the video recording is an “o” action. Since this application scenario needs to display the body action that meets the action prompt information, it can be determined that the body action contained in the partial content of the picture photographed by the video recording does not match the body action of the action prompt information, and the filling area can not display the partial content.
[0080] In this embodiment, by displaying the related video content in the filling area according to whether the action prompt information matches, the user can more intuitively understand whether his action is correct, thereby improving the experience and participation of the user in generating the expression in the video recording process, and also increasing the interactivity between the user and the application program, so that the user is more actively involved in the video recording process, and the enthusiasm of the user to record the video is improved.
[0081] In one embodiment, the body action includes a facial expression; the method further comprises:
[0082] Identifying facial feature points of the facial expression contained in the partial content of the video picture photographed by the video recording.
[0083] The facial feature points of the facial expression refer to key points in a human face image or video that are used to describe and identify facial expressions. These feature points are usually some obvious landmark positions on the face, such as the outlines of the eyes, the positions of the eyebrows, the tips of the nose, the outlines of the lips, etc. Common facial feature points include: eyes: upper and lower edges of the eye socket, inner and outer corners of the eye, etc. Eyebrows: brow, eyebrow peak, and eyebrow tail, etc. Nose: tip of the nose, nose wings, nose bridge, etc. Lips: outline of the lips, corners of the lips, etc. Chin: outline of the chin and center point of the chin, etc.
[0084] Specifically, the facial feature points of the facial expression contained in the local content of the video picture captured by the video recording can be identified using computer vision and image processing techniques. A facial key point detection model can be trained using deep learning techniques such as convolutional neural networks (CNN). By inputting the corresponding image, the model can output the position coordinates of each key point, thereby realizing the identification of facial feature points. It is also possible to first perform face detection and then use specific algorithms to locate the feature points of the face within the detected face region, such as shape model-based methods or regression-based methods. It should be noted that in actual applications, those skilled in the art can select the most suitable method to identify the facial feature points of the face according to the specific situation. The way to identify the facial feature points is not limited in some embodiments of the present disclosure.
[0085] Calculate the confidence of the facial feature points;
[0086] According to the confidence, the facial expression is classified and identified to determine the first expression type of the facial expression.
[0087] The confidence of the facial feature points refers to the degree of confidence or reliability of the system in the accuracy of each detected facial feature point. In the task of facial key point detection, the system usually assigns a confidence score to each detected facial feature point, indicating the reliability of the feature point. The confidence score is usually a value between 0 and 1, indicating the probability or confidence that the feature point is correctly detected. A higher confidence score indicates that the system is more confident about the position of the feature point, and a lower confidence score may indicate that the system is not sure about the position of the feature point or has some error.
[0088] Specifically, the confidence of the facial feature points for each expression category is calculated, and the facial expression is classified and identified according to the confidence to determine the category to which the facial expression belongs. Taking the happy expression and the sad expression as an example, if the confidence of the facial feature points corresponding to the happy expression is 0.5 and the confidence of the facial feature points for the sad expression is 0.9, then the first expression type of the facial expression can be determined as the sad expression.
[0089] In some exemplary embodiments, the facial feature points can be classified and recognized using, for example, a method based on an SVM classifier, so as to determine the first expression type of the facial expression.
[0090] The second expression type of the facial expression indicated by the action prompt information is determined.
[0091] When the first expression type and the second expression type are the same, the facial expression contained in the partial content of the video picture shot by the video recording is filled into the filling area of the displayed at least one expression template, so as to generate an expression.
[0092] Specifically, the second expression type of the facial expression indicated by the action prompt information is determined. When the second expression type and the first expression type are the same, the facial expression contained in the partial content of the video picture shot by the video recording is filled into the filling area of the displayed at least one expression template, and after the filling is completed, at least one expression is generated according to the at least one expression template obtained after the filling.
[0093] In the embodiment, by recognizing the first expression type of the facial expression in the partial content of the video picture, when the first expression type and the second expression type of the facial expression indicated by the action prompt information are the same, an expression is generated, which can ensure that when a happy or positive expression needs to be generated, no sad or unhappy expression is displayed in the filling area, and no sad or unhappy expression is generated, so as to ensure the accuracy of the generated expression.
[0094] In one embodiment, during the video recording, some video templates can also be used for video recording, the video templates can be provided with background music, or a certain music can be set as the background music during the video recording. Therefore, the background music can also be played during the video recording. The preset condition can also include a rhythm constraint condition for the background music played during the video recording. The rhythm constraint condition can generally ensure the condition that the rhythm of the music is strong. The freeze point is a rhythm card point that makes the rhythm of the background music meet the rhythm constraint condition. As shown in FIG. 8, it is a rhythm diagram of the background music, in which the rhythm at points A, B and C is strong. The points A, B and C can be the freeze points (rhythm card points) that meet the rhythm constraint condition. Figure 9
[0095] In some exemplary embodiments, the intensity and variation of the musical rhythm can be determined using a music rhythm tonality extraction algorithm, such as a rhythm tracking method based on dynamic programming, or by providing a visual rhythm display of background music through music software or rhythm tools. This allows for the extraction of rhythm points, and the determination of rhythm stops based on these rhythm points and rhythm constraints. For example, if 10 rhythm points are identified, including A1, A2...A10, where the rhythm corresponding to rhythm point A2 is the strongest, then A2 can be considered a rhythm stop (freeze point).
[0096] The response to reaching a freeze point that meets preset conditions during video recording, displaying at least one expression template from the expression template group, includes:
[0097] In response to reaching a rhythmic beat that makes the rhythm of the background music conform to the rhythmic constraints, at least one emoji template from the emoji template group is displayed in the video recording interface.
[0098] Among them, rhythmic constraints are the conditions that determine the strength of the music.
[0099] Specifically, when a rhythmic beat is reached during video recording that makes the background music's rhythm conform to the rhythmic constraints, at least one emoji template from the emoji template group is displayed in the video recording interface for each rhythmic beat that conforms to the rhythmic constraints.
[0100] For example, such as Figure 10 As shown, continuing with the example of displaying one set of emoji templates at a time, with each set containing only one emoji template, we can first obtain the background music and determine its rhythm. Then, based on the rhythm and constraints of the background music, we determine the rhythm points. For example, the determined rhythm points are C1, C2, and C3. During video recording, the background music will also play simultaneously. When the rhythm of the background music reaches rhythm point C1, the "bouncing" emoji template will be displayed on the video recording interface. When the rhythm of the background music reaches rhythm point C2, the "cheering" emoji template will be displayed on the video recording interface. When the rhythm of the background music reaches rhythm point C3, the "NO.1" emoji template will be displayed on the video recording interface.
[0101] Before displaying the next set of emoji templates, at least one emoji template is hidden in the video recording interface, wherein the emoji template displayed in the video recording interface is different each time.
[0102] Specifically, regarding how to hide the displayed emoji templates, please refer to the above embodiments, which will not be repeated here.
[0103] In one exemplary embodiment, before recording a video, if the user uses a video template (which includes background music) or has already determined the background music to be used during the recording process, the timing for displaying the emoji template can be determined based on the background music. Typically, the background music used for recording a video is known, and therefore its rhythmic spectrum can also be predetermined. Figure 11 As shown, after determining the rhythm spectrum of the background music, rhythmic timing points are determined using the rhythm spectrum and rhythmic constraints. These timing points are C1, C2, C3, and C4. Emoji template A can be displayed at rhythmic timing point C1, emoji template B at C2, emoji template C at C3, and emoji template D at C4. Furthermore, corresponding time points can be determined based on the rhythmic timing points: time point T1 is determined based on rhythmic timing point C1, time point T2 based on rhythmic timing point C2, time point T3 based on rhythmic timing point C3, and time point T4 based on rhythmic timing point C4. During video recording, when the recording duration or background music playback duration reaches time point T1, emoji template A is displayed; when it reaches time point T2, emoji template B is displayed; when it reaches time point T3, emoji template C is displayed; and when it reaches time point T4, emoji template D is displayed.
[0104] In this embodiment, in response to reaching a rhythmic beat that makes the rhythm of the background music conform to the rhythmic constraints, at least one emoticon template from the emoticon template group is displayed in the video recording interface. This allows the display of the emoticon template to match the rhythm of the background music, thereby increasing the interactivity between video recording and emoticon template display.
[0105] In one embodiment, the preset conditions include: content constraints on the recorded content of the video recording. Content constraints are typically conditions that ensure the recorded video content is the same as the preset content. The recorded video content may include: faces, bodies, backgrounds, etc. A freeze point is a recording node where the body movements or background information in the recorded content meet the content constraints. For example, if the currently recorded content contains a red background, the expression template corresponding to the generated expression also needs to use a red background. Therefore, the preset content can be a red background, and the recording node currently recording to the red background can be a freeze point. The step of displaying at least one expression template from the expression template group in response to reaching a freeze point that meets the preset conditions during video recording includes:
[0106] In response to reaching a recording node that makes the body movements or background information in the recorded content conform to the content constraints, at least one expression template from the expression template group is displayed in the video recording interface.
[0107] Specifically, when a recording node is reached in the video recording process, in which the body movement or background information contained in the recording content meets the content constraint condition, one of the expression templates in the expression template group is displayed in the interface of the video recording each time a recording node meeting the content constraint condition is reached.
[0108] For example, before the video recording, several facial expressions, body movements and background colors can be preset. The timing of displaying the expression template is determined according to the preset facial expression, body movement and background color, so as to determine the timing of generating the expression. During the video recording process, the recording content in the video picture obtained by the video recording can be acquired. As shown in FIG. 1, the recording content is detected, and when the recording content matches any one of the preset facial expression, body movement and background color, the recording node corresponding to the recording content is determined, for example, the recording content is determined to match the preset facial expression at B1, B2, B3 and B4, and different expression templates in the expression template group are displayed at the B1, B2, B3 and B4 respectively. Figure 12
[0109] Before displaying the next group of expression templates, the at least one expression template displayed in the interface of the video recording is hidden, wherein the expression template displayed in the interface of the video recording is different each time.
[0110] Specifically, how to hide the displayed expression template can be referred to the above-mentioned embodiments, which will not be repeated here.
[0111] In this embodiment, when a recording node is reached in which the body movement or background information in the recording content meets the content constraint condition, the expression template is displayed, which can increase the interactivity with the user during the video recording process. The user can control the display of the expression template according to different recording content, which can add interest and interactivity to the video recording. The user can also dynamically select the display timing of the expression template, so as to control the expression generation timing.
[0112] In one embodiment, the content constraint condition includes a preset expression. The body movement includes a facial expression. The freeze point is a recording node in which the facial expression matches the preset expression. Here, the facial expression matches the preset expression can be that the expression type of the facial expression matches. For example, the facial expression is a happy expression, and the preset expression is also a happy expression, in which case the facial expression matches the preset expression.
[0113] The displaying at least one expression template in the expression template group in response to reaching the freeze point meeting the preset condition in the video recording process includes:
[0114] In response to reaching the recording node that makes the facial expression in the recording content match the preset expression, at least one expression template in the expression template group indicated by the preset expression is displayed in the interface of the video recording.
[0115] The at least one expression template in the expression template group indicated by the preset expression can be an expression template associated with the preset expression. For example, when the preset expression is a happy expression, the expression template can be an expression template expressing a happy emotion, and when the preset expression is an angry expression, the expression template can be an expression template expressing an angry emotion.
[0116] Specifically, during the video recording process, in response to reaching the recording node that makes the facial expression in the recording content match the preset expression, the expression template indicated by the preset expression is determined, and the expression template indicated by the preset expression is displayed in the interface of the video recording.
[0117] In some exemplary embodiments, taking the preset expression as a surprised expression as an example, when the video recording process contains a facial expression in the recording content, and the facial expression is also a surprised expression, it can be determined that the time when the surprised facial expression appears is the recording node that makes the facial expression in the recording content match the preset expression, and the time when the expression template indicated by the preset expression is displayed in the interface of the video recording.
[0118] In this embodiment, when the recording node that makes the facial expression in the recording content match the preset expression is reached, the expression template indicated by the preset expression is displayed, which can accurately display the associated expression template according to the emotion of the current facial expression, so that the expression generated according to the displayed expression template is more consistent with the emotion of the user during the current video recording, and the displayed expression template can be combined with the emotion generated in the recording content, so that the generated expression template is more vivid.
[0119] In one embodiment, the method further comprises:
[0120] In response to a change in the facial expression in the recording content, the at least one expression template indicated by the displayed preset expression is switched to an expression template in the expression template group that has not been displayed.
[0121] Specifically, when the at least one expression template indicated by the preset expression is displayed, the facial expression in the recording content changes, at this time, the displayed expression template also needs to change accordingly, and the at least one expression template indicated by the displayed preset expression needs to be switched to the expression template in the expression template group that has not been displayed. In some scenarios, if there is an expression template associated with the changed facial expression in the expression template group, the switching can switch the expression template to the expression template associated with the changed facial expression. For example, if the facial expression changes from a surprised expression to a happy expression, the changed facial expression can be identified at this time. If there is an expression template associated with the happy expression in the expression template group, the expression template indicated by the surprised expression can be switched to the expression template associated with the happy expression. If there is no expression template associated with the changed facial expression, at least one expression template in the expression template group that has not been displayed can be switched.
[0122] In this embodiment, when the facial expression in the recording content changes, the displayed expression template can be changed in time, so that the expression template can be accurately switched. The user can adjust the facial expression in the recording content, so as to control the displayed expression template, and further control the generated expression.
[0123] In one embodiment, the method further comprises: switching at least one expression template in the expression template group displayed in the video recording interface to an expression template in the expression template group that has not been displayed, after the expression template is displayed for a preset duration, or in response to an expression template switching instruction.
[0124] The expression template switching instruction is an instruction indicating that the expression indicated by the displayed at least one expression template is successful. The preset duration can be a maximum expression switching waiting time.
[0125] Specifically, after the expression template is displayed, the expression corresponding to the expression template is generated successfully, and the server or the terminal outputs an expression switching instruction. In response to the expression switching instruction, at least one expression template in the expression template group displayed in the video recording interface can be switched to an expression template that has not been displayed. For example, the A expression template is currently displayed, and in response to the expression switching instruction, it is determined that the expression indicated by the A expression template is generated successfully, and the displayed A expression template can be switched to the B expression template in the expression template group. Alternatively, when the expression template is displayed for a preset duration, for example, has been displayed for 3s, at least one expression template in the expression template group displayed in the video recording interface can be switched to an expression template that has not been displayed.
[0126] In addition, it should be noted that the switching of the expression template mentioned in some embodiments of the present disclosure is taken as an example of switching an A expression template to a B expression template. The process is generally as follows: first, hide the displayed A expression template, and then display the B expression template after the A expression template is hidden, thereby completing the switching of the A expression template and the B expression template.
[0127] In the present embodiment, after the expression template has been displayed for a preset time length, it can be determined that the expression template has been displayed for a long time. In response to an expression switching instruction, it can be determined that the expression indicated by the currently displayed expression template has been successfully generated. In order to ensure the efficiency of expression generation, at least one expression template in the expression template group displayed in the interface for video recording can be switched to an expression template in the expression template group that has not been displayed.
[0128] In one embodiment, the display of the partial content of the video picture captured by the video recording in the filling area comprises:
[0129] Determining the material type specified by the filling area in the expression template.
[0130] The material type can be the material type of the partial content that needs to be displayed in the filling area, such as face material type, background material type, body material type, etc.
[0131] Specifically, the expression template usually specifies the material type of the content displayed in the filling area during the generation process. Therefore, the material type specified by the pre-determined filling area in the expression template can be determined.
[0132] Obtaining the partial content in the video picture captured by the video recording that matches the material type specified by the filling area.
[0133] Specifically, during the video recording process, since the partial content of the captured video picture needs to be displayed in the filling area, the material type of the partial content also needs to match the material type specified by the filling area. For example, the partial content of the captured video picture is a hand, and the material type specified by the filling area is also a hand. At this time, it can be determined that the material type of the partial content matches the material type specified by the filling area. That is, the hand is the matching partial content.
[0134] Displaying the matching partial content in the filling area
[0135] Specifically, after determining the partial content that matches the material type specified by the filling area, the matching partial content is displayed in the filling area. Taking the matching partial content as a facial expression (smiling face expression) as an example, after displaying the smiling face expression in the filling area, an expression as shown in Figure 13 can be obtained.
[0136] In the embodiment, the partial content displayed in the filling area is usually the partial content matching the material type specified by the filling area in the video picture captured during video recording. The partial content is the content matching the material type specified by the filling area, and the content displayed in the filling area does not mismatch the expression template (for example, a body part is displayed instead of a face), so that the accuracy of the generated expression can be ensured.
[0137] In one embodiment, as shown in Figure 14 , the method further includes:
[0138] S302, in response to an expression template generation event, entering a template generation interface.
[0139] The expression template generation event is usually an event requiring generation of an expression template. For example, clicking an expression template generation button in a program can trigger the expression template generation event. The template generation interface can usually be an interface for editing and generating an expression template. As shown in Figure 15 , the template generation interface can include 1, a toolbar, which can include, from top to bottom, a polyline lasso (mouse-drawn polyline area), a curve lasso (mouse-drawn curve area), intelligent extraction (clicking a picture range to intelligently identify the area range by algorithm, similar to the magic wand tool), and chroma key extraction (clicking a color to select the surrounding similar color area). 2, a preview canvas including a displayed picture material and a filling area selected in the picture material. 3, a timeline for displaying each picture material of a gif (Graphics Interchange Format) animation. 4, an attribute panel for specifying the material type specified by the filling area as a certain type of body / facial / background. In Figure 15 , the material type specified by the filling area can be specified as a face.
[0140] Specifically, in response to the expression template generation event, the template generation interface is entered.
[0141] S304, displaying a picture material in the template generation interface.
[0142] Specifically, the picture material for generating the expression template can be selected from a database or other storage, and one picture material is selected from the picture material and displayed in the preview canvas area of the template generation interface.
[0143] S306, in response to a selection operation and a type specification operation on the picture material.
[0144] S308, selecting the fill region in the picture material according to the selecting operation, determining the material type specified by the fill region according to the type specifying operation, and displaying the expression template in the template generation interface.
[0145] The selecting operation can be an operation of selecting a region in the picture material by using a tool, such as a lasso tool or a selection tool. The type specifying operation can be an operation of specifying the material type of the region selected by the selecting operation, such as selecting a region and specifying the material type of the region as a face material type.
[0146] Specifically, when an expression template needs to be generated, the fill region in the picture material can be selected by using the selecting operation, and the material type specified by the fill region can be specified by using the type specifying operation, and then the generated expression template can be displayed in the template generation interface.
[0147] In some exemplary embodiments, as shown in Figure 15 The fill region in the picture material can be selected by using a curved lasso tool or a polyline lasso tool, and then the face option in the attribute panel can be used to determine the material type specified by the selected fill region as a face material type.
[0148] In this embodiment, by displaying the picture material in the template generation interface, the picture material can be processed by using the selecting operation and the type specifying operation, so that the user can customize the fill region and the material type of the expression template. The template generation interface can also display the obtained expression template, so that the user can preview and confirm the final effect.
[0149] In one embodiment, the template generation interface also has a timeline, when the picture material is a dynamic picture material. The dynamic picture material can be a gif type picture material. The displaying of the picture material in the template generation interface includes:
[0150] Displaying one frame of picture material included in the dynamic picture material in the template generation interface.
[0151] Specifically, the dynamic picture material can be composed of multiple frames of picture material. Therefore, when the picture material is a dynamic picture material, any frame of picture material included in the dynamic picture material can be displayed in the template generation interface.
[0152] Displaying each frame of picture material included in the dynamic picture material in the timeline.
[0153] Specifically, since the template generation interface also has a time axis, each frame of picture material contained in the dynamic picture material can also be displayed in the time axis, and each frame of picture material displayed in the time axis can be arranged and displayed in sequence in general. For example, there are 3 frames of picture material in the dynamic picture material, and the first frame of picture material, the second frame of picture material, and the third frame of picture material in the dynamic picture material are sequentially displayed in the time axis in the order of frame rate.
[0154] In response to a click operation on the picture material displayed in the time axis, one frame of picture material in the displayed dynamic picture material is switched to the frame of picture material indicated by the click operation.
[0155] The click operation can be an operation of clicking on a frame of picture material displayed in the time axis.
[0156] Specifically, when a user needs to view a frame of picture material displayed in the time axis, the user can click on the picture material. In response to the click operation on the picture material, the frame of picture material displayed in the template generation interface is switched to the picture material. For example, there are three frames of picture material in the time axis, and the first frame of picture material is currently displayed in the template generation interface. In response to a click operation on the second frame of picture material, the first frame of picture material displayed in the template generation interface is switched to the second frame of picture material.
[0157] In the embodiment, when the picture material is a dynamic picture material, each frame of picture material contained in the dynamic picture material can be displayed in the time axis, and the content of the entire dynamic picture material can be more intuitively connected. By performing a click operation on the picture material displayed in the time axis, the user can conveniently switch to a specified frame of picture material, view the specified frame of picture material in more detail, and subsequently process the specified frame of picture material. In addition, when the picture material is a dynamic picture material, the expression template generated based on the dynamic picture material is usually a dynamic expression template.
[0158] In one embodiment, the selecting the filling area in the picture material according to the selection operation, determining the material type of the filling area according to the type designation operation, and displaying the expression template in the template generation interface include:
[0159] For a frame of picture material displayed in the template generation interface, the filling area in the frame of picture material in the template generation interface is selected according to a selection operation.
[0160] The material type of the filling area in the frame of picture material in the template generation interface is determined according to the type designation operation.
[0161] displaying one frame of expression template generated for one frame of picture material in the template generation interface in the timeline.
[0162] displaying the one frame of expression template in the timeline.
[0163] Specifically, when the picture material to be processed is dynamic material picture, the generated expression template is usually dynamic expression template. Therefore, at least one frame of picture material needs to be processed to generate dynamic expression template. When processing at least one frame of picture material, one frame of picture material displayed in the template generation interface is usually processed. Therefore, in this embodiment, only one frame of picture material is processed to illustrate that, for one frame of picture material displayed in the template generation interface, one frame of picture material in the template generation interface is selected according to the selection operation, and the material type of the filling area in the one frame of picture material in the template generation interface is determined according to the type designation operation. After processing, one frame of expression template generated for one frame of picture material in the template generation interface can be displayed in the template generation interface. One frame of expression template is displayed in the timeline. When other frames of picture material need to be processed, other frames of picture material displayed in the timeline can be clicked to switch one frame of picture material displayed in the template generation interface to other frames of picture material, and then the processing is performed.
[0164] In this embodiment, by processing one frame of picture material displayed in the template generation interface, one frame of expression template is displayed in the timeline after processing, so that the difference between generating one frame of expression template and other frames of picture material can be more intuitively seen.
[0165] In one embodiment, the method further includes: in response to the batch processing operation, applying the selection operation and the type designation operation on one frame of picture material in the template generation interface to each of the remaining frames of picture material.
[0166] displaying each frame of expression template generated for each frame of picture material in the timeline.
[0167] The batch processing operation can usually be an operation of batch processing each frame of picture material. In some embodiments of the present disclosure, the batch processing operation can be realized by SIFT (Scale-Invariant Feature Transform) matching algorithm.
[0168] Specifically, after processing a single frame of image material using selection and type specification operations, the SIFT matching algorithm can be used to apply the filled area selected by the current selection operation to every remaining frame of image material. Similarly, the material type specified by the filled area determined by the type specification operation can be applied to the filled area indicated by each remaining frame of image material. In this way, each frame of image material generates a corresponding emoji template. For example... Figure 15 As shown, the timeline displays each frame of the emoji template generated for each frame of image material.
[0169] As one implementation method, such as Figure 15 As shown, you can click to apply it to all frames, thereby applying the selection operation of one frame of image material in the template generation interface and the type specification operation to each of the remaining frame of image material, thus generating a dynamic expression template.
[0170] In this embodiment, through batch processing, only one frame of image material can be processed, and the processing operation can be applied to other frames of image material, thus improving the efficiency of processing dynamic image material.
[0171] In one embodiment, the method further includes:
[0172] In response to a test trigger operation, a test emoticon is displayed in the template generation interface. The test emoticon includes the emoticon template and a preset image that fills the fill area of the emoticon template. The material type of the preset image is the same as the material type specified by the fill area.
[0173] Among them, the test trigger operation can be an operation to test the generated emoji template.
[0174] Specifically, in response to a test-triggered operation, the generated emoji template can be tested to preview the final effect of the generated emoji. For example... Figure 15 As shown, you can click to fill in the test content, and the template generation interface will display it as follows. Figure 16 The test emoticon shown is an example of a test emoticon. It includes an emoticon template and a preset image to fill the fill area of the emoticon template. The material type of the preset image is the same as the material type specified in the fill area; for example, if a facial expression needs to be filled, the preset image must also be a facial expression.
[0175] In this embodiment, by triggering a test operation, a test emoticon is displayed in the template generation interface, which allows for a convenient preview of the final effect of the generated emoticon.
[0176] In one embodiment, such as Figure 17 As shown, this disclosure also provides another information processing method, including:
[0177] Template creation process: A template maker can create a short video template and an expression template in a template editing server. The short video template and the expression template created are uploaded to a template storage server. In the process of uploading to the template storage server, the template storage server will verify whether the user music meets the length requirement and the compliance requirement through an algorithm or manual means. When the length requirement and the compliance requirement are met, the short video template and the expression template can be uploaded to the template storage server. When the template maker creates an expression template, in response to an expression template generation event, the template maker enters a template generation interface. In the template generation interface, a predetermined picture material is displayed. When the picture material is a dynamic picture material, for a frame of picture material displayed in the template generation interface, a filling area in the frame of picture material in the template generation interface is selected according to a selection operation; a material type of the filling area in the frame of picture material in the template generation interface is determined according to a type designation operation; a frame of expression template generated for the frame of picture material in the template generation interface is displayed in the template generation interface; and the frame of expression template is displayed in a timeline. In response to a batch processing operation, the selection operation and the type designation operation on the frame of picture material in the template generation interface are applied to each of the remaining frames of picture material, thereby generating a dynamic expression template; and each of the frames of expression template generated for each of the frames of picture material is displayed in the timeline. After the expression template is generated, in response to a test trigger operation, a test expression is displayed in the template generation interface, the test expression including the expression template and a preset picture filled in the filling area of the expression template, and the material type of the preset picture is the same as the material type designated by the filling area. Whether the generated expression template meets the requirement can be determined according to the test expression. When the expression template meets the requirement, the expression template is uploaded to the template storage server.
[0178] Video recording process: A short video creator can select and load an expression template and a short video template in a short video client of a terminal. A video is then shot in the short video client. In the process of shooting the video, in response to reaching a time point at which a recording length meets a time constraint condition, at least one expression template in an expression template group is displayed in an interface for video recording, together with action prompt information, the action prompt information being used to prompt a body action to be performed during the video recording. In response to a body action included in a local content of a video shot during the video recording matching the body action prompted by the action prompt information, the local content of the video shot during the video recording is displayed in the filling area.
[0179] When there is background music in the short video template, the background music is also played during the video recording. In response to reaching a rhythm beat point at which a rhythm of the background music meets a rhythm constraint condition, at least one expression template in an expression template group is displayed in an interface for video recording.
[0180] In response to reaching a recording node that causes the facial expression in the recording content to match the preset expression, at least one expression template in the expression template group indicated by the preset expression is displayed in the interface of the video recording. In response to a change in the facial expression in the recording content, the at least one expression template indicated by the displayed preset expression is switched to an expression template in the expression template group that has not been displayed.
[0181] After the expression template is displayed for a preset duration, or in response to an expression template switching instruction, at least one expression template in the expression template group is displayed in the interface of the video recording is switched to an expression template in the expression template group that has not been displayed, wherein the expression template switching instruction is an instruction for generating a successful expression indicated by the displayed at least one expression template.
[0182] The displayed expression template has a fill area. Therefore, the material type specified by the fill area in the expression template can be determined; the local content matching the material type specified by the fill area in the video captured by the video recording is obtained; the matching local content is displayed in the fill area to generate an expression. In response to the end of the video recording, a recording video and an expression group containing multiple expressions are generated in the short video client.
[0183] Application process: the recording video can be uploaded to the short video server, and the expression group can be uploaded to the sticker server. Subsequently, other users (viewers) can view the recording video through the short video client in the terminal, and view the stickers in the sticker group corresponding to the recording video. If the short video client also contains a chat client, the stickers in the sticker group can also be used for chatting.
[0184] In addition, it should be noted that the viewer, the short video creator, and the template maker can be the same person.
[0185] The application also provides some application scenarios of the information processing method. Specifically, the information processing method can also be applied to scenarios of generating dynamic expression templates or static expression templates, and can also be applied to video recording scenarios.
[0186] It should be understood that although each step in the flowchart involved in each embodiment as described above is shown in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, there is no strict order limitation for the execution of these steps, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.
[0187] Based on the same inventive concept, the embodiments of the present disclosure also provide an information processing device for implementing the information processing method as described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more information processing device embodiments provided below can refer to the limitations of the information processing method described above, which will not be repeated here.
[0188] In one embodiment, as shown in Figure 18 An information processing device 400 is provided, comprising a video recording module 402, a template display module 404 and a data acquisition module 406, wherein:
[0189] The video recording module 402 is configured to start video recording in response to a video recording trigger event;
[0190] The template display module 404 is configured to display at least one expression template in an expression template group in response to reaching a freeze point in the video recording process that meets a preset condition during the video recording process, the expression template having a fill area; and display partial content of a video picture captured during video recording in the fill area.
[0191] The data acquisition module 406 is configured to obtain a recorded video and an expression group containing multiple expressions in response to the end of video recording, each expression in the expression group containing at least one expression template displayed during the video recording process and partial content displayed in the fill area of the displayed expression template; the expression group is used to select an expression from the expression group in response to a user operation in a conversation and send in the conversation, and the recorded video is used for playback browsing.
[0192] In the embodiment, at least one expression template in the expression template group is displayed in response to reaching the freeze frame point meeting the preset condition in the video recording process, which can increase the interest in the video recording process and make the recorded video more lively and interesting. The partial content of the video picture captured in the video recording is displayed in the filling area, so that the user can accurately see the content in the filling area of the current expression template, thereby the user can adjust the partial content of the video picture captured in the video recording in a targeted manner, the generated expression can be more in line with the user's demand, and the user only needs to adjust the partial content of the video picture captured in the video recording in a targeted manner, without other operations, thereby saving the time and effort of the user. Since the expression group containing multiple expressions is generated in the video recording process, the user does not need to manually select the material and manually process, the process of user operation and processing is reduced, the difficulty of user operation is reduced, the expression group containing multiple expressions can be generated at one time, and the efficiency of expression generation is improved.
[0193] In an embodiment of the apparatus, the preset condition includes a time constraint condition for a recording time length of the video recording, the freeze frame point is a time point at which the recording time length meets the time constraint condition, and the template display module 404 is further configured to display at least one expression template in the expression template group in the interface of the video recording in response to reaching the time point at which the recording time length meets the time constraint condition, and hide the at least one expression template displayed in the interface of the video recording before displaying the next group of expression templates, wherein the expression templates displayed in the interface of the video recording are different each time.
[0194] In an embodiment of the apparatus, the template display module 404 is further configured to display action prompt information in the interface of the video recording and display at least one expression template in the expression template group indicated by the action prompt information, and the action prompt information is used to prompt a body action to be performed during the video recording.
[0195] In an embodiment of the apparatus, the template display module 404 is further configured to display the partial content of the video picture captured in the video recording in the filling area in response to the body action contained in the partial content of the video picture captured in the video recording matching the body action prompted by the action prompt information.
[0196] In an embodiment of the apparatus, the body action includes a facial expression; the apparatus further includes an expression generation module configured to identify facial feature points of the facial expression included in the partial content of the video frame captured by the video recording; calculate a confidence of the facial feature points; classify and identify the facial expression according to the confidence to determine a first expression type of the facial expression; determine a second expression type of the facial expression indicated by the action prompt information; and when the first expression type and the second expression type are the same, fill the facial expression included in the partial content of the video frame captured by the video recording into a filling area of at least one expression template displayed to generate an expression.
[0197] In an embodiment of the apparatus, background music is further played during the video recording; the preset condition includes a rhythm constraint condition for the background music played during the video recording; and the freeze point is a rhythm card point that makes the rhythm of the background music comply with the rhythm constraint condition. The template display module 404 is further configured to, in response to reaching the rhythm card point that makes the rhythm of the background music comply with the rhythm constraint condition, display at least one expression template in the expression template group in the interface of the video recording; and hide the at least one expression template displayed in the interface of the video recording before displaying the next group of expression templates, wherein the expression templates displayed in the interface of the video recording are different each time.
[0198] In an embodiment of the apparatus, the preset condition includes a content constraint condition for the recording content of the video recording, and the freeze point is a recording node that makes the body action or background information in the recording content comply with the content constraint condition. The template display module 404 is further configured to, in response to reaching the recording node that makes the body action or background information in the recording content comply with the content constraint condition, display at least one expression template in the expression template group in the interface of the video recording; and hide the at least one expression template displayed in the interface of the video recording before displaying the next group of expression templates, wherein the expression templates displayed in the interface of the video recording are different each time.
[0199] In an embodiment of the apparatus, the content constraint condition includes a preset expression, the body action includes a facial expression, and the freeze point is a recording node that makes the facial expression match the preset expression. The template display module 404 is further configured to, in response to reaching the recording node that makes the facial expression in the recording content match the preset expression, display at least one expression template in the expression template group indicated by the preset expression in the interface of the video recording.
[0200] In an embodiment of the apparatus, the apparatus further includes an expression switching module configured to switch the displayed at least one expression template indicated by the preset expression to an expression template in the expression template group that has not been displayed, in response to a change in facial expression in the recorded content.
[0201] In an embodiment of the apparatus, the expression switching module is further configured to switch the displayed at least one expression template in the video recording interface to an expression template in the expression template group that has not been displayed, after the expression template has been displayed for a preset time length or in response to an expression template switching instruction, wherein the expression template switching instruction is an instruction indicating that the expression of the generated and displayed at least one expression template is successful.
[0202] In an embodiment of the apparatus, the template display module 404 is further configured to determine a material type specified by a fill area in the expression template, acquire local content in a video picture captured during video recording that matches the material type specified by the fill area, and display the matched local content in the fill area.
[0203] In an embodiment of the apparatus, the apparatus further includes an expression template display module configured to enter a template generation interface in response to an expression template generation event, display picture material in the template generation interface, and display an expression template in the template generation interface in response to a selection operation on the picture material and a type specification operation.
[0204] In an embodiment of the apparatus, the template generation interface has a time axis, and when the picture material is dynamic picture material, the expression template display module is further configured to display one frame of picture material included in the dynamic picture material in the template generation interface, display each frame of picture material included in the dynamic picture material in the time axis, and switch the displayed one frame of picture material in the dynamic picture material to the frame of picture material indicated by a click operation on the picture material displayed in the time axis.
[0205] In an embodiment of the apparatus, the expression template display module is further configured to select a fill area in one frame of picture material in the template generation interface according to the selection operation for the one frame of picture material displayed in the template generation interface.
[0206] The type designation operation determines a material type of a fill area in a frame of picture material in the template generation interface; a frame of expression template generated for the frame of picture material in the template generation interface is displayed in the template generation interface; and the frame of expression template is displayed in a timeline.
[0207] In an embodiment of the apparatus, the apparatus further comprises a batch processing module configured to, in response to a batch processing operation, apply the selection operation and the type designation operation to each of the remaining frames of picture material in the template generation interface; and display each of the frames of expression template generated for each of the frames of picture material in the timeline.
[0208] In an embodiment of the apparatus, the apparatus further comprises a test expression display module configured to, in response to a test trigger operation, display a test expression in the template generation interface, the test expression comprising the expression template and a preset picture filled in a fill area of the expression template, a material type of the preset picture being the same as a material type designated by the fill area.
[0209] Each of the above modules in the information processing apparatus can be implemented in whole or in part by software, hardware, and a combination thereof. Each of the above modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to be invoked and executed by a processor to perform operations corresponding to each of the above modules.
[0210] In an embodiment, a computer device, which can be a server, is provided, and an internal structure diagram of the computer device can be as shown in Figure 19 The computer device includes a processor, a memory, and a network interface connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store a recorded video and an expression group. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement an information processing method.
[0211] In an embodiment, a computer device, which can be a terminal, is provided, and an internal structure diagram of the computer device can be as shown in Figure 20As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used for wired or wireless communication with external terminals. Wireless communication can be achieved through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to implement an information processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad provided on the shell of the computer device. It can also be an external keyboard, touchpad or mouse, etc.
[0212] Those skilled in the art can understand that, Figure 19 The structure shown in 20 is only a block diagram of part of the structure related to the present scheme, and does not constitute a limitation on the computer device to which the present scheme is applied. A specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0213] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in any of the above method embodiments.
[0214] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in any of the above method embodiments.
[0215] In one embodiment, a computer program product is provided, including a computer program, and the computer program is executed by a processor to implement the steps in any of the above method embodiments.
[0216] It should be noted that the body movements (such as facial expressions) contained in the local content of the video screen and the recorded video in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards.
[0217] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of each method can be included. Any reference to memory, database or other medium used in each embodiment provided by the present disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in each embodiment provided by the present disclosure can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in each embodiment provided by the present disclosure can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0218] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of each technical feature in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present disclosure.
[0219] The above embodiments only express several implementation manners of the present disclosure, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present disclosure. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present disclosure, a number of modifications and improvements can be made, which are within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the appended claims.
Claims
1. An information processing method characterized by comprising: The method comprises: starting video recording in response to a video recording trigger event; in response to reaching a preset condition-compliant freeze point in the video recording process, displaying at least one expression template in an expression template group, the expression template having a fill area; in the fill area, displaying partial content of a video picture captured during the video recording; in response to the end of the video recording, obtaining a recorded video and an expression group containing multiple expressions, each expression in the expression group containing the partial content displayed in the fill area of the displayed expression template.
2. The method of claim 1, wherein, The preset condition comprises a time constraint condition for the recording duration of the video recording, and the freeze point is a time point at which the recording duration meets the time constraint condition, and the response to reaching the freeze point in the video recording process that meets the preset condition to display at least one expression template in the expression template group comprises: in response to reaching the time point at which the recording duration meets the time constraint condition, displaying at least one expression template in the expression template group in the interface of the video recording; before displaying the next group of expression templates, hiding the at least one expression template displayed in the interface of the video recording, wherein the expression template displayed in the interface of the video recording is different each time.
3. The method of claim 2, wherein, The display of at least one expression template in the expression template group in the interface of the video recording comprises: displaying action prompt information in the interface of the video recording and displaying at least one expression template in the expression template group indicated by the action prompt information, the action prompt information being used to prompt a body movement during the video recording.
4. The method of claim 3, wherein, The display of the partial content of the video picture captured during the video recording in the fill area comprises: in response to the body movement contained in the partial content of the video picture captured during the video recording matching the body movement prompted by the action prompt information, displaying the partial content of the video picture captured during the video recording in the fill area.
5. The method of claim 4, wherein, The body movement contains facial expressions; the method further comprises: identifying facial feature points of the facial expressions contained in the partial content of the video picture captured during the video recording; calculating the confidence of the facial feature points; classifying and identifying the facial expressions according to the confidence to determine a first expression type of the facial expressions; determining a second expression type of the facial expressions indicated by the action prompt information; when the first expression type and the second expression type are the same, filling the facial expressions contained in the partial content of the video picture captured during the video recording into the fill area of the displayed at least one expression template to generate an expression.
6. The method of claim 1, wherein, Background music is also played during the video recording process; the preset condition comprises a rhythm constraint condition for the background music played during the video recording process; the freeze point is a rhythm card point at which the rhythm of the background music meets the rhythm constraint condition; The response to reaching the freeze point in the video recording process that meets the preset condition to display at least one expression template in the expression template group comprises: In response to reaching the rhythm beat point that makes the rhythm of the background music consistent with the rhythm constraint condition, display at least one expression template in the interface for video recording from the group of expression templates; Before displaying the next group of expression templates, hide the at least one expression template displayed in the interface for video recording, wherein the expression template displayed in the interface for video recording is different each time.
7. The method of claim 1, wherein, The preset condition includes a content constraint condition for the recording content of the video recording, and the beat point is a recording node that makes a body movement or background information in the recording content consistent with the content constraint condition; and the displaying at least one expression template from the group of expression templates in response to reaching the beat point that meets the preset condition during the video recording process includes: In response to reaching the recording node that makes the body movement or background information in the recording content consistent with the content constraint condition, display at least one expression template in the interface for video recording from the group of expression templates; Before displaying the next group of expression templates, hide the at least one expression template displayed in the interface for video recording, wherein the expression template displayed in the interface for video recording is different each time.
8. The method of claim 7, wherein, The content constraint condition includes a preset expression, and the body movement includes a facial expression; and the beat point is a recording node that makes the facial expression match the preset expression; The displaying at least one expression template from the group of expression templates in response to reaching the beat point that meets the preset condition during the video recording process includes: In response to reaching the recording node that makes the facial expression in the recording content match the preset expression, display at least one expression template in the interface for video recording from the group of expression templates indicated by the preset expression.
9. The method of claim 8, wherein, The method further includes: In response to a change in the facial expression in the recording content, switch the at least one expression template indicated by the preset expression that is displayed to an expression template that has not been displayed from the group of expression templates.
10. The method according to any one of claims 1 to 8, characterized in that, The method further includes: After the expression template is displayed for a preset duration or in response to an expression template switching instruction, switch the at least one expression template displayed in the interface for video recording from the group of expression templates to an expression template that has not been displayed from the group of expression templates, wherein the expression template switching instruction is a successful instruction for generating the expression indicated by the at least one expression template displayed.
11. The method of claim 1, wherein, The displaying, in the fill area, partial content of a video picture captured by video recording includes: Determining a material type specified by a fill area in the expression template; Obtaining partial content in a video picture captured by video recording that matches the material type specified by the fill area; Displaying the matching partial content in the fill area.
12. The method of claim 1, wherein, The method further includes: In response to an expression template generation event, entering a template generation interface; Displaying picture material in the template generation interface; In response to a selection operation and a type specification operation for the picture material; Selecting a fill area in the picture material according to the selection operation and determining a material type specified by the fill area according to the type specification operation, and displaying an expression template in the template generation interface.
13. The method of claim 12, wherein, The template generation interface has a time axis; when the picture material is dynamic picture material; The displaying of the picture material in the template generation interface comprises: Displaying a frame of picture material contained in the dynamic picture material in the template generation interface; Displaying each frame of picture material contained in the dynamic picture material in the time axis; In response to a click operation on the picture material displayed in the time axis, switching a frame of picture material in the displayed dynamic picture material to a frame of picture material indicated by the click operation.
14. The method of claim 13, wherein, The displaying of the expression template in the template generation interface according to the selecting operation of selecting the filling area in the picture material and the type specifying operation of determining the material type specified by the filling area comprises: According to the selecting operation, selecting a filling area in a frame of picture material displayed in the template generation interface; According to the type specifying operation, determining the material type of the filling area in the frame of picture material in the template generation interface; Displaying a frame of expression template generated for the frame of picture material in the template generation interface in the template generation interface; Displaying the frame of expression template in the time axis.
15. The method of claim 14, wherein, The method further comprises: In response to a batch processing operation, applying the selecting operation and the type specifying operation on each frame of picture material in the template generation interface; Displaying each frame of expression template generated for each frame of picture material in the time axis.
16. The method of claim 12, wherein, The method further comprises: In response to a test triggering operation, displaying a test expression in the template generation interface, the test expression containing the expression template and a preset picture filled in the filling area of the expression template, the material type of the preset picture being the same as the material type specified by the filling area.
17. An information processing apparatus comprising: The device comprises: A video recording module configured to start video recording in response to a video recording triggering event; A template display module configured to display at least one expression template in an expression template group in response to reaching a freeze point in a video recording process that meets a preset condition, the expression template having a filling area in which partial content of a video screen captured during the video recording is displayed; A data acquisition module configured to obtain a recorded video and an expression group containing multiple expressions in response to the end of the video recording, each expression in the expression group containing the partial content displayed in the filling area of the displayed expression template. 18.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-17. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 16.
19. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 16.
20. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 16.