Data processing methods, apparatus, electronic devices and storage media

By detecting the 3D pose of the hand in the camera device and rendering matching 3D text element effects, the problem of monotonous effects in camera applications is solved, improving fun and usability.

CN116112618BActive Publication Date: 2026-07-31BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2023-01-17
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, camera applications have fixed and monotonous special effects, which are not very interesting and have a low usage rate.

Method used

By detecting the hand in real-time footage when the camera device is turned on, the three-dimensional pose information of the hand is obtained, and a three-dimensional text element effect matching the hand is rendered in the image. The pose of the effect is controlled based on the three-dimensional pose information of the hand.

Benefits of technology

It enhances the diversity and fun of special effects during the shooting process, increasing user engagement and the usage rate of special effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116112618B_ABST
    Figure CN116112618B_ABST
Patent Text Reader

Abstract

This disclosure relates to a data processing method, apparatus, electronic device, and storage medium. The method includes, in response to an activation command from a camera device, performing hand detection on a real-time image captured by the camera device; if at least one hand is detected in the real-time image, acquiring three-dimensional pose information of at least one hand; rendering a three-dimensional text element effect matching the at least one hand in the real-time image, and controlling the pose of the three-dimensional text element effect based on the three-dimensional pose information. Using embodiments of this disclosure, effects can be selected in real-time based on the detected hand, greatly enhancing the diversity and fun of effects during shooting, as well as user engagement, and consequently significantly increasing the usage rate of text element effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of Internet technology, and in particular to a data processing method, apparatus, electronic device and storage medium. Background Technology

[0002] With the development of internet and camera technologies, a large number of camera applications or applications with camera functions have become an indispensable part of people's daily lives. Some applications provide 3D effects while offering camera functions. However, in these technologies, users often pre-select the corresponding effect templates, resulting in fixed and monotonous effects with poor fun, and consequently, low usage rates. Summary of the Invention

[0003] This disclosure provides a data processing method, apparatus, electronic device, and storage medium to at least solve the technical problems in related technologies, such as fixed and monotonous special effects, poor entertainment value, and low usage rate of special effects. The technical solution of this disclosure is as follows:

[0004] According to a first aspect of the present disclosure, a data processing method is provided, comprising:

[0005] In response to the camera device's activation command, hand detection is performed on the real-time footage captured by the camera device;

[0006] If at least one hand is detected in the real-time image, the three-dimensional pose information of the at least one hand is obtained;

[0007] In the real-time image, a 3D text element effect matching the at least one hand is rendered, and the pose of the 3D text element effect is controlled based on the 3D pose information.

[0008] In an optional embodiment, when the at least one hand is two hands, the three-dimensional text element effect is a first preset effect, the first preset effect having two relatively moving ends, and the amount of text elements in the first preset effect is positively correlated with the distance between the two ends.

[0009] In an optional embodiment, the method further includes:

[0010] Determine the origin hand and the target hand among the two hand parts;

[0011] Obtain the direction information between the origin hand and the target hand, as well as the distance information between the origin hand and the target hand;

[0012] The step of rendering a 3D text element effect that matches the at least one hand in the real-time image, and controlling the pose of the 3D text element effect based on the 3D pose information, includes:

[0013] In the real-time scene, the first end of the first preset special effect is aligned with the origin hand, and the pose of the first end is controlled based on the three-dimensional pose information of the origin hand; the second end of the first preset special effect is aligned with the target hand, and in the process of controlling the pose of the second end based on the three-dimensional pose information of the target hand, the text element corresponding to the first preset special effect is controlled to be displayed from the first end to the second end based on the direction information and the distance information;

[0014] Wherein, the first end is either of the two ends, and the second end is the other end of the two ends besides the first end.

[0015] In an optional embodiment, rendering a 3D text element effect that matches the at least one hand in the real-time image, and controlling the pose of the 3D text element effect based on the 3D pose information, further includes:

[0016] From the set of preset effects, determine the first preset effect that matches the two hands;

[0017] The preset effects set includes multiple preset 3D text element effects corresponding to the number of hands. The multiple preset 3D text element effects include the first preset effect, and the number of hands corresponding to the first preset effect is two.

[0018] In an optional embodiment, the preset effect set includes a preset effect subset, which is a set of preset 3D text element effects corresponding to two hands; determining the first preset effect matching the two hands from the preset effect set includes:

[0019] Determine the angle information of the line connecting the origin hand and the target hand relative to the direction of global gravity;

[0020] Based on the angle information, the first preset effect is determined from the preset effect subset.

[0021] In an optional embodiment, determining the first preset effect from the preset effect subset based on the angle information includes:

[0022] If the angle information is greater than the first preset angle threshold, the preset horizontal special effect in the preset special effect subset is taken as the first preset special effect.

[0023] In an optional embodiment, determining the first preset effect from the preset effect subset based on the angle information includes:

[0024] If the angle information is less than or equal to the second preset angle threshold, the preset vertical effect in the preset effect subset is taken as the first preset effect.

[0025] In an optional embodiment, the preset effects subset includes a first vertical effect corresponding to the first origin hand category and a second vertical effect corresponding to the second origin hand category; the method further includes:

[0026] Determine the target hand category to which the origin hand belongs;

[0027] When the angle information is less than or equal to a second preset angle threshold, using the preset vertical effects in the preset effects subset as the first preset effect includes:

[0028] If the angle information is less than or equal to the second preset angle threshold, the first preset effect is determined from the first vertical effect and the second vertical effect according to the target hand category.

[0029] In an optional embodiment, the preset effects subset includes a third vertical effect and a fourth vertical effect; the method further includes:

[0030] The relative position information of the origin hand in the real-time image is determined, and the relative position information indicates that the origin hand is located to the left or right of the corresponding center position in the real-time image;

[0031] When the angle information is less than or equal to a second preset angle threshold, using the preset vertical effects in the preset effects subset as the first preset effect includes:

[0032] If the angle information is less than or equal to the second preset angle threshold, the first preset effect is determined from the third vertical effect and the fourth vertical effect based on the relative position information.

[0033] The third vertical effect corresponds to the first relative position information, which indicates that the origin hand is located to the left of the center position of the real-time image; the fourth vertical effect corresponds to the second relative position information, which indicates that the origin hand is located to the right of the center position of the real-time image.

[0034] In an optional embodiment, the preset effects subset includes alternating fifth and sixth vertical effects; the method further includes:

[0035] Determine the first occurrence count of the fifth vertical effect and the second occurrence count of the sixth vertical effect during the operation of the camera device;

[0036] When the angle information is less than or equal to a second preset angle threshold, using the preset vertical effects in the preset effects subset as the first preset effect includes:

[0037] If the angle information is less than or equal to the second preset angle threshold, the vertical effect with the fewer occurrences among the fifth and sixth vertical effects is selected as the first preset effect based on the first occurrence count and the second occurrence count.

[0038] In an optional embodiment, determining the origin hand and the target hand among the two hands includes:

[0039] Determine the relative positional relationship between the two hands;

[0040] Based on the relative positional relationship, the two hands are divided into the origin hand and the target hand.

[0041] In an optional embodiment, the first end is the end closest to the starting text element corresponding to the first preset effect, and the second end is the other end of the two ends besides the first end.

[0042] In an optional embodiment, when the at least one hand is a single hand, the 3D text element effect is a second preset effect, which is an effect with fixed text elements; rendering the 3D text element effect that matches the at least one hand in the real-time image, and controlling the pose of the 3D text element effect based on the 3D pose information includes:

[0043] From the set of preset effects, determine the second preset effect that matches the hand.

[0044] The second preset effect is rendered in the real-time image, and the pose of the second preset effect is controlled based on the three-dimensional pose information;

[0045] The preset effects set includes multiple preset 3D text element effects corresponding to the number of hands, and the multiple preset 3D text element effects include the second preset effect, the number of hands corresponding to the second preset effect is one.

[0046] In an optional embodiment, the three-dimensional pose information comprises the first pose information of at least one hand in the current frame corresponding to the real-time image and the second pose information of at least one hand in a preset number of frames prior to the current frame; the rendering of a three-dimensional text element effect matching the at least one hand in the real-time image, and controlling the pose of the three-dimensional text element effect based on the three-dimensional pose information, includes:

[0047] The first pose information and the second pose information are weighted and fused to obtain the target pose information;

[0048] The 3D text element effect is rendered in the real-time image, and the pose of the 3D text element effect is controlled based on the target pose information.

[0049] In an optional embodiment, the method further includes:

[0050] In response to a video synthesis command, a target special effects video is generated based on the special effects footage during the camera device's activation process;

[0051] The special effects include the real-time footage during the camera's activation process and the rendered 3D text element effects within the real-time footage.

[0052] According to a second aspect of the present disclosure, a data processing apparatus is provided, comprising:

[0053] The hand detection module is configured to perform hand detection on the real-time images captured by the camera device in response to the camera device's activation command;

[0054] The three-dimensional pose information acquisition module is configured to acquire the three-dimensional pose information of at least one hand when the real-time image is detected to include at least one hand.

[0055] The 3D text element effects processing module is configured to render 3D text element effects that match the at least one hand in the real-time image, and to control the pose of the 3D text element effects based on the 3D pose information.

[0056] In an optional embodiment, when the at least one hand is two hands, the three-dimensional text element effect is a first preset effect, the first preset effect having two relatively moving ends, and the amount of text elements in the first preset effect is positively correlated with the distance between the two ends.

[0057] In an optional embodiment, the apparatus further includes:

[0058] The hand determination module is configured to determine the origin hand and the target hand among the two hands;

[0059] The information acquisition module is configured to acquire the direction information between the origin hand and the target hand, as well as the distance information between the origin hand and the target hand;

[0060] The 3D text element special effects processing module includes:

[0061] The first special effects processing unit is configured to, in the real-time frame, align the first end of the first preset special effect with the origin hand, and control the pose of the first end based on the three-dimensional pose information of the origin hand; align the second end of the first preset special effect with the target hand, and, in the process of controlling the pose of the second end based on the three-dimensional pose information of the target hand, control the text element corresponding to the first preset special effect to be displayed from the first end to the second end based on the direction information and the distance information;

[0062] Wherein, the first end is either of the two ends, and the second end is the other end of the two ends besides the first end.

[0063] In an optional embodiment, the 3D text element effects processing module further includes:

[0064] The first preset effect determination unit is configured to perform an action to determine the first preset effect that matches the two hands from the preset effect set;

[0065] The preset effects set includes multiple preset 3D text element effects corresponding to the number of hands. The multiple preset 3D text element effects include the first preset effect, and the number of hands corresponding to the first preset effect is two.

[0066] In an optional embodiment, the preset special effects set includes a preset special effects subset, which is a set of preset 3D text element special effects corresponding to two hand gestures; the first preset special effects determining unit includes:

[0067] An angle information determination unit is configured to determine the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity.

[0068] The first preset effect determination subunit is configured to determine the first preset effect from the preset effect subset based on the angle information.

[0069] In an optional embodiment, the first preset effect determination subunit is specifically configured to perform the following operation: when the angle information is greater than a first preset angle threshold, the preset horizontal effect in the preset effect subset is taken as the first preset effect.

[0070] In an optional embodiment, the first preset effect determination subunit is specifically configured to perform the following operation: when the angle information is less than or equal to a second preset angle threshold, the preset vertical effect in the preset effect subset is taken as the first preset effect.

[0071] In an optional embodiment, the preset effects subset includes a first vertical effect corresponding to the first origin hand category and a second vertical effect corresponding to the second origin hand category; the device further includes:

[0072] The target hand category determination module is configured to determine the target hand category to which the origin hand belongs;

[0073] The first preset effect determination subunit is specifically configured to determine the first preset effect from the first vertical effect and the second vertical effect according to the target hand category when the angle information is less than or equal to the second preset angle threshold.

[0074] In an optional embodiment, the preset effects subset includes a third vertical effect and a fourth vertical effect; the device further includes:

[0075] The relative position information determination module is configured to determine the relative position information of the origin hand in the real-time image, wherein the relative position information indicates that the origin hand is located to the left or right of the corresponding center position in the real-time image;

[0076] The first preset effect determination subunit is specifically configured to determine the first preset effect from the third vertical effect and the fourth vertical effect based on the relative position information when the angle information is less than or equal to the second preset angle threshold.

[0077] The third vertical effect corresponds to the first relative position information, which indicates that the origin hand is located to the left of the center position of the real-time image; the fourth vertical effect corresponds to the second relative position information, which indicates that the origin hand is located to the right of the center position of the real-time image.

[0078] In an optional embodiment, the preset effects subset includes alternating fifth and sixth vertical effects; the device further includes:

[0079] The occurrence count determination unit is configured to determine the first occurrence count of the fifth vertical effect and the second occurrence count of the sixth vertical effect during the operation of the camera device;

[0080] The first preset effect determination subunit is specifically configured to, when the angle information is less than or equal to the second preset angle threshold, select the vertical effect with fewer occurrences between the fifth vertical effect and the sixth vertical effect as the first preset effect based on the first occurrence count and the second occurrence count.

[0081] In an optional embodiment, the hand determination module includes:

[0082] A relative position relationship determination unit is configured to determine the relative position relationship between the two hands;

[0083] The hand segmentation unit is configured to perform the action of dividing the two hands into the origin hand and the target hand based on the relative positional relationship.

[0084] In an optional embodiment, the first end is the end closest to the starting text element corresponding to the first preset effect, and the second end is the other end of the two ends besides the first end.

[0085] In an optional embodiment, when the at least one hand is a single hand, the 3D text element effect is a second preset effect, which is an effect with fixed text elements; the 3D text element effect processing module includes:

[0086] The second preset effect determination unit is configured to perform an action to determine a second preset effect that matches the hand from a preset effect set;

[0087] The second special effects processing unit is configured to render the second preset special effects in the real-time image and control the pose of the second preset special effects based on the three-dimensional pose information.

[0088] The preset effects set includes multiple preset 3D text element effects corresponding to the number of hands, and the multiple preset 3D text element effects include the second preset effect, the number of hands corresponding to the second preset effect is one.

[0089] In an optional embodiment, the three-dimensional pose information comprises the first pose information of at least one hand in the current frame corresponding to the real-time image and the second pose information of at least one hand in a preset number of frames prior to the current frame; the three-dimensional text element special effects processing module includes:

[0090] The weighted fusion processing unit is configured to perform weighted fusion processing on the first pose information and the second pose information to obtain the target pose information;

[0091] The data processing unit is configured to render the 3D text element effect in the real-time screen and control the pose of the 3D text element effect based on the target pose information.

[0092] In an optional embodiment, the apparatus further includes:

[0093] The target special effects video generation module is configured to execute a video compositing instruction in response to generate a target special effects video based on the special effects footage during the camera device's activation process;

[0094] The special effects include the real-time footage during the camera's activation process and the rendered 3D text element effects within the real-time footage.

[0095] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method as described in any one of the first aspects above.

[0096] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any one of the data processing methods of the present disclosure.

[0097] According to a fifth aspect of the present disclosure, a computer program product including instructions is provided that, when run on a computer, causes the computer to perform the method as described in any one of the first aspects above.

[0098] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0099] When the camera device is activated, hand detection is performed on the real-time footage captured by the camera device. If at least one hand is detected in the real-time footage, the 3D pose information of at least one hand is obtained. 3D text element effects matching at least one hand are rendered in the real-time footage. This allows for real-time selection of effects based on the detected hand, greatly enhancing the diversity and fun of effects during shooting. Furthermore, controlling the pose of 3D text element effects based on the 3D pose information of at least one hand can significantly improve user engagement and, consequently, greatly increase the usage rate of text element effects.

[0100] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0101] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0102] Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment;

[0103] Figure 2 This is a flowchart illustrating a data processing method according to an exemplary embodiment;

[0104] Figure 3 This is a schematic diagram illustrating a real-time view of a hand, according to an exemplary embodiment.

[0105] Figure 4 This is a partial schematic diagram of a process for displaying a second preset special effect in a real-time video, according to an exemplary embodiment.

[0106] Figure 5 This is a flowchart illustrating another data processing method according to an exemplary embodiment;

[0107] Figure 6 This is a schematic diagram illustrating a real-time view of two hands according to an exemplary embodiment;

[0108] Figure 7 This is a partial schematic diagram of a process for displaying a preset horizontal effect in a real-time video, according to an exemplary embodiment.

[0109] Figure 8 This is a schematic diagram illustrating a partial process of converting a preset horizontal effect into a preset vertical effect in a real-time video, according to an exemplary embodiment.

[0110] Figure 9 This is a schematic diagram illustrating a portion of the process of displaying a preset vertical special effect in a real-time video, according to an exemplary embodiment.

[0111] Figure 10 This is a partial schematic diagram of a process for displaying two vertical special effects in a real-time video, according to an exemplary embodiment.

[0112] Figure 11 This is a block diagram of a data processing apparatus according to an exemplary embodiment;

[0113] Figure 12This is a block diagram illustrating an electronic device for data processing according to an exemplary embodiment. Detailed Implementation

[0114] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0115] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0117] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment, which may include a terminal 100 and a server 200.

[0118] In an optional embodiment, terminal 100 can be used to provide special effects processing services during the video recording process to any user. Specifically, terminal 100 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices, or software running on the aforementioned electronic devices, such as applications. Optionally, the operating system running on the electronic device can be, but is not limited to, Android, iOS, Linux, Windows, etc.

[0119] In an optional embodiment, server 200 can provide background services to terminal 100. Specifically, server 200 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.

[0120] In addition, it should be noted that, Figure 1 The example shown is merely one application environment provided by this disclosure. In practical applications, other application environments may also be included, such as more terminals.

[0121] In the embodiments described in this specification, the terminal 100 and the server 200 can be directly or indirectly connected through wired or wireless communication, and this disclosure does not impose any restrictions.

[0122] Figure 2 This is a flowchart illustrating a data processing method according to an exemplary embodiment. This method can be applied to electronic devices such as terminals. Figure 2 As shown, the method may include the following steps:

[0123] In step S201, in response to the camera device's activation command, hand detection is performed on the real-time image captured by the camera device.

[0124] In one optional embodiment, the camera activation command can be a startup command in a camera app, or a shooting page loading command triggered by a shooting page loading control set on a specific page. In a specific embodiment, for example, a short video browsing page often includes a shooting page loading control for users to access the short video shooting page. In another specific embodiment, for example, a photography page (a page for capturing still images) often includes a switching control (i.e., a shooting page loading control) to switch from the photography page to the video recording page (a page for capturing dynamic videos). Accordingly, the shooting page loading command can be triggered by clicking the switching control or similar operations.

[0125] In one specific embodiment, the aforementioned real-time image can be an image captured in real-time by a camera device; optionally, a pre-trained hand detection model can be used to detect hands in the real-time image. Specifically, the hand detection model can be trained on a first preset deep learning model based on sample hand images with hand position information annotations. Correspondingly, the image of the real-time image can be input into the hand detection model for hand detection. Optionally, if the hand detection model detects a hand in the real-time image, it can output the position information of the detected hand. Conversely, if no hand is detected in the real-time image, it may not output hand position information or may output empty position information. Accordingly, the position information output by the hand detection model can be used to determine whether the real-time image contains a hand.

[0126] In step S203, if at least one hand is detected in the real-time image, the three-dimensional pose information of at least one hand is acquired.

[0127] In an optional embodiment, when at least one hand is detected in the real-time image, a pre-trained 3D pose detection model can be used to obtain the 3D pose information of at least one hand. Specifically, the 3D pose detection model can be trained on a second preset deep learning model using sample hand images labeled with 3D pose information. Correspondingly, an image of the real-time image including at least one hand can be input into the 3D pose detection model for 3D pose detection to obtain the 3D pose information of at least one hand.

[0128] In one specific embodiment, the three-dimensional pose information of any hand may include the position information and rotation information of the hand. Optionally, depending on the actual application requirements, the three-dimensional pose information of the center position of the hand can be used as the three-dimensional pose information of the hand; alternatively, the three-dimensional pose information of a specified joint of the hand can be used as the three-dimensional pose information of the hand; or, the three-dimensional pose information corresponding to the line connecting the specified joint of the hand and the center position of the hand can be used as the three-dimensional pose information of the hand.

[0129] In step S205, a 3D text element effect matching at least one hand is rendered in the real-time screen, and the pose of the 3D text element effect is controlled based on the 3D pose information.

[0130] In one specific embodiment, the 3D text element effect can be a 3D effect with text elements. In an optional embodiment, when at least one hand is a single hand, the aforementioned 3D text element effect is a second preset effect, which is an effect with fixed text elements; correspondingly, rendering the 3D text element effect matching at least one hand in the real-time image, and controlling the pose of the 3D text element effect based on 3D pose information, can include:

[0131] From the set of preset effects, determine a second preset effect that matches a hand;

[0132] The second preset effect is rendered in real time, and the pose of the second preset effect is controlled based on the three-dimensional pose information.

[0133] In a specific embodiment, the above preset special effect set may be a set of a large number of pre-set three-dimensional special effects with text elements (pre-set three-dimensional text element special effects). Specifically, the preset special effect set may include multiple pre-set three-dimensional text element special effects corresponding to the number of hands, and the multiple pre-set three-dimensional text element special effects include the second preset special effect; correspondingly, the matching of special effects can be performed in combination with the number of hands. Correspondingly, determining the second preset special effect that matches one hand from the preset special effect set may include: using the pre-set three-dimensional text element special effect with the corresponding number of hands being one in the preset special effect set as the second preset special effect, that is, the number of hands corresponding to the second preset special effect is one.

[0134] In a specific embodiment, the three-dimensional pose information of one hand detected in the real-time screen can be assigned to the second preset special effect (that is, using the three-dimensional pose information of the hand as the three-dimensional pose information of the second preset special effect), and thus the pose of the second preset special effect can be controlled based on the pose of the hand.

[0135] In a specific embodiment, as Figure 3 shown, Figure 3 FIG. is a schematic diagram of a real-time screen showing one hand provided according to an exemplary embodiment; further, real-time hand detection can be performed on the content of the real-time screen. Correspondingly, when it is detected that the real-time screen includes one hand, the three-dimensional pose information of one hand can be obtained, and then the rendering and pose control of the second preset special effect that matches one hand can be performed; specifically, as Figure 4 shown, Figure 4 FIG. is a partial process schematic diagram of showing the second preset special effect in the real-time screen. Among them, Figure 4 a in FIG. shows a schematic diagram of showing the second preset special effect in the real-time screen when one hand is detected, Figure 4 b in FIG. shows a schematic diagram after the pose of the second preset special effect in the real-time screen changes correspondingly with the change of the three-dimensional pose information of the hand. Combining Figure 4 it can be seen that the fixed text element in the second preset special effect is the character "Fu". Optionally, in practical applications, in addition to the fixed text element "Fu", the second preset special effect may also include other elements, which can be set according to actual application requirements, such as a preset pattern around the character "Fu".

[0136] In the above embodiments, when a hand is included in the real-time image, a second preset effect with fixed text elements is matched from the preset effect set and rendered based on the hand. The pose of the second preset effect is controlled based on the three-dimensional pose information of the hand. This allows for the selection of different types of current rendering features based on the number of hands, greatly enhancing the diversity of effects and the fun of using the effects. Furthermore, by combining the three-dimensional pose information of the hand to control the pose of the second preset effect, the second preset effect can move, rotate, and flip accordingly as the hand moves, rotates, and flips. This can greatly enhance user engagement and thus improve user participation rate.

[0137] In an optional embodiment, when at least one hand is two hands, the above-mentioned three-dimensional text element effect is a first preset effect, the first preset effect has two relatively movable ends, and the amount of text elements in the first preset effect is positively correlated with the distance between the two ends.

[0138] In the above embodiments, when at least one hand in the real-time image is two hands, the first preset feature that has two relatively moving ends and the amount of text elements in the first preset effect is positively correlated with the distance between the two ends is used as the three-dimensional text element effect to be rendered, which can improve the matching between the hand and the three-dimensional text effect element.

[0139] In an optional embodiment, such as Figure 5 As shown, the above method may further include the following steps:

[0140] In step S207, the origin hand and the target hand are determined among the two hands;

[0141] In step S209, the direction information between the origin hand and the target hand, as well as the distance information between the origin hand and the target hand, are obtained;

[0142] Accordingly, in step S205 above, rendering a 3D text element effect that matches at least one hand in the real-time frame, and controlling the pose of the 3D text element effect based on 3D pose information, may include:

[0143] In the real-time video, the first end of the first preset effect is aligned with the origin hand, and the pose of the first end is controlled based on the three-dimensional pose information of the origin hand; the second end of the first preset effect is aligned with the target hand, and in the process of controlling the pose of the second end based on the three-dimensional pose information of the target hand, the text element corresponding to the first preset effect is controlled to be displayed from the first end to the second end based on the direction information and distance information.

[0144] In an optional embodiment, determining the origin hand and the target hand among the two hands may include:

[0145] Determine the relative positional relationship between the two hands;

[0146] Based on their relative positions, the two hands are divided into the origin hand and the target hand.

[0147] In one specific embodiment, the origin hand can be a hand aligned with the first end of the first preset effect. The target hand can be a hand aligned with the second segment of the first preset effect. The first end can be any one of the two ends, and the second end can be the other end of the two ends besides the first end.

[0148] In one specific embodiment, the division between the origin hand and the target hand can be based on the relative vertical or horizontal relationship between the two hands. Optionally, taking the division between the origin hand and the target hand based on the relative vertical relationship between the two hands as an example, the above-mentioned division of the two hands into the origin hand and the target hand based on the relative position relationship may include: designating the hand that is above the other hand as the origin hand, and designating the hand that is below the other hand as the target hand.

[0149] Optionally, when the two hands are on the same horizontal line, the division between the origin hand and the target hand can be made by combining the relative left and right relationship between the two hands. Optionally, the above division of the two hands into the origin hand and the target hand based on the relative position relationship can include taking the left hand as the origin hand and the right hand as the target hand.

[0150] In the above embodiments, by combining the relative positional relationship between the two hands, the origin hand and the target hand can be quickly and effectively divided, thereby greatly improving the smoothness of the special effects rendering process.

[0151] In an optional embodiment, the first end can be the end closest to the starting text element corresponding to the first preset effect, and the second end can be the other end besides the first end. In practical applications, text information often has a sequential order; correspondingly, the starting text element can be the first text element among the text elements corresponding to the first preset effect. For example, if the text elements corresponding to the first preset effect are, in order, "New," "Spring," "Happy," and "Joyful," then the starting text element can be: "New."

[0152] Furthermore, it should be noted that in the actual special effects control process, not only can the text elements in the first preset special effects be displayed sequentially starting from the starting text element, but the text elements in the first preset special effects can also be displayed sequentially starting from any specified text element, such as starting from the ending text element, or starting from any non-combined beginning and end text element and displaying them sequentially to both sides.

[0153] In the above embodiments, taking the end closest to the starting text element corresponding to the first preset effect as the first end makes it easier to display the text elements corresponding to the first preset effect sequentially from the starting text element during the display of the first preset effect. This can greatly improve the readability of text elements during the effect rendering control process, and thus better enhance the user experience.

[0154] In an optional embodiment, the above-described rendering of a 3D text element effect matching at least one hand in a real-time frame, and the control of the pose of the 3D text element effect based on 3D pose information, may further include:

[0155] From the set of preset effects, determine the first preset effect that matches the two hands;

[0156] In one specific embodiment, the aforementioned preset effect set includes multiple preset 3D text element effects corresponding to the number of hands. These multiple preset 3D text element effects include a first preset effect. Accordingly, the effect matching can be performed based on the number of hands. That is, determining the first preset effect that matches two hands from the preset effect set can include: using the 3D effect in the preset effect set corresponding to two hands as the first preset effect. In other words, the number of hands corresponding to the first preset effect is two.

[0157] In the above embodiments, by pre-setting a set of preset effects that includes multiple preset 3D text element effects corresponding to the number of hands, it is possible to select different types of current rendering features based on the number of hands, which greatly enhances the diversity of effects and the fun of using the effects.

[0158] In another optional embodiment, the preset effect set may contain multiple preset 3D text element effects corresponding to the number of hands being two. Optionally, when there are two hands in the real-time image, the effect selection can be made by combining the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity. Accordingly, the aforementioned preset effect set may include a subset of preset effects, which is a set of preset 3D text element effects corresponding to the number of hands being two. The determination of the first preset effect matching the two hands from the preset effect set may include:

[0159] Determine the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity;

[0160] Based on the angle information, the first preset effect is determined from the preset effect subset.

[0161] In one specific embodiment, the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity can be less than or equal to 90 degrees. In an optional embodiment, determining the first preset effect from a preset effect subset based on the angle information may include:

[0162] If the angle information is greater than the first preset angle threshold, the preset horizontal effect in the preset effect subset will be used as the first preset effect.

[0163] In an optional embodiment, the first preset angle threshold can be the lower limit of the filtering angle corresponding to the preset horizontal effect, which can be preset according to the actual application, such as 85 degrees. Optionally, if the angle information is greater than the first preset angle threshold, it can be determined that the line connecting the origin hand and the target hand is horizontal or close to horizontal; accordingly, the preset horizontal effect can be used as the first preset effect. Specifically, the preset horizontal effect can have a three-dimensional effect of text elements displayed sequentially from left to right.

[0164] In the above embodiments, when the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity is greater than the first preset angle threshold, the preset horizontal special effect in the preset special effect subset is used as the first preset special effect. This can make the displayed special effect more compatible with the posture of the two hands, thereby better improving the fun of the special effect and the user experience.

[0165] In a specific embodiment, such as Figure 6 As shown, Figure 6 This is a schematic diagram illustrating a real-time image displaying two hands, provided according to an exemplary embodiment. Further, real-time hand detection can be performed on the content of the real-time image. Correspondingly, when two hands are detected in the real-time image, the three-dimensional pose information of the two hands can be obtained, thereby enabling the rendering and pose control of a first preset effect matching the two hands. Specifically, when the two hands are on the same horizontal line, the division between the origin hand and the target hand can be determined by combining the relative left-right relationship between the two hands. Optionally, the hand on the left can be taken as the origin hand, and the hand on the right of the two hands can be taken as the target hand. Further, the angle information of the line connecting the two hands relative to the direction of world gravity (vertically downward) can be 90 degrees. Optionally, assuming the first preset angle threshold is 85 degrees, if 90 degrees is greater than 85 degrees, a preset horizontal effect can be determined as the first preset effect (the effect currently to be rendered). Figure 7 As shown, Figure 7 This is a partial schematic diagram of a process for displaying a preset horizontal effect in a real-time video, according to an exemplary embodiment. Figure 7Figure a illustrates a method for displaying preset horizontal effects in real-time video; furthermore, as the two hands move, Figure 7 Figure b shows a schematic diagram illustrating the real-time changes in the preset lateral effect as the 3D pose information of the two hands changes, along with the distance between the two hands and the direction information between the origin hand and the target hand. Combined with... Figure 7 As can be seen, the text elements corresponding to the first preset effect are "New", "Spring", "Happy", and "Joy". Optionally, in actual application, the pose of the first preset effect and the amount of text elements displayed in real time will change accordingly as the hand moves.

[0166] In an optional embodiment, determining the first preset effect from a subset of preset effects based on angle information may include:

[0167] If the angle information is less than or equal to the second preset angle threshold, the preset vertical effect in the preset effect subset will be used as the first preset effect.

[0168] In one specific embodiment, the second preset angle threshold is less than or equal to the first preset angle threshold. The second preset angle threshold can be the upper limit of the filtering angle corresponding to the preset vertical effect, and can be preset according to actual application, such as 80 degrees. When the angle information is less than or equal to the second preset angle threshold, it can be determined that the line connecting the origin hand and the target hand is vertical or close to vertical; accordingly, the preset vertical effect can be used as the first preset effect. Specifically, the preset vertical effect can have a three-dimensional effect of text elements displayed sequentially from top to bottom.

[0169] In the above embodiments, when the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity is less than or equal to the second preset angle threshold, the preset vertical special effect in the preset special effect subset is used as the first preset special effect. This can make the displayed special effect more compatible with the posture of the two hands, thereby better improving the fun of the special effect and the user experience.

[0170] In another specific embodiment, when the real-time image detects two hands, the division between the origin hand and the target hand can be determined by combining the relative vertical relationship between the two hands. Optionally, the hand located above can be taken as the origin hand, and the hand located below can be taken as the target hand. Further, assuming the angle information of the line connecting the two hands relative to the world's gravity direction (vertically downward) can be 78 degrees, and optionally, assuming the second preset angle threshold is 80 degrees, since 78 degrees is less than 80 degrees, the preset vertical effect can be determined as the first preset effect (the effect that needs to be rendered currently). Figure 8 As shown, Figure 8It is a partial process schematic diagram of converting a preset horizontal special effect into a preset vertical special effect in a real-time screen according to an exemplary embodiment. Figure 8 A schematic diagram showing a preset horizontal special effect in a real-time screen as shown in a; further, as the two hands move, Figure 8 A schematic diagram of a preset vertical special effect as shown in b.

[0171] In a specific embodiment, in combination with the above Figure 8 The schematic diagram of the preset vertical special effect as shown in b. Figure 9 It is a partial process schematic diagram of showing a preset vertical special effect in a real-time screen according to an exemplary embodiment; Figure 9 A schematic diagram showing a preset horizontal special effect in a real-time screen as shown in a, further, as the two hands move, Figure 9 A schematic diagram of a preset vertical special effect after real-time change of the distance information between the two hands and the direction information between the origin hand and the target hand among the two hands as the three-dimensional pose information of the two hands changes, as shown in b; in combination with Figure 9 It can be seen that the text elements corresponding to the first preset special effect are joy, arrival, happiness, arrival, joy, arrival; optionally, in practical applications, as the two hands move, the pose of the first preset special effect and the amount of text elements shown in real time will also change accordingly.

[0172] In the above embodiment, in combination with the angle information of the line connecting the origin hand and the target hand relative to the world gravity direction, the first preset special effect is determined from the preset special effect subset. When the number of hands is two, the special effect selection can be further performed in combination with the angle information, which can improve the diversity of the special effects, and thus can also better improve the fun of the user during the use of the special effects.

[0173] In an optional embodiment, there may be multiple (at least two) vertical special effects in the preset special effect set (three-dimensional special effects with text elements shown sequentially from top to bottom). Optionally, the special effect selection can be further performed in combination with the origin hand category; optionally, the above preset special effect subset includes a first vertical special effect corresponding to the first origin hand category and a second vertical special effect corresponding to the second origin hand category; correspondingly, the above method may further include:

[0174] Determine the target hand category to which the origin hand belongs;

[0175] Correspondingly, the above step of taking the preset vertical special effect in the preset special effect subset as the first preset special effect when the angle information is less than or equal to the second preset angle threshold may include:

[0176] If the angle information is less than or equal to the second preset angle threshold, the first preset effect is determined from the first vertical effect and the second vertical effect according to the target hand category.

[0177] In one specific embodiment, the hand category may include a first origin hand category and a second origin hand category; in an optional embodiment, the first origin hand category may be a left hand, and the second origin hand category may be a right hand; in another optional embodiment, the first origin hand category may be a right hand, and correspondingly, the second origin hand category may be a left hand; when the target category is the first origin hand category, the first vertical effect can be used as the first preset effect; when the target category is the second origin hand category, the second vertical effect can be used as the first preset effect. Specifically, both the first and second vertical effects are three-dimensional effects with text elements displayed sequentially from top to bottom, but the text elements in the first and second vertical effects are different.

[0178] In the above embodiments, based on the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity for selecting vertical special effects, and further combined with the category of the origin hand, the first preset special effect can be selected from different vertical special effects, which can improve the diversity of special effects and thus better enhance the fun of users in using special effects.

[0179] In an optional embodiment, there may be multiple (at least two) vertical effects (three-dimensional effects with text elements displayed sequentially from top to bottom) in the preset effects set. Optionally, the selection of effects can be further combined with the relative position information of the origin hand in the real-time image. Optionally, the above-mentioned preset effects subset includes a third vertical effect and a fourth vertical effect. The above method may also include:

[0180] Determine the relative position of the hand at the origin in the real-time image;

[0181] Accordingly, when the angle information is less than or equal to the second preset angle threshold, using the preset vertical effects in the preset effects subset as the first preset effect may include:

[0182] If the angle information is less than or equal to the second preset angle threshold, the first preset effect is determined from the third and fourth vertical effects based on the relative position information.

[0183] In one specific embodiment, the aforementioned relative position information can indicate that the origin hand is located to the left or right of the corresponding center position in the real-time image. Specifically, the third vertical effect corresponds to the first relative position information, which indicates that the origin hand is located to the left of the corresponding center position in the real-time image; the fourth vertical effect corresponds to the second relative position information, which indicates that the origin hand is located to the right of the corresponding center position in the real-time image. Both the third and fourth vertical effects are 3D effects with text elements displayed sequentially from top to bottom, but the text elements in the third and fourth vertical effects are different. Optionally, when the relative position information is the first relative position information, the third vertical effect can be used as the first preset effect; when the relative position information is the second relative position information, the fourth vertical effect can be used as the first preset effect.

[0184] In the above embodiments, based on the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity for selecting vertical effects, and further combined with the relative position information of the origin hand in the real-time image, the first preset effect can be selected from different vertical effects. This can improve the diversity of effects and thus better enhance the fun of using effects.

[0185] In an optional embodiment, there may be multiple (at least two) vertical effects (three-dimensional effects with text elements displayed sequentially from top to bottom) in the preset effect set. Optionally, the vertical effects can be displayed alternately. Optionally, the aforementioned preset effect subset includes alternating fifth and sixth vertical effects. The method may also include:

[0186] Determine the number of times the fifth vertical effect appears first and the number of times the sixth vertical effect appears second during the camera's operation.

[0187] Accordingly, when the angle information is less than or equal to the second preset angle threshold, using the preset vertical effects in the preset effects subset as the first preset effect may include:

[0188] If the angle information is less than or equal to the second preset angle threshold, the vertical effect that appears less frequently between the fifth and sixth vertical effects is selected as the first preset effect based on the first and second occurrence counts.

[0189] In one specific embodiment, the first occurrence count can be the number of times the fifth vertical effect appears during the camera's activation. The second occurrence count can be the number of times the sixth vertical effect appears during the camera's activation. Optionally, the vertical effect that is displayed first can be pre-selected from the fifth and sixth vertical effects. By determining the occurrence counts of the fifth and sixth vertical effects during the camera's activation, the alternating display of the fifth and sixth vertical effects can be controlled.

[0190] In one specific embodiment, both the fifth and sixth vertical effects are 3D effects featuring text elements displayed sequentially from top to bottom, but the text elements in the fifth and sixth vertical effects are different. Optionally, if the first occurrence frequency is less than the second occurrence frequency, the fifth vertical effect can be used as the first preset effect; if the first occurrence frequency is greater than the second occurrence frequency, the sixth vertical effect can be used as the first preset effect.

[0191] In the above embodiments, based on the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity for selecting vertical effects, and further considering the occurrence frequency of the fifth and sixth vertical effects that alternate during the camera device's operation, the first preset effect to be rendered and displayed can be selected from different vertical effects. This can enhance the diversity of effects and thus improve the user's enjoyment of using the effects.

[0192] In a specific embodiment, the direction information between the origin hand and the target hand can be the direction vector from the origin hand to the target hand; the distance information between the origin hand and the target hand can be the magnitude of the distance between them. Specifically, the three-dimensional pose information of the origin hand can be assigned to the first end so that the first end of the first preset effect moves with the movement of the origin hand, and the three-dimensional pose information of the target hand can be assigned to the second end so that the second end of the first preset effect moves with the movement of the origin hand. Further, the text element corresponding to the first preset effect can include multiple text elements. Correspondingly, the greater the current distance between the two hands, the more text elements are currently displayed; conversely, the smaller the current distance between the two hands, the fewer text elements are currently displayed. Accordingly, during the movement of the second end based on the three-dimensional pose information of the target hand, the amount of text elements displayed in real time can be controlled by combining the distance information, and the text elements corresponding to the amount of text elements can be displayed sequentially from the first end to the second end by combining the direction information.

[0193] In the above embodiments, when at least one of the hands in the real-time screen is two hands, the first preset feature with two ends having relative movement and the amount of text elements in the first preset special effect being positively correlated with the distance between the two ends is used as the three-dimensional text element special effect to be rendered currently, which can improve the matching between the hand and the three-dimensional text special effect element. By dividing the origin hand and the target hand, while controlling the poses of both ends of the three-dimensional text element special effect based on the three-dimensional pose information of the hand, the direction information between the origin hand and the target hand and the distance information between the origin hand and the target hand can be combined to control the text element corresponding to the first preset special effect to be displayed from the first end to the second end. Furthermore, on the basis of ensuring the real-time nature of the hand's movement control over the three-dimensional text element special effect, the control of the amount of text elements in the special effect can be achieved, greatly enhancing the authenticity of the special effect.

[0194] In a specific embodiment, taking the two vertical special effect pose controls based on two hands in the real-time screen as an example, assume that the text elements corresponding to the two vertical special effects are the upper and lower couplets of a pair of couplets. Among them, the text elements corresponding to one vertical special effect include "Spring", "Returns", "To", "The", "Earth", "People", "Warm" (upper couplet); the text elements corresponding to the other vertical special effect include "Good", "Fortune", "Descends", "To", "The", "Land", "Joy" (lower couplet). As Figure 10 shown, Figure 10 is a partial process schematic diagram of displaying two vertical special effects in a real-time screen provided according to an exemplary embodiment. Optionally, during the process of two vertical special effect pose controls based on two hands in the real-time screen, the vertical special effects in the real-time screen can be switched by combining "the angle information of the line connecting the origin hand and the target hand relative to the world gravity direction and the target hand category to which the origin hand belongs", or "the angle information of the line connecting the origin hand and the target hand relative to the world gravity direction and the relative position information of the origin hand in the real-time screen", or "the angle information of the line connecting the origin hand and the target hand relative to the world gravity direction and the appearance times of the two alternately appearing vertical special effects". Optionally, Figure 10 is a schematic diagram of a vertical special effect shown in a). Figure 8 a shows some of the text elements of the vertical special effect corresponding to the upper couplet: "Spring", "Returns", "To", "The", "Earth"; further, as the poses of the two hands change, Figure 8 is a schematic diagram of another vertical special effect shown in b); Figure 10 b shows some of the text elements of the vertical special effect corresponding to the lower couplet: "Good", "Fortune", "Descends", "To", "The", "Land".

[0195] In addition, it should be noted that during the process of controlling the pose of two vertical special effects based on two hands in the real-time image, if one hand is detected, the second preset special effect corresponding to that hand can also be displayed. If two hands are detected, and the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity is combined, the first preset special effect matching the two hands is determined to be a preset horizontal special effect, and the preset horizontal special effect can also be displayed.

[0196] In an optional embodiment, the aforementioned three-dimensional pose information may be the first pose information of at least one hand in the current frame corresponding to the real-time image and the second pose information of at least one hand in a preset number of frames prior to the current frame; correspondingly, the first pose information and the second pose information may be obtained by combining the real-time image corresponding to the current frame and the real-time image corresponding to the preset number of frames prior to the current frame. Accordingly, rendering a three-dimensional text element effect matching at least one hand in the real-time image, and controlling the pose of the three-dimensional text element effect based on the three-dimensional pose information, may include:

[0197] The first pose information and the second pose information are weighted and fused to obtain the target pose information;

[0198] Render 3D text element effects in real-time, and control the pose of the 3D text element effects based on the target pose information.

[0199] In a specific embodiment, the preset number can be set according to actual application requirements. The current frame can be the frame corresponding to the current real-time screen. When performing weighted fusion processing on the first pose information and the second pose information, the weight information corresponding to the first pose information and the second pose information can be preset. Optionally, when there are multiple preset numbers, the weight information corresponding to the second pose information corresponding to different frames can be different. Optionally, the sum of the weights corresponding to the first pose information and the second pose information is equal to 1, with the first pose information having the largest weight. The weight information corresponding to the second pose information of the preset number of frames is negatively correlated with the interval between the corresponding frame and the current frame; that is, the smaller the interval, the larger the corresponding weight information. Further, for each hand, the first pose information of the hand in the current frame, the weight information corresponding to the first pose information, the second pose information of the hand in the preset number of frames, and the weight information corresponding to the second pose information in the preset number of frames can be weighted and summed to obtain the target pose information corresponding to each hand. The target pose information can be the 3D pose information used to control the effects of the hand object and 3D text elements in the current frame (the current real-time screen).

[0200] In the above embodiments, by performing weighted fusion processing on the first pose information of at least one hand in the current frame and the second pose information of at least one hand in a preset number of frames before the current frame, the target pose information of each hand in the current real-time screen can be generated. This can reduce the instability caused by rapid hand movement while ensuring the real-time control of the hand's movement of 3D text element effects.

[0201] In an optional embodiment, the above method may further include:

[0202] In response to video compositing instructions, a target special effects video is generated based on the special effects footage during the camera device's activation process;

[0203] In one specific embodiment, the special effects footage includes real-time footage during the camera's startup process and rendered 3D text element effects within that real-time footage. Optionally, the video compositing instruction can be triggered by combining the end-of-shoot control on the shooting interface corresponding to the real-time footage, thereby generating a target special effects video based on the real-time footage during the camera's startup process and rendered 3D text element effects within that real-time footage. Specifically, the target special effects video is a video with 3D text element effects.

[0204] In the above embodiments, when a video synthesis command is triggered, a target special effects video is generated based on the real-time footage during the camera device's startup process and the rendered 3D text element effects in the real-time footage. This can help users record the special effects footage during the camera device's startup process.

[0205] As can be seen from the technical solutions provided in the embodiments of this specification above, when the camera device is activated, hand detection is performed on the real-time image captured by the camera device; when at least one hand is detected in the real-time image, the three-dimensional pose information of at least one hand is obtained; and a three-dimensional text element effect matching at least one hand is rendered in the real-time image. This allows for the selection of effects in real time according to the detected hand, greatly improving the diversity and fun of effects during shooting. Furthermore, controlling the pose of the three-dimensional text element effect based on the three-dimensional pose information of at least one hand can greatly enhance user engagement and thus significantly increase the usage rate of text element effects.

[0206] Figure 11 This is a block diagram of a data processing apparatus according to an exemplary embodiment. (Refer to...) Figure 11 The device includes:

[0207] The hand detection module 1110 is configured to perform hand detection on the real-time image captured by the camera device in response to the camera device's activation command.

[0208] The three-dimensional pose information acquisition module 1120 is configured to acquire the three-dimensional pose information of at least one hand when at least one hand is detected in the real-time image.

[0209] The 3D text element effects processing module 1130 is configured to render 3D text element effects that match at least one hand in a real-time scene, and to control the pose of the 3D text element effects based on 3D pose information.

[0210] In an optional embodiment, when at least one hand is two hands, the three-dimensional text element effect is a first preset effect, the first preset effect has two relatively movable ends, and the amount of text elements in the first preset effect is positively correlated with the distance between the two ends.

[0211] In an optional embodiment, the above-described apparatus further includes:

[0212] The hand determination module is configured to determine the origin hand and the target hand among two hands;

[0213] The information acquisition module is configured to acquire the direction information between the origin hand and the target hand, as well as the distance information between the origin hand and the target hand.

[0214] The 3D text element special effects processing module 1130 includes:

[0215] The first special effects processing unit is configured to execute in the real-time screen, align the first end of the first preset special effect with the origin hand, and control the pose of the first end based on the three-dimensional pose information of the origin hand; align the second end of the first preset special effect with the target hand, and in the process of controlling the pose of the second end based on the three-dimensional pose information of the target hand, control the text element corresponding to the first preset special effect to be displayed from the first end to the second end based on direction information and distance information;

[0216] The first end is either of the two ends, and the second end is the other end of the two ends besides the first end.

[0217] In an optional embodiment, the 3D text element effects processing module 1130 further includes:

[0218] The first preset effect determination unit is configured to perform a first preset effect that matches the two hands from the preset effect set;

[0219] The preset effects set includes multiple preset 3D text element effects corresponding to the number of hands. The multiple preset 3D text element effects include a first preset effect, which corresponds to two hands.

[0220] In an optional embodiment, the preset effects set includes a preset effects subset, which is a set of preset 3D text element effects corresponding to two hand gestures; the first preset effects determining unit includes:

[0221] Angle information determination unit is configured to determine the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity.

[0222] The first preset effect determination subunit is configured to determine the first preset effect from the preset effect subset based on the angle information.

[0223] In an optional embodiment, the first preset effect determination subunit is specifically configured to perform the following: when the angle information is greater than a first preset angle threshold, the preset horizontal effect in the preset effect subset is taken as the first preset effect.

[0224] In an optional embodiment, the first preset effect determination subunit is specifically configured to perform the following: when the angle information is less than or equal to a second preset angle threshold, the preset vertical effect in the preset effect subset is taken as the first preset effect.

[0225] In an optional embodiment, the preset effects subset includes a first vertical effect corresponding to the first origin hand category and a second vertical effect corresponding to the second origin hand category; the device further includes:

[0226] The target hand category determination module is configured to determine the target hand category to which the hand at the origin belongs;

[0227] The first preset effect determination subunit is specifically configured to determine the first preset effect from the first vertical effect and the second vertical effect based on the target hand category when the angle information is less than or equal to the second preset angle threshold.

[0228] In an optional embodiment, the preset effects subset includes a third vertical effect and a fourth vertical effect; the above-mentioned device further includes:

[0229] The relative position information determination module is configured to determine the relative position information of the origin hand in the real-time image. The relative position information indicates whether the origin hand is located to the left or right of the corresponding center position in the real-time image.

[0230] The first preset effect determination subunit is specifically configured to determine the first preset effect from the third and fourth vertical effects based on the relative position information when the angle information is less than or equal to the second preset angle threshold.

[0231] The third vertical effect corresponds to the first relative position information, which indicates that the origin hand is located to the left of the center position of the real-time image; the fourth vertical effect corresponds to the second relative position information, which indicates that the origin hand is located to the right of the center position of the real-time image.

[0232] In an optional embodiment, the preset effects subset includes alternating fifth and sixth vertical effects; the device further includes:

[0233] The occurrence count determination unit is configured to determine the first occurrence count corresponding to the fifth vertical effect and the second occurrence count corresponding to the sixth vertical effect during the camera device activation process;

[0234] The first preset effect determination subunit is specifically configured to, when the angle information is less than or equal to the second preset angle threshold, select the vertical effect that appears less frequently between the fifth and sixth vertical effects as the first preset effect based on the first and second occurrence counts.

[0235] In an optional embodiment, the hand determination module includes:

[0236] The relative position relationship determination unit is configured to determine the relative position relationship between two hands;

[0237] The hand segmentation unit is configured to perform the task of segmenting two hands into an origin hand and a target hand based on their relative positional relationship.

[0238] In an optional embodiment, the first end is the end closest to the starting text element corresponding to the first preset effect, and the second end is the other end of the two ends besides the first end.

[0239] In an optional embodiment, when at least one hand is a single hand, the 3D text element effect is a second preset effect, which is an effect with fixed text elements; the 3D text element effect processing module 1030 includes:

[0240] The second preset effect determination unit is configured to perform a second preset effect that matches a hand from the preset effect set;

[0241] The second special effects processing unit is configured to render a second preset special effect in the real-time screen and control the pose of the second preset special effect based on the three-dimensional pose information.

[0242] The preset effects set includes multiple preset 3D text element effects corresponding to the number of hands. The multiple preset 3D text element effects include a second preset effect, which corresponds to one hand.

[0243] In an optional embodiment, the three-dimensional pose information consists of the first pose information of at least one hand in the current frame corresponding to the real-time image and the second pose information of at least one hand in a preset number of frames prior to the current frame; the three-dimensional text element special effects processing module 1130 includes:

[0244] The weighted fusion processing unit is configured to perform weighted fusion processing on the first pose information and the second pose information to obtain the target pose information;

[0245] The data processing unit is configured to render 3D text element effects in real-time and control the pose of the 3D text element effects based on the target pose information.

[0246] In an optional embodiment, the above-described apparatus further includes:

[0247] The target special effects video generation module is configured to execute in response to video compositing instructions and generate a target special effects video based on the special effects footage during the camera device's activation process;

[0248] The special effects include real-time footage during the camera's activation and the rendered 3D text element effects within that footage.

[0249] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0250] Figure 12 This is a block diagram illustrating an electronic device for data processing according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 12 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a data processing method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0251] Those skilled in the art will understand that Figure 12The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0252] In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the data processing method as described in the embodiments of this disclosure.

[0253] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the data processing method of the present disclosure embodiments.

[0254] In an exemplary embodiment, a computer program product including instructions is also provided, which, when run on a computer, causes the computer to perform the data processing method of the present disclosure embodiments.

[0255] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0256] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0257] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A data processing method, characterized in that, include: In response to the camera device's activation command, hand detection is performed on the real-time footage captured by the camera device; If at least one hand is detected in the real-time image, the three-dimensional pose information of the at least one hand is obtained, and the three-dimensional pose information includes the position information and rotation information of the corresponding hand in three-dimensional space; In the real-time image, a 3D text element effect that matches the at least one hand is rendered, and the pose of the 3D text element effect is controlled based on the 3D pose information so that the pose of the 3D text element effect is synchronized and matched with the 3D pose of the corresponding hand in real time. When at least one hand is two hands, the 3D text element effect is a first preset effect. The first preset effect has two relatively movable ends, and the amount of text elements in the first preset effect is positively correlated with the distance between the two ends. Rendering the 3D text element effect matching the at least one hand in the real-time image, and controlling the pose of the 3D text element effect based on the 3D pose information, includes: determining the first preset effect matching the two hands from a preset effect set; wherein the preset effect set contains multiple preset 3D text element effects corresponding to the number of hands. The effect includes multiple preset 3D text element effects, including the first preset effect, where the number of hands corresponding to the first preset effect is two; the preset effect set includes a preset effect subset, which is a set of preset 3D text element effects corresponding to two hands; determining the first preset effect matching the two hands from the preset effect set includes: determining the origin hand and the target hand among the two hands; determining the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity; and determining the first preset effect from the preset effect subset based on the angle information. When the at least one hand is a single hand, the 3D text element effect is a second preset effect, which is an effect with fixed text elements; rendering the 3D text element effect that matches the at least one hand in the real-time image, and controlling the pose of the 3D text element effect based on the 3D pose information, includes: determining the second preset effect that matches the single hand from a preset effect set; wherein the preset effect set contains multiple preset 3D text element effects corresponding to the number of hands, the multiple preset 3D text element effects include the second preset effect, and the number of hands corresponding to the second preset effect is one.

2. The data processing method according to claim 1, characterized in that, The method further includes: Obtain the direction information between the origin hand and the target hand, as well as the distance information between the origin hand and the target hand; The step of rendering a 3D text element effect that matches the at least one hand in the real-time image, and controlling the pose of the 3D text element effect based on the 3D pose information, includes: In the real-time scene, the first end of the first preset special effect is aligned with the origin hand, and the pose of the first end is controlled based on the three-dimensional pose information of the origin hand; the second end of the first preset special effect is aligned with the target hand, and in the process of controlling the pose of the second end based on the three-dimensional pose information of the target hand, the text element corresponding to the first preset special effect is controlled to be displayed from the first end to the second end based on the direction information and the distance information; Wherein, the first end is either of the two ends, and the second end is the other end of the two ends besides the first end.

3. The data processing method according to claim 1, characterized in that, The step of determining the first preset effect from the preset effect subset based on the angle information includes: If the angle information is greater than the first preset angle threshold, the preset horizontal special effect in the preset special effect subset is taken as the first preset special effect.

4. The data processing method according to claim 1, characterized in that, The step of determining the first preset effect from the preset effect subset based on the angle information includes: If the angle information is less than or equal to the second preset angle threshold, the preset vertical effect in the preset effect subset is taken as the first preset effect.

5. The data processing method according to claim 4, characterized in that, The preset special effects subset includes a first vertical special effect corresponding to the first origin hand category and a second vertical special effect corresponding to the second origin hand category; the method further includes: Determine the target hand category to which the origin hand belongs; When the angle information is less than or equal to a second preset angle threshold, using the preset vertical effects in the preset effects subset as the first preset effect includes: If the angle information is less than or equal to the second preset angle threshold, the first preset effect is determined from the first vertical effect and the second vertical effect according to the target hand category.

6. The data processing method according to claim 4, characterized in that, The preset special effects subset includes a third vertical special effect and a fourth vertical special effect; the method also includes: The relative position information of the origin hand in the real-time image is determined, and the relative position information indicates that the origin hand is located to the left or right of the corresponding center position in the real-time image; When the angle information is less than or equal to a second preset angle threshold, using the preset vertical effects in the preset effects subset as the first preset effect includes: If the angle information is less than or equal to the second preset angle threshold, the first preset effect is determined from the third vertical effect and the fourth vertical effect based on the relative position information. The third vertical effect corresponds to the first relative position information, which indicates that the origin hand is located to the left of the center position of the real-time image; the fourth vertical effect corresponds to the second relative position information, which indicates that the origin hand is located to the right of the center position of the real-time image.

7. The data processing method according to claim 4, characterized in that, The preset special effects subset includes alternating fifth and sixth vertical special effects; the method further includes: Determine the first occurrence count of the fifth vertical effect and the second occurrence count of the sixth vertical effect during the operation of the camera device; When the angle information is less than or equal to a second preset angle threshold, using the preset vertical effects in the preset effects subset as the first preset effect includes: If the angle information is less than or equal to the second preset angle threshold, the vertical effect with the fewer occurrences among the fifth and sixth vertical effects is selected as the first preset effect based on the first occurrence count and the second occurrence count.

8. The data processing method according to claim 1, characterized in that, Determining the origin hand and the target hand among the two hands includes: Determine the relative positional relationship between the two hands; Based on the relative positional relationship, the two hands are divided into the origin hand and the target hand.

9. The data processing method according to claim 2, characterized in that, The first end is the end closest to the starting text element corresponding to the first preset effect, and the second end is the other end of the two ends besides the first end.

10. The data processing method according to claim 1, characterized in that, The step of rendering a 3D text element effect that matches the at least one hand in the real-time image, and controlling the pose of the 3D text element effect based on the 3D pose information, includes: The second preset effect is rendered in the real-time image, and the pose of the second preset effect is controlled based on the three-dimensional pose information.

11. The data processing method according to any one of claims 1 to 10, characterized in that, The three-dimensional pose information comprises the first pose information of at least one hand in the current frame corresponding to the real-time image and the second pose information of at least one hand in a preset number of frames prior to the current frame; the rendering of a three-dimensional text element effect matching the at least one hand in the real-time image, and the control of the pose of the three-dimensional text element effect based on the three-dimensional pose information, includes: The first pose information and the second pose information are weighted and fused to obtain the target pose information; The 3D text element effect is rendered in the real-time image, and the pose of the 3D text element effect is controlled based on the target pose information.

12. The data processing method according to any one of claims 1 to 10, characterized in that, The method further includes: In response to a video synthesis command, a target special effects video is generated based on the special effects footage during the camera device's activation process; The special effects include the real-time footage during the camera's activation process and the rendered 3D text element effects within the real-time footage.

13. A data processing apparatus, characterized in that, include: The hand detection module is configured to perform hand detection on the real-time images captured by the camera device in response to the camera device's activation command; The three-dimensional pose information acquisition module is configured to acquire the three-dimensional pose information of the at least one hand when the real-time image is detected to include at least one hand. The three-dimensional pose information includes the position information and rotation information of the corresponding hand in three-dimensional space. The 3D text element effect processing module is configured to render a 3D text element effect that matches the at least one hand in the real-time screen, and control the pose of the 3D text element effect based on the 3D pose information so that the pose of the 3D text element effect is synchronized and matched with the 3D pose of the corresponding hand in real time. When at least one hand is two hands, the three-dimensional text element effect is a first preset effect, the first preset effect has two relatively moving ends, and the amount of text elements in the first preset effect is positively correlated with the distance between the two ends; The data processing device further includes: a hand determination module, configured to determine the origin hand and the target hand among the two hands; the three-dimensional text element effect processing module includes: a first preset effect determination unit, configured to determine the first preset effect matching the two hands from a preset effect set; wherein, the preset effect set includes multiple preset three-dimensional text element effects corresponding to the number of hands, the multiple preset three-dimensional text element effects include the first preset effect, and the number of hands corresponding to the first preset effect is two; the preset effect set includes a preset effect subset, the preset effect subset being a set of preset three-dimensional text element effects corresponding to the number of hands of two; the first preset effect determination unit includes: an angle information determination unit, configured to determine the angle information of the line connecting the origin hand and the target hand relative to the direction of world gravity; the first preset effect determination subunit is configured to determine the first preset effect from the preset effect subset based on the angle information; When at least one hand is a single hand, the three-dimensional text element effect is a second preset effect, which is an effect with fixed text elements; the three-dimensional text element effect processing module includes: a second preset effect determination unit, configured to determine, from a preset effect set, a second preset effect that matches the single hand; wherein, the preset effect set contains multiple preset three-dimensional text element effects corresponding to the number of hands, the multiple preset three-dimensional text element effects include the second preset effect, and the number of hands corresponding to the second preset effect is one.

14. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the data processing method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the data processing method as described in any one of claims 1 to 12.