Data processing method and device, equipment, medium and product

By performing facial and body detection processing on images taken by a monocular camera and updating the body detection results to ensure that the overlapping area is greater than the threshold, the problem of difficulty in obtaining motion capture data in multi-object scenes is solved, and more accurate motion capture data acquisition and virtual image restoration is achieved.

CN120014683APending Publication Date: 2025-05-16BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510072367.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively obtain motion capture data, especially in multi-object scenarios, where the face and body detection results may not belong to the same object, resulting in increased difficulty in obtaining motion capture data.

Method used

The face and body information of the object in the image is determined by performing facial and body detection processing on the images taken by the monocular camera. If the overlap area between the face and body area indicated by the detection result is not greater than the preset threshold, the body detection result is updated to ensure that the overlap area is greater than the threshold, thereby determining the motion capture data.

Benefits of technology

Improve the accuracy of obtaining motion capture data in multi-object scenes, ensuring that the motion capture data can more accurately represent the actions of objects in the image, thereby improving the action restoration effect of virtual images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014683A_ABST
    Figure CN120014683A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, equipment, a medium and a product, and the method comprises the steps: obtaining a first image used for describing some objects, carrying out the face detection processing and body detection processing of the image, and obtaining a face detection result and a body detection result, when it is detected that the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result is not larger than a preset area threshold value, the body detection result is updated; the method comprises the following steps: updating a body detection result of an image to enable an overlapping area between a body region indicated by the updated body detection result and a face region indicated by a face detection result to be greater than a preset area threshold value, and determining motion capture data corresponding to the image according to the updated body detection result and the face detection result, therefore, the motion capture data can more accurately represent the motion of a certain object presented by the image, so that the motion capture data can be determined from one image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, device, equipment, medium, and product. Background Art

[0002] Motion capture, abbreviated as motion capture, is a technology used to record and process the movements of an object (such as a person, animal, or some object) so that it can be widely used in many fields such as entertainment, sports, computer vision, and robotics.

[0003] However, how to obtain motion capture data has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] In order to solve the above technical problems, the present application provides a data processing method, device, equipment, medium, and product.

[0005] In order to achieve the above objectives, the technical solutions provided by this application are as follows:

[0006] The present application provides a data processing method, the method comprising: acquiring a first image, the first image being used to describe at least one object; performing face detection processing on the first image to obtain a face detection result, the face detection result being used to indicate face information of the object, the face information including a face area, and performing body detection processing on the first image to obtain a body detection result, the body detection result being used to indicate body information of the object, the body information including a body area, the body including a head; in response to an overlapping area between a body area indicated by the body detection result and a face area indicated by the face detection result being not greater than a preset area threshold, updating the body detection result so that an overlapping area between a body area indicated by the updated body detection result and a face area indicated by the face detection result is greater than a preset area threshold, and determining motion capture data corresponding to the first image based on the updated body detection result and the face detection result.

[0007] In one possible implementation, the method further includes: in response to an overlapping area between a body area indicated by the body detection result and a facial area indicated by the facial detection result being greater than a preset area threshold, determining motion capture data corresponding to the first image based on the body detection result and the facial detection result.

[0008] In one possible implementation, the process of determining the body detection result includes: acquiring a first area, where the first area is the body area indicated by the body detection result; performing bone point detection processing on the first image based on the first area to obtain bone point information, where the bone point information is used to describe the state of each bone point in the first area; and determining the body detection result based on the first area and the bone point information.

[0009] In one possible implementation, if a multiplexing condition is met, the first region is obtained by multiplexing a body region recognition result of a second image, the second image is captured by a monocular camera earlier than the first image is captured by the monocular camera, the motion capture data corresponding to the second image is determined earlier than the motion capture data corresponding to the first image, and during the determination of the motion capture data corresponding to the second image, body region recognition processing is performed on the second image to obtain a body region recognition result of the second image; if the multiplexing condition is not met, the first region is obtained by performing body region recognition processing on the first image.

[0010] In one possible implementation, if the first region is obtained by reusing the body region recognition result of the second image, the process of determining the updated body detection result includes: performing body region recognition processing on the first image to obtain the second region; performing bone point detection processing on the first image based on the second region to obtain a processing result, and the processing result is used to describe the status of each bone point in the second region; and determining the updated body detection result based on the second region and the processing result.

[0011] In one possible implementation, after obtaining the face detection result, the method further includes: if the area of ​​the face region indicated by the face detection result does not exceed a first threshold, adjusting the face detection result using preset face information, the adjusted face detection result being the preset face information, and the motion capture data being determined based on the adjusted face detection result.

[0012] In one possible implementation, after obtaining the body detection result, the method further includes: if the area of ​​the body region indicated by the body detection result does not exceed a second threshold, adjusting the body detection result using preset body information, the adjusted body detection result being the preset body information, and the motion capture data being determined based on the adjusted body detection result.

[0013] In one possible implementation, the body detection result includes a hand detection result; after obtaining the body detection result, the method further includes at least one of the following steps: if the hand detection result is used to indicate that the hand is in an invisible state, the hand detection result in the body detection result is replaced by preset hand information, and the replaced body detection result includes the preset hand information; if the hand detection result is used to indicate that the hand is in a visible state, and there is a gesture matching the gesture indicated by the hand detection result among the pre-set multiple candidate gestures, the hand detection result in the body detection result is replaced by hand information pre-configured for the matching gesture, and the replaced body detection result includes the configured hand information; the motion capture data is determined based on the replaced body detection result.

[0014] In one possible implementation, the method further includes: performing interpolation processing on the motion capture data corresponding to the first image and the motion capture data corresponding to the third image according to a preset frame rate to obtain a motion capture data sequence, the frame rate of the motion capture data sequence being the preset frame rate, the image sequence captured by the monocular camera including the first image and the third image, the arrangement position of the third image in the image sequence being adjacent to the arrangement position of the first image in the image sequence; and driving the virtual image using the motion capture data sequence.

[0015] In one possible implementation, the method further includes: determining a usage time of each motion capture data in the motion capture data sequence based on a time when the first image is captured, a time when the third image is captured, and a preset delay time; using the motion capture data sequence to drive a virtual image includes: for any motion capture data in the motion capture data sequence, using the motion capture data to drive the virtual image at the usage time of the motion capture data.

[0016] In a possible implementation manner, a sum of a time duration consumed in a process of determining the motion capture data corresponding to the first image and a time duration consumed in a process of determining the motion capture data corresponding to the third image is less than the preset delay duration.

[0017] In one possible implementation, the motion capture data corresponding to the first image includes the motion capture data corresponding to each part under the first image, and the motion capture data corresponding to the third image includes the motion capture data corresponding to each part under the third image, and the motion capture data sequence is used to indicate the interpolation processing results of each part; for any of the parts, the interpolation processing result of the part is obtained by interpolating the motion capture data corresponding to the part under the first image and the motion capture data corresponding to the part under the third image based on the interpolation parameters corresponding to the part, and the interpolation parameters are used to indicate the smoothness of the interpolation processing result of the part.

[0018] In one possible implementation, for any of the parts, the motion capture data corresponding to the part under the first image and the motion capture data corresponding to the part under the third image are both determined using a target detection algorithm, and the target detection algorithm is used to implement the face detection processing or the body detection processing, and the interpolation parameters corresponding to the part are determined based on the performance of the target detection algorithm and the variation amplitude of the motion capture data corresponding to the part under the image sequence.

[0019] In a possible implementation manner, the first image is a two-dimensional image captured in real time using a monocular camera.

[0020] The present application provides a data processing device, including: an acquisition unit, used to acquire a first image, the first image is used to describe at least one object; a detection unit, used to perform face detection processing on the first image to obtain a face detection result, the face detection result is used to indicate face information of the object, the face information includes a face area, and perform body detection processing on the first image to obtain a body detection result, the body detection result is used to indicate body information of the object, the body information includes a body area, and the body includes a head; an update unit, used to update the body detection result in response to the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result being not greater than a preset area threshold, so that the overlapping area between the body area indicated by the updated body detection result and the face area indicated by the face detection result is greater than the preset area threshold, and determine the motion capture data corresponding to the first image based on the updated body detection result and the face detection result.

[0021] The present application provides an electronic device, which includes: a processor and a memory; the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device performs the data processing method provided by the present application.

[0022] The present application provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes the data processing method provided by the present application.

[0023] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the data processing method provided by the present application.

[0024] Compared with the related art, this application has at least the following advantages:

[0025] In the technical solution provided by the present application, after obtaining a first image for describing some objects (such as a two-dimensional image for describing multiple people taken in real time using a monocular camera), first, face detection processing is performed on the first image to obtain a face detection result, so that the face detection result is used to indicate the facial information of an object (such as the most conspicuous object) among these objects, such as facial expression coefficient, facial area, key point information of the face, and other information, and body detection processing is performed on the first image to obtain a body detection result, so that the body detection result is used to indicate the body information of an object among these objects, such as body area, three-dimensional position, rotation and confidence of each bone point of the whole body, so that when it is detected that the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result is not greater than a preset area threshold, it can be determined that the body detection result is the body information of the object. The body described by the first image and the face described by the face detection result do not belong to the same object, so the body detection result is updated so that the overlapping area between the body area indicated by the updated body detection result and the face area indicated by the face detection result is greater than a preset area threshold, so that the body described by the updated body detection result and the face described by the face detection result belong to the same object; then, based on the updated body detection result and the face detection result, the motion capture data corresponding to the first image is determined, so that the motion capture data can more accurately represent the action of a certain object presented by the first image, so that when the motion capture data is used to drive the virtual image later, the virtual image can accurately restore the action presented by the object in the first image, so that it is possible to determine the motion capture data from one image, thereby effectively reducing the difficulty of obtaining the motion capture data. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0027] Figure 1 A flowchart of a data processing method provided in an embodiment of the present application;

[0028] Figure 2 A schematic diagram of an update process provided in an embodiment of the present application;

[0029] Figure 3 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;

[0030] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] Research has found that in some scenarios, in order to better reduce the difficulty of obtaining motion capture data, motion capture data can be obtained by analyzing the two-dimensional images taken by a monocular camera.

[0032] Research has also found that one way to implement the motion capture data acquisition process shown in the previous paragraph can be: after obtaining the two-dimensional image taken by a monocular camera, first perform face detection and body detection on the image to obtain face detection results and body detection results; then use these two detection results as motion capture data to drive the virtual image.

[0033] Research has also found that because the face detection processing and the body detection processing are independent of each other in the implementation shown in the previous paragraph, when multiple objects appear in the image, the face determined by the face detection processing may not belong to the same object as the body determined by the body detection processing, thus affecting the accuracy.

[0034] Based on the above research, in order to better improve the accuracy, the present application provides a data processing method, which includes: after obtaining a first image used to describe some objects (such as a two-dimensional image used to describe multiple people taken in real time using a monocular camera), first, performing face detection processing on the first image to obtain a face detection result, and performing body detection processing on the first image to obtain a body detection result, so that when it is detected that the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result is not greater than a preset area threshold, it can be determined that the body described by the body detection result and the face described by the face detection result do not belong to the same object, so the body detection result is updated so that the updated body detection result The method further comprises the step of: determining, according to the first image, motion capture data corresponding to the first image, determining the body area indicated by the updated body detection result and the face area indicated by the face detection result greater than a preset area threshold, so that the body described by the updated body detection result and the face described by the face detection result belong to the same object; then, based on the updated body detection result and the face detection result, determining the motion capture data corresponding to the first image, so that the motion capture data can more accurately represent the action of an object presented by the first image, so that when the motion capture data is used to drive the virtual image later, the virtual image can accurately restore the action presented by the object in the first image, thereby effectively overcoming the defect caused by the detected face and body belonging to different objects in a multi-object scene, thereby facilitating improving the accuracy.

[0035] In addition, the present application does not limit the execution subject of the data processing method. For example, the method can be applied to a terminal device or a server. For another example, the method can also be implemented by means of a data interaction process between a terminal device and a server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server, or a cloud server.

[0036] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0037] In order to better understand the technical solution provided by this application, the data processing method provided by this application is described below with reference to some drawings. Figure 1 As shown, the data processing method provided in the embodiment of the present application includes the following S1-S3.

[0038] S1: Acquire a first image, where the first image is used to describe at least one object.

[0039] The first image refers to a two-dimensional image that needs to be processed for motion capture data determination (such as Figure 2 The second frame image shown in the figure) is used to describe the state of at least one object in the two-dimensional space, such as the state of action. It should be noted that the at least one object is the foreground of the first image; and the present application does not limit the implementation method of the object, for example, it can be implemented by a person, an animal, or an object.

[0040] In addition, the present application does not limit the method for acquiring the first image. For example, it may be: using a monocular camera, such as a monocular RGB (Red Green Blue) camera, to photograph at least one object to obtain the first image.

[0041] For example, in some scenarios, such as converting a captured image into motion capture data in real time, the process of acquiring the above-mentioned first image can be: acquiring a two-dimensional image captured in real time by a monocular camera as the first image, and using the data processing method provided in the present application to timely convert the image into motion capture data to meet the real-time requirements in these scenarios.

[0042] S2: performing face detection processing on the first image to obtain a face detection result, where the face detection result is used to indicate face information of an object, where the face information includes a face area; and performing body detection processing on the first image to obtain a body detection result, where the body detection result is used to indicate body information of an object, where the body information includes a body area, where the body includes a head.

[0043] The face detection result is obtained by performing face detection processing on the first image, so that the face detection result can describe an object (such as Figure 2 The facial state of the object 2) shown, such as the facial expression of the object, the position of each key point in the face of the object, the position of the face of the object in the image, etc.

[0044] It can be seen that for the face detection result determined from the first image, the result is used to indicate the face information of an object (such as the most conspicuous object) in at least one object described by the image, so that the face information can describe the state of the face of the object in the image. Among them, the face information may include one or more of the face area, the face expression coefficient and the face key point information. The face area is used to describe the position of the face of the object in the image; and the face area can be implemented by a face mask area, or by other area representation methods (such as a face bounding box), which is not limited in this application. The face expression coefficient is used to describe the expression of the face of the object in the image; and the face expression coefficient can be implemented by BS (Blend Shape) information, or by other expression representation methods, which is not limited in this application. The face key point information is used to describe the position of each key point in the face of the object in the image; and the face key point information can be implemented by any face key point representation method, which is not limited in this application.

[0045] In addition, the present application does not limit the implementation method of the above-mentioned face detection processing. For example, it can adopt any method that can perform face detection processing on an image, such as implementing it with the help of a machine learning model with face detection function.

[0046] For another example, in some scenarios, such as multi-object scenarios, the above-mentioned face detection result determination process can be: after acquiring the first image, first detect the face information of each object from the image; then filter out the face information that meets certain conditions from these face information, so that the face area described by the "face information that meets certain conditions" is as close to the center of the image as possible, and the area of ​​the face area described by the "face information that meets certain conditions" is as large as possible, so that the face described by the "face information that meets certain conditions" is as conspicuous as possible, so the "face information that meets certain conditions" is used as the face detection result. Among them, the condition is used to indicate the constraints that the most conspicuous face needs to meet.

[0047] The body detection result is obtained by performing body detection processing on the first image, so that the body detection result can describe a certain object (such as Figure 2 The body state of the object 1) shown, such as the three-dimensional position of each bone point in the body of the object, the rotation of each bone point in the body of the object, the position of the body of the object in the image, etc. It should be noted that the body includes parts such as the head, torso and limbs.

[0048] It can be seen that for the body detection result determined for the first image, the result is used to indicate the body information of an object (such as the most conspicuous object or the object that has been tracked and locked) in at least one object described by the image, so that the body information can describe the state of the object's body in the image. Among them, the body information may include body area and bone point information. The body area is used to describe the position of the object's body in the image; and the body area can be implemented using a body bounding box (Box) area, or can be implemented using other area representation methods (such as a body mask area), which is not limited in this application. The bone point information is used to describe the state of each bone point in the object's body in the image, such as the three-dimensional position of each bone point, the rotation of each bone point, the confidence of each bone point, whether each bone point is out of the picture, and other states.

[0049] It should be noted that, for any skeleton point, the three-dimensional position of the skeleton point is used to describe the position of the skeleton point in the three-dimensional space; the confidence of the skeleton point is used to describe whether the information obtained from the detection of the skeleton point (such as three-dimensional position, rotation, whether it is out of the frame, etc.) is accurate, so that the confidence can indicate to a certain extent whether the skeleton point is visible in the first image; whether the skeleton point is out of the frame is used to indicate whether the skeleton point is located in the first image, so that it can indicate whether the shooting range circled by the camera when the first image is shot by a monocular camera includes the skeleton point.

[0050] In addition, since the body may include multiple parts (such as the head, hands, etc.), in order to better improve the effect, the above body detection results may include detection results of various parts, such as hand detection results, head detection results, etc. Among them, the i-th part detection result is used to describe the state of the i-th part in the first image, i is a positive integer, i≤N, and N represents the number of parts in the multiple parts.

[0051] In addition, the present application does not limit the implementation method of the above-mentioned body detection processing. For example, it can adopt any method that can perform body detection processing on an image, such as implementing it with the help of a machine learning model with body detection function.

[0052] For example, in some scenarios, such as multi-object scenarios, the determination process of the above body detection result can be: after acquiring the first image, first detect the body information of each object from the image; then select the body information that meets certain conditions from these body information, so that the body area described by the "body information that meets certain conditions" is as close to the center of the image as possible, and the area of ​​the body area described by the "body information that meets certain conditions" is as large as possible, and the number of valid points of the body described by the "body information that meets certain conditions" in the image is as large as possible, so that the body described by the "body information that meets certain conditions" is as conspicuous as possible, so the "body information that meets certain conditions" is used as the body detection result. Among them, the condition is used to indicate the constraints that the most conspicuous body needs to meet. Among them, the valid points refer to some relatively important bone points that are pre-set, such as three bone points that can describe the facial state as completely as possible, and multiple bone points (such as six bone points) that can describe other parts except the head as completely as possible; and the valid points can be manually set according to the actual application scenario.

[0053] In addition, in order to better improve the detection effect, the above body detection process may include the following steps 11 to 13.

[0054] Step 11: Acquire a first region, so that the first region is the body region indicated by the above body detection result, so that the first region can indicate the position of the object requiring body detection processing in the first image.

[0055] It should be noted that the present application does not limit the implementation of the above step 11. For example, it can be specifically as follows: after acquiring the first image, first perform body region recognition processing on the image to obtain the body region of each object; then search for a body region that meets certain conditions (such as being as close to the center of the image as possible, having as large an area as possible, and having as many valid points as possible) from these body regions as the first region. It should be noted that the present application does not limit this condition. For example, the condition can be set according to the actual application scenario.

[0056] Step 12: Based on the first region, skeleton point detection is performed on the first image to obtain skeleton point information, so that the skeleton point information is used to indicate the state of each skeleton point in the first region of the image.

[0057] Among them, the skeleton point detection process is used to detect the status of the skeleton points in a certain area in real time, such as position, rotation, confidence, whether out of the picture, etc.

[0058] In addition, the present application does not limit the implementation method of the above-mentioned skeleton point detection processing. For example, it can adopt any method that can detect the state of skeleton points in real time from an image, such as implementing it with the help of a machine learning model with skeleton point state detection function.

[0059] The skeleton point information is used to describe the state of each skeleton point in the body of the object circled by the first area in the first image, such as the three-dimensional position of each skeleton point, the rotation of each skeleton point, the confidence of each skeleton point, whether each skeleton point is out of the picture, etc.

[0060] In addition, in order to better improve the accuracy, the above-mentioned step 12 can be specifically as follows: after obtaining the above-mentioned first area, firstly perform body key point recognition processing within the range circled by the area in the first image to obtain key point information, so that the key point information can represent the body state of the object circled by the area in the image; then, based on the key point information, solve the preset three-dimensional skeleton (such as the Ybot skeleton) to obtain the state information of each bone point in the preset three-dimensional skeleton, such as rotation angle and other information; then, use these state information to drive the preset three-dimensional model to obtain bone point information, so that the bone point information can more accurately describe the body state of the object.

[0061] It should be noted that the present application is not limited to a preset three-dimensional model. For example, it can be implemented using any statistical three-dimensional human body model, such as the SMPL (Skinned Multi-Person Linear Model) model.

[0062] Step 13: Determine a body detection result corresponding to the first image based on the first area and the skeleton point information, so that the body detection result can represent the body state of the object encircled by the area in the first image.

[0063] It should be noted that the present application does not limit the implementation method of the above-mentioned step 13. For example, it can specifically be: aggregating the above-mentioned first area and the above-mentioned bone point information to obtain a body detection result corresponding to the first image, so that the body detection result can represent the body state of the object circled by the area in the first image.

[0064] Based on the relevant contents of steps 11 to 13 above, it can be known that for some scenarios, after acquiring the first image, the most conspicuous body area is first identified from the image; then, the image is subjected to skeleton point detection processing according to the area to obtain skeleton point information; then, the area and the information are aggregated to obtain the body detection result corresponding to the image.

[0065] S3: In response to the overlapping area between the body area indicated by the body detection result and the facial area indicated by the face detection result being not greater than a preset area threshold, the body detection result is updated so that the overlapping area between the body area indicated by the updated body detection result and the facial area indicated by the face detection result is greater than the preset area threshold, and based on the updated body detection result and the face detection result, the motion capture data corresponding to the first image is determined.

[0066] In the present application, for the body detection results and face detection results obtained by two completely independent processes, if the overlapping area between the body region indicated by the body detection result and the face region indicated by the face detection result is less than or equal to a preset area threshold (such as 50% of the area of ​​the face region), it can be determined that the intersection between the body region and the face region is relatively small, so that it can be determined that the relative position relationship between the body described by the body detection result and the face described by the face detection result does not satisfy the relative position relationship between the body and the face of the same object, and further it can be determined that the object encircled by the body region is different from the object encircled by the face region. object, so it can be determined that the body detection result and the face detection result do not match, so the body detection result can be updated, so that the overlapping area between the body area indicated by the updated body detection result and the face area indicated by the face detection result is greater than the preset area threshold, so that the updated body detection result and the face detection result belong to the same object, so that the motion capture data corresponding to the first image can be determined based on the updated body detection result and the face detection result, so that the motion capture data can more accurately describe the action state of an object in the first image, so that the virtual image driven by the motion capture data can accurately restore the action of the object.

[0067] It should be noted that the present application does not limit the implementation method of the virtual image. For example, the virtual image can be implemented using a three-dimensional deformable model so that the virtual image driven by the motion capture data can accurately restore the action shot for a certain object in a three-dimensional space. For another example, the virtual image can be implemented using a two-dimensional deformable model so that the virtual image driven by the motion capture data can accurately restore the action shot for the object in a two-dimensional space.

[0068] In addition, the present application does not limit the implementation method of the step of "updating physical examination results" in S3 above. For example, it can be implemented in a human-computer interaction manner to enable relevant personnel to correct the current physical examination results through some interactive operations.

[0069] For example, in order to improve flexibility, the updated body detection result can be obtained by re-performing body detection processing on the first image based on the facial area indicated by the face detection result, so as to ensure that the overlapping area between the body area indicated by the updated body detection result and the facial area indicated by the face detection result is greater than a preset area threshold.

[0070] Based on the relevant contents of S1 to S3 above, it can be known that in the determination scheme of motion capture data provided by the present application, after obtaining a first image for describing some objects, first, face detection processing is performed on the image to obtain a face detection result, and body detection processing is performed on the image to obtain a body detection result, so that when it is detected that the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result is not greater than a preset area threshold, it can be determined that the body described by the body detection result and the face described by the face detection result do not belong to the same object, so the body detection result is updated so that the body area indicated by the updated body detection result overlaps with the face area indicated by the face detection result. The overlapping area between the facial regions indicated by the detection results is greater than a preset area threshold, so that the body described by the updated body detection result and the face described by the face detection result belong to the same object; then, based on the updated body detection result and the face detection result, the motion capture data corresponding to the image is determined, so that the motion capture data can more accurately represent the action of a certain object presented by the image, so that when the motion capture data is used to drive the virtual image later, the virtual image can accurately restore the action presented by the object in the image, which can effectively overcome the defects caused by the face and body detected belonging to different objects in a multi-object scene, thereby facilitating improving accuracy.

[0071] Through research, it is found that for some algorithms with body detection function, such as skeleton detection algorithm with detection function, the working principle of the algorithm is: when starting the algorithm, only the image of the current frame (such as Figure 1 The first frame image shown in the figure is used for body area recognition processing to obtain the body area required for detection and locking when starting this algorithm, so that the body area can be directly reused when processing subsequent images, so as to achieve skeleton point detection and processing for the same object during one operation of the algorithm, so as to achieve the purpose of tracking the changes of the skeleton points of the object.

[0072] Based on the above research, it can be known that in some scenarios, such as converting a real-time acquired image stream into motion capture data, in order to better improve the effect, the present application also provides a possible implementation of the above-mentioned first area, in which the first area can at least meet the following constraints: if the reuse condition is met (such as the condition that the Skeleton detection algorithm has been started and used from the second image before the first image is acquired to determine the body area required for tracking and locking when the algorithm is started), then the first area is obtained by reusing the body area recognition result of the second image (such as Figure 2However, if the reuse condition is not met, the first area is obtained by performing body area recognition processing on the first image. The moment when the second image is captured by the monocular camera is earlier than the moment when the first image is captured by the monocular camera, the process of determining the motion capture data corresponding to the second image is earlier than the process of determining the motion capture data corresponding to the first image, and the body area recognition result of the second image is obtained by performing body area recognition processing on the second image during the process of determining the motion capture data corresponding to the second image.

[0073] It should be noted that the above-mentioned multiplexing conditions can be set according to the actual application scenario. For example, it may specifically include: before acquiring the first image, the Skeleton detection algorithm has been started and used to determine from the second image the body area required for tracking and locking when starting this algorithm, and the Skeleton detection algorithm has not been restarted for the first image.

[0074] It should also be noted that the present application does not limit the implementation method of the body region recognition processing implemented by the Skeleton detection algorithm. For example, the body region recognition processing can at least meet the following constraints: if multiple bones are detected from an image, it is necessary to select a most conspicuous bone from the multiple bones based on the position of each bone in the image, the area occupied by each bone in the image, and the number of valid points of each bone appearing in the image, so that the most conspicuous bone is as close to the center of the image as possible, the most conspicuous bone occupies as large an area as possible, and the number of valid points of the most conspicuous bone appearing in the image is greater than a preset value (such as 6), so as to ensure that all valid points in the most conspicuous bone used to describe the face appear in the image, thereby ensuring that the object indicated by the most conspicuous bone can display the face normally in the image.

[0075] Based on the above three sections, it can be seen that after obtaining the first image captured by the monocular camera in real time, the process of determining the body detection result corresponding to the first image may at least include the following:

[0076] If the first image is the first frame image captured by the monocular camera (such as Figure 2 If the first frame image is the one shown in the figure, it can be determined that the Skeleton detection algorithm has not been started yet, so it can be determined that the multiplexing condition is not met, so the Skeleton detection algorithm can be started for the first image, so as to use the algorithm to perform body region recognition processing on the first image to obtain the first region, so that the first region can accurately represent the more conspicuous body region detected from the first image, so that the body detection result corresponding to the first image can be determined based on the first region later;

[0077] If the first image is a non-first frame image taken by the monocular camera (such as Figure 2 The second frame image shown in FIG. 1 ), and before acquiring the first image, the Skeleton detection algorithm has been started and used to detect the second image (such as Figure 2 If the body region to be tracked and locked in the current algorithm startup is determined in the first frame image shown in the figure, it can be determined that the reuse condition is met, so the body region that has been detected from the second image in the historical time period can be directly reused as the first region, so that the first region can represent the body region that still needs to be tracked and locked as of the first image, so that the body detection result corresponding to the first image can be determined based on the first region later;

[0078] If the first image is not the first frame image taken by the monocular camera, but the Skeleton detection algorithm needs to be restarted for the first image, it can be determined that the multiplexing condition is not met. Therefore, the algorithm is used to perform body area recognition processing on the first image to obtain the first area, so that the first area can accurately represent the more conspicuous body area detected from the first image, so that the body detection result corresponding to the first image can be determined based on the first area later.

[0079] Based on the above-mentioned Skeleton detection algorithm, it can be known that for some scenes, such as multi-person scenes, when the Skeleton detection algorithm is used in these scenes to implement body detection processing, the Skeleton detection algorithm includes two stages of body region recognition processing and bone point detection processing that are performed sequentially, so that after obtaining a first image taken by a monocular camera in real time, if it is detected that the Skeleton detection algorithm is started (or restarted) for the first image, the algorithm can be used to sequentially perform body region recognition processing and bone point detection processing on the first image to obtain the first region and the bone point information within the first region; however, if it is detected that Before acquiring the first image, the Skeleton detection algorithm has been started and used to determine the body area that needs to be tracked and locked when starting the algorithm this time from a frame image (such as the second image) historically taken by the monocular camera. It can be determined that there is no need to perform the body area recognition processing under the first image, and only the skeleton point detection processing needs to be performed. Therefore, the locked body area can be regarded as the first area, and based on the locked body area, the skeleton point detection processing is performed on the first image to obtain the skeleton point information within the first area, so that the first area and the skeleton point information can be summarized later to obtain the body detection result corresponding to the first image, which is conducive to improving efficiency.

[0080] After research, it was found that for the processing process shown in the previous paragraph, if the first area is obtained by reusing the body area recognition result of the second image, it can be determined that the first area refers to the body area that needs to be tracked and locked obtained by processing the historical image (such as the second image) under the current startup of the Skeleton detection algorithm, so that it can be determined that the first area may be correct or inaccurate. Therefore, when it is detected that the overlapping area between the first area and the face area indicated by the face detection result determined from the first image is not greater than the preset area threshold, it can be accurately determined that the first area is incorrect, so the problem can be solved by restarting the Skeleton detection algorithm.

[0081] Based on the above research, it can be known that in a possible implementation, when the above-mentioned first area is obtained by reusing the body area recognition result of the second image (that is, the body area that has been determined within the historical time period is directly reused under the first image), the determination process of the above-mentioned updated body detection result may include the following steps 21-23.

[0082] Step 21: Perform body region recognition processing on the first image to obtain a second region, so that the second region is used to indicate the body region detected from the first image, so that the object circled by the body region indicated by the second region in the first image belongs to an object with a relatively prominent body whose face can be displayed normally, and further the overlapping area between the second region and the face region indicated by the face detection result determined from the first image is greater than a preset area threshold, so that the two regions belong to the same object.

[0083] Step 22: Based on the second area, the first image is subjected to skeleton point detection processing to obtain a processing result, so that the processing result is used to indicate the status of each skeleton point in the second area in the first image, such as the three-dimensional position of each skeleton point, the rotation of each skeleton point, the confidence of each skeleton point, whether each skeleton point is out of the picture, etc.

[0084] Step 23: Based on the second area and the processing result, determine an updated body detection result, so that the updated body detection result includes the second area and the processing result, so that the updated body detection result is the result obtained by restarting the Skeleton detection algorithm for the first image, thereby making the updated body detection result more accurate.

[0085] Based on the relevant contents of steps 21 to 23 above, it can be known that for the first image (such as Figure 2 For example, if the body regions determined by the Skeleton detection algorithm in the historical time period are reused (such as the second frame image shown in FIG. Figure 2 The body detection result corresponding to the first image is determined by the method of the area 1 shown in FIG. Figure 2 1 shown in the body detection result), then after detecting the body area indicated by the body detection result (such as Figure 2 The area 1 shown in FIG. 1 is the same as the face detection result determined from the first image (eg, Figure 2 When the overlapping area between the facial areas indicated by the facial detection result of object 2 shown in the figure is not greater than the preset area threshold, it can be determined that the two detection results belong to different objects. Therefore, the Skeleton detection algorithm can be restarted for the first image to enable the algorithm to detect a body area that is as complete and conspicuous as possible from the first image, so that the updated body detection result determined based on the body area is as accurate as possible, which is conducive to improving the detection effect.

[0086] In addition, in order to better improve the determination effect of motion capture data, the present application provides a possible implementation of a data processing method, in which the method may include the following steps 31 to 35.

[0087] Step 31: Acquire a first image, where the first image is used to describe at least one object.

[0088] It should be noted that for the relevant content of step 31, please refer to the relevant content of S1 above.

[0089] Step 32: Performing face detection processing on the first image to obtain a face detection result, and performing body detection processing on the first image to obtain a body detection result.

[0090] It should be noted that for the relevant content of step 32, please refer to the relevant content of S2 above.

[0091] Step 33: Determine whether the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result is greater than a preset area threshold. If so, execute the following step 34; if not, execute the following step 35.

[0092] Step 34: If the overlapping area between the body area indicated by the above body detection result and the facial area indicated by the above face detection result is greater than a preset area threshold, the motion capture data corresponding to the first image is determined based on the body detection result and the facial detection result, so that the motion capture data includes the body detection result and the facial detection result.

[0093] In the present application, for the body detection results and face detection results obtained by two completely independent processes, if the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result is greater than a preset area threshold, it can be determined that the intersection between the body area and the face area is relatively large, so that it can be determined that the relative position relationship between the body circled by the body area in the first image and the face circled by the face area in the first image satisfies the relative position relationship between the body and the face of the same object, and further it can be determined that the object circled by the body area and the object circled by the face area are the same object, so it can be determined that the body detection result matches the face detection result, so the motion capture data corresponding to the first image can be obtained by summarizing the body detection result and the face detection result.

[0094] Step 35: If the overlapping area between the body area indicated by the above body detection result and the facial area indicated by the above face detection result is not greater than the preset area threshold, the body detection result is updated, the overlapping area between the body area indicated by the updated body detection result and the facial area indicated by the face detection result is greater than the preset area threshold, and the motion capture data corresponding to the first image is determined based on the updated body detection result and the face detection result.

[0095] It should be noted that for the relevant content of step 35, please refer to the relevant content of S3 above.

[0096] Based on the relevant contents of the above steps 31 to 35, it can be known that for an image used to describe multiple objects, if the body detection result determined from the image using the Skeleton detection algorithm and the body detection result determined from the image using a certain face detection method both belong to the same object, it can be determined that the two detection results are relatively accurate, so the two detection results can be directly used to determine the motion capture data corresponding to the image; however, if the two detection results belong to different objects respectively, it can be determined that the two detection results do not match, and thus it can be determined that the body detection result determined by the body area determined by reusing the history is inaccurate, so the body detection result can be updated by restarting the Skeleton detection algorithm so that the updated body detection result and the body detection result both belong to the same object, thereby overcoming the defect caused by the detection results determined by two independent processing processes not belonging to the same object, thereby helping to better improve the accuracy.

[0097] Through research, it is found that for the above-mentioned face detection results, when the face described by the face detection results occupies a larger area in the first image, it can be determined that the visible area (also known as the effective area) of the face is relatively large, and thus it can be determined that the facial state described by the face detection results is relatively accurate; however, when the face occupies a very small area in the first image, it can be determined that the visible area of ​​the face is relatively small, and thus it can be determined that the facial state described by the face detection results may present some unreasonable visual effects, and thus the motion capture data determined based on the face detection results may be distorted.

[0098] Based on the above research, in order to avoid driving distortion as much as possible, the present application provides a possible implementation of the data processing method. In this implementation, the method may further include the following step 41.

[0099] Step 41: If the area of ​​the facial region indicated by the facial detection result does not exceed the first threshold, the facial detection result is adjusted using the preset facial information so that the adjusted facial detection result is the preset facial information, thereby enabling the motion capture data corresponding to the first image to be determined subsequently based on the adjusted facial detection result.

[0100] Among them, the preset facial information refers to pre-configured replacement information used when the facial detection result is in an invalid state, so that the preset facial information is used to describe a specific and reasonable facial state, such as a facial state used to describe a smiling expression, so that the preset facial information can replace the facial detection result in an invalid state, thereby effectively avoiding the distortion defects caused by using the facial detection result in an invalid state.

[0101] In addition, the present application does not limit the implementation method of the above-mentioned step 41. For example, in some scenarios, if the preset facial information is pre-set as the initial facial state, then the step 41 can be specifically: if the area of ​​the facial region indicated by the above-mentioned facial detection result does not exceed the first threshold, then the facial detection result is reset so that the facial detection result can be restored to the preset facial information used to describe the initial facial state, which is conducive to improving efficiency.

[0102] Based on the relevant contents of the above step 41, it can be known that after determining the face detection result from the first image, it can be determined whether the area of ​​the face region indicated by the face detection result exceeds the first threshold. If it exceeds, it can be determined that the effective area of ​​the face described by the face detection result is relatively large, and it can be determined that the area blocked by the face is relatively small, so that it can be determined that the face state described by the face detection result is relatively accurate, and then it can be determined that the face detection result is in a valid state, and it can be determined that when the virtual image is driven according to the face detection result, almost no distortion will occur; however, if it does not exceed, it can be determined that the effective area of ​​the face is relatively small, and it can be determined that the area blocked by the face is relatively large, so that it can be determined that the face state described by the face detection result is easy to be distorted. Errors may occur, such as some unreasonable visual effects, etc., and it can be determined that the face detection result is in an invalid state, and it can be determined that distortion is likely to occur when the virtual image is driven based on the face detection result. Therefore, in order to overcome the distortion phenomenon, the face detection result can be reset to achieve adjustment processing for the face detection result to ensure that the adjusted face detection result is the preset face information used to describe the initial face state, so that the face state described by the adjusted face detection result must be reasonable, and then when the virtual image is driven based on the adjusted face detection result, no distortion will occur. In this way, the defects caused by the excessively large occlusion area of ​​the face in the first image can be flexibly solved, thereby effectively reducing the occurrence of driving distortion.

[0103] After research, it was found that for the above-mentioned body detection results, when the body described by the body detection results occupies a larger area in the first image, it can be determined that the visible area of ​​the body (also known as the effective area) is relatively large, and thus it can be determined that the body state described by the body detection results is relatively accurate; however, when the body occupies a very small area in the first image, it can be determined that the effective area of ​​the body is relatively small, and thus it can be determined that the body state described by the body detection results may present some unreasonable visual effects, and thus the motion capture data determined based on the body detection results may be distorted.

[0104] Based on the above research, in order to avoid driving distortion as much as possible, the present application provides a possible implementation of the data processing method. In this implementation, the method may further include the following step 51.

[0105] Step 51: If the area of ​​the body region indicated by the above body detection result does not exceed the second threshold, the body detection result is adjusted using the preset body information so that the adjusted body detection result is the preset body information, so that the motion capture data corresponding to the first image can be determined subsequently based on the adjusted body detection result.

[0106] Among them, the preset body information refers to pre-configured replacement information used when the body detection result is in an invalid state, so that the preset body information is used to describe a specific and reasonable body state, such as used to describe the body state of being in a certain posture (such as standing posture), so that the preset body information can replace the body detection result in an invalid state, and thus can effectively avoid the distortion defects caused by the use of the body detection result in an invalid state.

[0107] In addition, the present application does not limit the implementation method of the above-mentioned step 51. For example, in some scenarios, if the preset body information is pre-set as the initial body state, then the step 51 can be specifically as follows: if the area of ​​the body region indicated by the above-mentioned body detection result does not exceed the second threshold, then the body detection result is reset so that the body detection result can be restored to the preset body information used to describe the initial body state, which is conducive to improving efficiency.

[0108] Based on the relevant content of the above step 51, it can be known that after the body detection result is determined based on the first image, it can be determined whether the area of ​​the body region indicated by the body detection result exceeds the second threshold. If it exceeds, it can be determined that the effective area of ​​the body described by the body detection result is relatively large, and it can be determined that the area blocked by the body is relatively small, so that it can be determined that the body state described by the body detection result is relatively accurate, and then it can be determined that the body detection result is in a valid state, and it can be determined that when the virtual image is driven according to the body detection result, there will be almost no distortion; however, if it does not exceed, it can be determined that the effective area of ​​the body is relatively small, and it can be determined that the area blocked by the body is relatively large, so that it can be determined that the body state described by the body detection result is easy to Errors may occur, such as some unreasonable visual effects, etc., and it can be determined that the body detection result is invalid, and it can be determined that distortion is likely to occur when the virtual image is driven based on the body detection result. Therefore, in order to overcome the distortion, the body detection result can be reset to adjust the body detection result to ensure that the adjusted body detection result is the preset body information used to describe the initial body state, so that the body state described by the adjusted body detection result must be reasonable, and then when the virtual image is driven based on the adjusted body detection result, no distortion will occur. In this way, the defects caused by the excessively large occlusion area of ​​the body in the first image can be flexibly solved, thereby effectively reducing the occurrence of driving distortion.

[0109] Through research, it is found that in some scenarios, if the hand is invisible in the captured image, the predicted information for the hand is inaccurate, so that some problems may occur when the information is used to drive the virtual image, such as presenting unreasonable hand postures or the hands constantly changing without regularity, which can easily cause user confusion. Therefore, in order to overcome this problem, the present application provides a possible implementation method of the data processing method. In this method, when the above-mentioned body detection results include hand detection results, the method may also include the following step 61.

[0110] Step 61: If the hand detection result is used to indicate that the hand is in an invisible state, the preset hand information is used to replace the hand detection result in the body detection result, so that the replaced body detection result includes the preset hand information, thereby making the hand state described by the replaced body detection result consistent with the hand state described by the preset hand information, thereby making it possible to subsequently determine the motion capture data corresponding to the first image based on the replaced body detection result.

[0111] Among them, the preset hand information refers to pre-configured replacement information used when the hand detection result is in an invalid state, so that the preset hand information is used to describe a specific and reasonable hand state, such as used to describe the hand state in a certain gesture (such as a fist gesture), so that the preset hand information can replace the hand detection result in an invalid state, and thus can effectively avoid the distortion defects caused by the use of the hand detection result in an invalid state.

[0112] In addition, the present application does not limit the implementation method of the above-mentioned step 61. For example, in some scenarios, if the preset hand information is pre-set as the initial hand state, then the step 61 can be specifically: if the above-mentioned hand detection result is used to indicate that the hand is in an invisible state, then the hand detection result in the above-mentioned body detection result is reset, so that the hand detection result in the body detection result can be restored to the preset hand information used to describe the initial hand state, which is conducive to improving efficiency.

[0113] Based on the relevant content of the above step 61, it can be known that after the hand detection result is determined based on the first image, it can be determined whether the hand described by the hand detection result is visible based on the confidence level in the hand detection result. If it is visible, it can be determined that the hand state described by the hand detection result is consistent with the hand state presented in the first image (that is, the hand state when the first image was taken), so that it can be determined that the hand detection result is in a valid state, and further it can be determined that there will be almost no distortion when the virtual image is driven based on the hand detection result; however, if it is not visible, it can be determined that the hand state described by the hand detection result is consistent with the hand state when the first image was taken. The hand state when capturing an image may be very different, so it can be determined that the hand detection result is in an invalid state, and further it can be determined that some problems are likely to occur when driving the virtual image based on the hand detection result. Therefore, in order to overcome this problem, the preset hand information configured in a certain way can be used to replace the hand detection result in the above-mentioned body detection result, so that the replaced body detection result includes the preset hand information, so that the hand state described by the replaced body detection result is consistent with the hand state described by the preset hand information. In this way, the problem caused by the invisible hand in the first image can be flexibly solved, which is beneficial to improving the motion capture experience.

[0114] It should be noted that the present application does not limit the implementation method of the above-mentioned step of "determining whether the hand described by the hand detection result is visible based on the confidence in the hand detection result". For example, it can be specifically: if the confidence is greater than a preset confidence threshold, it can be determined that the hand described by the hand detection result is in a visible state in the first image; if the confidence is less than or equal to the confidence threshold, it can be determined that the hand described by the hand detection result is in an invisible state in the first image.

[0115] Research has found that because different objects have different hand characteristics (such as different finger lengths), the hand detection results of different objects under the same gesture may be different. As a result, when the result is used to drive the virtual image, the gesture presented is different from the actual gesture of the object, thus affecting the motion capture effect.

[0116] Based on the above research, in order to better improve the motion capture effect, the present application provides a possible implementation of the data processing method. In this manner, when the above body detection results include hand detection results, the method may also include the following step 71.

[0117] Step 71: If the above-mentioned hand detection result is used to indicate that the hand is in a visible state, and there is a gesture matching the gesture indicated by the hand detection result among the pre-set multiple candidate gestures, then the hand detection result in the body detection result is replaced by the hand information pre-configured for the matching gesture, so that the replaced body detection result includes the configured hand information, so that the gesture described by the replaced body detection result is the matching gesture, and then the motion capture data corresponding to the first image can be determined based on the replaced body detection result.

[0118] Among them, multiple candidate gestures refer to pre-set gestures that are configured with hand information and are used to describe a normalized processing result, so as to ensure that the hand detection results of different objects under the same gesture can be normalized based on these candidate gestures and the hand information configured for these candidate gestures.

[0119] Based on the relevant content of the above step 71, it can be known that after the hand detection result is determined based on the first image, if the hand detection result is used to indicate that the hand is in a visible state, then a gesture library pre-set for storing a large number of gestures is searched to see whether there is a gesture matching the gesture indicated by the hand detection result. If not, the virtual image driven by the hand detection result can be used subsequently; however, if it exists, it can be determined that the matching gesture is used to represent the normalized processing result of the gesture indicated by the hand detection result, so the hand detection result in the above body detection result can be replaced by the hand information pre-configured for the matching gesture, so that the replaced body detection result includes the configured hand information, so that the gesture presented when the virtual image is subsequently driven according to the replaced body detection result is the matching gesture. In this way, the rendering of some specific gestures can be constrained by means of the gesture library, thereby effectively avoiding the defect of poor rendering effect of these specific gestures due to different hand characteristics of different objects.

[0120] Through research, it is found that for any image, the process of determining the hand detection result of the image can be: first determine the skeleton point information from the image; then determine the hand key points (such as wrist points) and the hand position based on the skeleton point information; then, predict the hand detection result from the image based on the hand key points and the hand position; finally, determine the body detection result based on the hand detection result and the skeleton point information.

[0121] Research has also found that for the body detection process shown above, because the hands need to use some additional steps for detection and processing, there may be a large gap between the time costs of determining the motion capture data from different images. For example, the time cost of determining the motion capture data from an image that includes the hands is 30 milliseconds, but the time cost of determining the motion capture data from an image that does not include the hands is 10 milliseconds. This makes it easy for the driven virtual image's movements to change less smoothly when these images are converted into motion capture data in real time, thereby affecting the determination effect of the motion capture data.

[0122] Based on the above research, in order to better improve the effect, the present application provides a possible implementation of the data processing method. In this manner, when the image sequence captured by a monocular camera includes a first image and a third image, and the arrangement position of the third image in the image sequence is adjacent to the arrangement position of the first image in the image sequence, the method may also include the following step 81.

[0123] Step 81: Based on a preset frame rate, interpolate the motion capture data corresponding to the first image and the motion capture data corresponding to the third image to obtain a motion capture data sequence, so that the frame rate of the motion capture data sequence is the preset frame rate, thereby enabling the motion capture data sequence to meet the frame rate constraint of the motion change of the virtual image, so that the motion capture data sequence can be used to drive the virtual image subsequently to ensure that the motion change of the virtual image is as smooth as possible.

[0124] The third image refers to an image in the image sequence captured by the monocular camera and adjacent to the arrangement position of the first image. It can be seen that the third image can be a frame image before the first image or a frame image after the first image, so that the interpolation processing between the motion capture data corresponding to the two frames of images can be used to obtain multiple frames of motion capture data that meet the preset frame rate.

[0125] The preset frame rate refers to a frame rate (such as 60 frames per second) set in advance for the driving process of the virtual image to ensure that the virtual image can smoothly perform movement changes.

[0126] In addition, the present application does not limit the implementation method of the above interpolation processing. For example, it can be implemented using any interpolation method, such as a bilinear interpolation method.

[0127] Based on the relevant contents of the above step 81, it can be known that for the i+1th frame image captured in real time by the monocular camera, after obtaining the motion capture data corresponding to the i+1th frame image, in order to avoid the driving being not smooth enough, the motion capture data corresponding to the i-th frame image obtained historically and the motion capture data corresponding to the i+1th frame image can be interpolated according to the preset frame rate to obtain [the shooting time T of the i-th frame image i, the shooting time of the i+1th frame image is T i+1 ], so that the number of frames of the motion capture data in the motion capture data sequence is greater than 2, and the frame rate of the motion capture data sequence is the preset frame rate, and the moments when each frame of the motion capture data in the motion capture data sequence is used to drive the virtual image all belong to [T i , T i+1 ] time period, so that the motion capture data sequence can more smoothly describe the object photographed by the monocular camera at the preset frame rate in [T i , T i+1 ] The motion changes within this time period can be restored more smoothly by the virtual image when the virtual image is driven according to the motion capture data sequence, which is beneficial to improving the determination effect of the motion capture data.

[0128] After research, it was found that for the solution shown in the previous paragraph, the interpolation process also takes time, so that the solution may have a certain degree of delay, thus affecting the driving effect.

[0129] Research has also found that the maximum time cost of determining motion capture data from each frame of image is about 30 milliseconds, so that the maximum time cost of determining motion capture data from two adjacent frames of image is about 60 milliseconds, and the maximum time cost of the process from shooting two adjacent frames of image to obtaining the corresponding interpolated motion capture data is about 70 milliseconds. Therefore, in order to avoid the delay defect caused by the interpolation processing, the movement changes of the virtual image can be configured to be 70 milliseconds later than the actual movement changes of the object captured by the monocular camera.

[0130] Based on the above research, in order to better improve the effect, the present application provides a possible implementation of the data processing method. In this way, when the image sequence captured by a monocular camera includes a first image and a third image, and the arrangement position of the third image in the image sequence is adjacent to the arrangement position of the first image in the image sequence, the method may at least include the following steps 91-93.

[0131] Step 91: According to a preset frame rate, interpolation processing is performed on the motion capture data corresponding to the first image and the motion capture data corresponding to the third image to obtain a motion capture data sequence, so that the frame rate of the motion capture data sequence is the preset frame rate.

[0132] It should be noted that for the relevant content of step 91, please refer to the relevant content of step 81 above.

[0133] Step 92: Determine the use time of each motion capture data in the motion capture data sequence according to the time of capturing the first image, the time of capturing the third image, and a preset delay time (eg, 70 milliseconds).

[0134] The preset delay duration refers to a preset duration used to describe the overall delay of the motion changes presented by driving the virtual image relative to the actual motion changes of the object photographed by the monocular camera, such as 70 milliseconds.

[0135] In addition, the present application does not limit the method for obtaining the preset delay time. For example, the preset delay time can be determined based on an actual application scenario.

[0136] For another example, in order to better improve the driving effect, the preset delay duration can be determined based on the sum of the time consumed in the process of determining the motion capture data corresponding to the first image and the time consumed in the process of determining the motion capture data corresponding to the third image, to ensure that the preset delay duration is greater than the sum.

[0137] In addition, the present application does not limit the implementation of the above step 92. For example, when the time of capturing the first image is T i , and the time of taking the third image is T i+1 , and the preset delay time is 70 milliseconds, if the motion capture data sequence obtained by interpolating the motion capture data corresponding to the first image and the motion capture data corresponding to the third image includes K frames of motion capture data, then the use time of the motion capture data at the kth arrangement position in the motion capture sequence = T i +(k-1)×the time interval corresponding to the preset frame rate+70 milliseconds, and the time interval is used to describe the time interval between two adjacent frames of data in the motion capture data sequence at the preset frame rate, k is a positive integer, k≤K, K is a positive integer, K is based on the preset frame rate, the T i+1 And the T i Determined.

[0138] It should be noted that in some scenarios, such as scenarios where images are captured at a certain frequency, if the frame rate used by the monocular camera when capturing images is regarded as the acquisition frame rate, it can be determined that the time difference between two adjacent frames of images captured by the monocular camera (e.g., the time interval corresponding to the acquisition frame rate) is fixed, and thus it can be determined that the above K is determined based on the acquisition frame rate and the above preset frame rate.

[0139] It should also be noted that the present application does not limit the relationship between the execution time of the above step 92 and the execution time of the above step 91, for example, the two are the same. Another example is that the two are different.

[0140] Step 93: For any motion capture data in the above motion capture data sequence, use the motion capture data to drive the virtual image at the time when the motion capture data is used.

[0141] Based on the relevant contents of the above steps 91 to 93, it can be known that for the i+1th frame image captured in real time by the monocular camera, after obtaining the motion capture data corresponding to the i+1th frame image, the motion capture data corresponding to the i-th frame image obtained historically and the motion capture data corresponding to the i+1th frame image can be interpolated according to the preset frame rate to obtain [the shooting time T of the i-th frame image i , the shooting time of the i+1th frame image is T i+1 ] The motion capture data sequence corresponding to this time period; then according to the preset delay time and [T i , T i+1 ] time period, determine the use time of each motion capture data in the motion capture data sequence, so that the use time belongs to [T i +Preset delay time, T i+1 + preset delay time] so that you can i +Preset delay time, T i+1 +preset delay duration], so that the virtual image can restore the motion changes of the object photographed by the monocular camera within the time period [Ti, Ti+1]. In this way, the motion changes of the virtual image can be delayed as a whole in the time dimension compared with the actual motion changes of the object photographed by the monocular camera, thereby ensuring that the virtual image can smoothly restore the actual motion changes, which is conducive to better improving the determination effect of the motion capture data.

[0142] Through research, it is found that, for an object, the face of the object presents a smaller change between two adjacent frames of images, so that the difference in the motion capture data corresponding to the two frames of images on the face is relatively small, so that a smaller smoothness can be used for the face during the interpolation process to ensure that the interpolation result retains facial details as much as possible; however, the hand of the object may present a larger change between two adjacent frames of images, so that the motion capture data corresponding to the two frames of images on the hand is relatively different, so that a larger smoothness can be used for the hand during the interpolation process to ensure that the interpolation result is as smooth as possible.

[0143] Based on the above research findings, in order to better improve the smoothing effect, the present application also provides a possible implementation method of the above interpolation processing. In this method, when the motion capture data corresponding to the first image includes the motion capture data corresponding to each part under the first image, and the motion capture data corresponding to the third image includes the motion capture data corresponding to each part under the third image, the motion capture data sequence obtained by interpolating the two motion capture data is used to indicate the interpolation processing results of each part; and for any part, the interpolation processing result of the part is obtained by interpolating the motion capture data corresponding to the part under the first image and the motion capture data corresponding to the part under the third image based on the interpolation parameters corresponding to the part, and the interpolation parameters are used to indicate the smoothness of the interpolation processing result of the part, so that interpolation processing with different smoothness can be achieved for different parts.

[0144] It can be seen that after obtaining the motion capture data corresponding to the jth part under the i-th frame image and the motion capture data corresponding to the jth part under the i+1-th frame image, the two kinds of motion capture data can be interpolated according to the interpolation parameters corresponding to the jth part to obtain the interpolation result of the jth part, so that the smoothness presented by the interpolation result is consistent with the smoothness indicated by the interpolation parameter, so that the motion capture data sequence including the interpolation result of the jth part can present the changes of the jth part according to the smoothness. Wherein, j is a positive integer, j≤J, J is a positive integer, and J represents the number of parts.

[0145] It should be noted that the motion capture data corresponding to the j-th part in the i-th frame image refers to the data determined from the i-th frame image and used to describe the state of the j-th part. For example, if the j-th part is the face, then the "motion capture data corresponding to the j-th part in the i-frame image" may be the face detection result determined from the i-frame image.

[0146] It should also be noted that the present application does not limit the implementation method of the interpolation parameter. For example, when the above interpolation process is implemented by bilinear interpolation, the interpolation parameter can be the weight involved in the bilinear interpolation. Among them, the weight can have some influence on the interpolation result obtained by bilinear interpolation, such as smoothness, detail retention, interpolation accuracy, image quality, computational complexity, etc.

[0147] It should also be noted that the present application does not limit the method for obtaining the interpolation parameters corresponding to the j-th part. For example, it can be provided by relevant personnel through human-computer interaction.

[0148] For another example, in order to improve flexibility, the interpolation parameter corresponding to the j-th part can be determined based on the variation range of the motion capture data corresponding to the j-th part in the image sequence captured by the monocular camera, so that the smoothness described by the interpolation parameter matches the variation range (such as positive correlation). It can be seen that in a possible implementation, the greater the variation range, the greater the smoothness.

[0149] It should be noted that, for the variation amplitude of the motion capture data corresponding to the j-th part in the image sequence taken by the monocular camera, the process of determining the variation amplitude can be: first, for any pair of adjacent images in the image sequence, calculate the difference between the motion capture data corresponding to the adjacent images; then, based on these differences, determine the variation amplitude (such as taking the maximum value of these differences as the variation amplitude, or drawing these differences into a curve as the variation amplitude, etc.), so that the variation amplitude can better represent the motion change characteristics of the j-th part.

[0150] Research has found that the motion capture data corresponding to different parts may be obtained by using different algorithms. However, due to the different performances of different algorithms, the amount of noise generated by different algorithms is different, which makes the motion capture data corresponding to different parts carry different amounts of noise. Therefore, in order to suppress the interference caused by noise as much as possible, different degrees of smoothness can be achieved when interpolating the motion capture data corresponding to different parts.

[0151] Based on the above research, it can be known that in order to better improve the smoothing effect, for any part, when the motion capture data corresponding to the part under the first image and the motion capture data corresponding to the part under the third image are both determined using a target detection algorithm (such as a face detection algorithm or a body detection algorithm), the interpolation parameter corresponding to the part can be determined based on the performance of the target detection algorithm and the change amplitude of the motion capture data corresponding to the part under the above image sequence, so that the smoothness described by the interpolation parameter not only matches the change amplitude, but also can suppress the noise generated by the algorithm as much as possible. Among them, the performance is used to describe the amount of noise data generated when the algorithm is used to process data; and this application does not limit the method of obtaining the performance, for example, it can be implemented using any algorithm evaluation method.

[0152] It should be noted that the present application does not limit the determination process of the above-mentioned interpolation parameters. For example, the interpolation parameters can be obtained by querying from a pre-constructed parameter library. Among them, the parameter library records the interpolation parameters corresponding to various <algorithm, variation range> combinations. For another example, the interpolation parameters can be determined by using a constructed machine learning model that can predict the interpolation parameters based on the algorithm + variation range. For another example, the interpolation parameters can be determined by using a pre-fitted function that is used to describe the interpolation parameters corresponding to various <algorithm, variation range> combinations.

[0153] Based on the relevant contents of the above data processing method, it can be known that the motion capture data determination scheme provided in this application has the advantages shown in ①-④ below.

[0154] ① This application is submitted through certain means (such as Figure 2 The method shown in the figure solves the defect caused by the face detection result and the body detection result not belonging to the same object in a multi-object scene.

[0155] ② This application optimizes the final motion capture data by resetting some detection results (such as face detection results, body detection results or hand detection results, etc.) to certain preset information, thereby reducing the occurrence of extreme phenomena.

[0156] ③ This application not only solves the problem of a large difference between the actual algorithm frame rate (such as the actual efficiency of converting images into motion capture data) and the user's expected rendering frame rate (such as the frequency of change of the virtual image) through interpolation processing, but also solves the delay problem caused by increased smoothing.

[0157] ④ The present application solves the jitter caused by driving the virtual image due to unstable algorithm data by performing interpolation processing on different parts according to different interpolation parameters.

[0158] Based on the data processing method provided in the embodiment of the present application, the embodiment of the present application also provides a data processing device. Figure 3 Explain and illustrate. Figure 3 This is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. It should be noted that for the technical details of the data processing device provided in an embodiment of the present application, please refer to the relevant content of the data processing method above.

[0159] like Figure 3 As shown, the data processing device 300 provided in the embodiment of the present application includes:

[0160] An acquisition unit 301 is configured to acquire a first image, where the first image is used to describe at least one object;

[0161] The detection unit 302 is configured to perform face detection processing on the first image to obtain a face detection result, wherein the face detection result is used to indicate face information of the object, wherein the face information includes a face area, and perform body detection processing on the first image to obtain a body detection result, wherein the body detection result is used to indicate body information of the object, wherein the body information includes a body area, and the body includes a head;

[0162] An updating unit 303 is used to update the body detection result in response to the overlapping area between the body area indicated by the body detection result and the facial area indicated by the facial detection result being not greater than a preset area threshold, so that the overlapping area between the body area indicated by the updated body detection result and the facial area indicated by the facial detection result is greater than the preset area threshold, and determine the motion capture data corresponding to the first image based on the updated body detection result and the facial detection result.

[0163] In a possible implementation manner, the data processing device 300 further includes:

[0164] A response unit is used to determine the motion capture data corresponding to the first image based on the body detection result and the face detection result in response to the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result being greater than a preset area threshold.

[0165] In a possible implementation, the detection unit 302 is specifically used to: obtain a first area, where the first area is the body area indicated by the body detection result; perform bone point detection processing on the first image based on the first area to obtain bone point information, where the bone point information is used to describe the state of each bone point in the first area; and determine the body detection result based on the first area and the bone point information.

[0166] In one possible implementation, if a multiplexing condition is met, the first region is obtained by multiplexing a body region recognition result of a second image, the second image is photographed by a monocular camera earlier than the first image is photographed by the monocular camera, a process of determining motion capture data corresponding to the second image is earlier than a process of determining the motion capture data corresponding to the first image, and during the process of determining the motion capture data corresponding to the second image, body region recognition processing is performed on the second image to obtain a body region recognition result of the second image; if the multiplexing condition is not met, the first region is obtained by performing body region recognition processing on the first image.

[0167] In a possible implementation, if the first region is obtained by reusing the body region recognition result of the second image, the process of determining the updated body detection result includes: performing body region recognition processing on the first image to obtain the second region; performing bone point detection processing on the first image based on the second region to obtain a processing result, and the processing result is used to describe the state of each bone point in the second region; and determining the updated body detection result based on the second region and the processing result.

[0168] In a possible implementation manner, the data processing device 300 further includes:

[0169] The first adjustment unit is used to adjust the face detection result using preset face information after obtaining the face detection result, if the area of ​​the face region indicated by the face detection result does not exceed the first threshold, the adjusted face detection result is the preset face information, and the motion capture data is determined based on the adjusted face detection result.

[0170] In a possible implementation manner, the data processing device 300 further includes:

[0171] The second adjustment unit is used to adjust the body detection result using preset body information after obtaining the body detection result, if the area of ​​the body region indicated by the body detection result does not exceed the second threshold, and the adjusted body detection result is the preset body information, and the motion capture data is determined based on the adjusted body detection result.

[0172] In a possible implementation manner, the body detection result includes a hand detection result; and the data processing device 300 further includes at least one of the following units:

[0173] A first replacement unit is used for replacing the hand detection result in the body detection result with preset hand information after obtaining the body detection result, if the hand detection result is used to indicate that the hand is in an invisible state, wherein the replaced body detection result includes the preset hand information, and the motion capture data is determined according to the replaced body detection result;

[0174] A second replacement unit is used to replace the hand detection result in the body detection result with the hand information pre-configured for the matching gesture after obtaining the body detection result, if the hand detection result is used to indicate that the hand is in a visible state, and there is a gesture matching the gesture indicated by the hand detection result among the pre-set multiple candidate gestures. The replaced body detection result includes the configured hand information, and the motion capture data is determined based on the replaced body detection result.

[0175] In a possible implementation manner, the data processing device 300 further includes:

[0176] an interpolation unit, configured to perform interpolation processing on the motion capture data corresponding to the first image and the motion capture data corresponding to the third image according to a preset frame rate to obtain a motion capture data sequence, wherein the frame rate of the motion capture data sequence is the preset frame rate, the image sequence captured by the monocular camera includes the first image and the third image, and the arrangement position of the third image in the image sequence is adjacent to the arrangement position of the first image in the image sequence;

[0177] A driving unit is used to drive the virtual image using the motion capture data sequence.

[0178] In a possible implementation manner, the data processing device 300 further includes:

[0179] a determining unit, configured to determine a use time of each motion capture data in the motion capture data sequence according to a time when the first image is captured, a time when the third image is captured, and a preset delay time;

[0180] The driving unit is specifically used to: for any motion capture data in the motion capture data sequence, drive the virtual image using the motion capture data at the time when the motion capture data is used.

[0181] In a possible implementation manner, the sum of the time consumed by a process of determining the motion capture data corresponding to the first image and the time consumed by a process of determining the motion capture data corresponding to the third image is less than the preset delay time.

[0182] In a possible implementation, the motion capture data corresponding to the first image includes the motion capture data corresponding to each part under the first image, and the motion capture data corresponding to the third image includes the motion capture data corresponding to each part under the third image, and the motion capture data sequence is used to indicate the interpolation processing results of each part; for any of the parts, the interpolation processing result of the part is obtained by interpolating the motion capture data corresponding to the part under the first image and the motion capture data corresponding to the part under the third image based on the interpolation parameters corresponding to the part, and the interpolation parameters are used to indicate the smoothness of the interpolation processing result of the part.

[0183] In a possible implementation, for any of the parts, the motion capture data corresponding to the part under the first image and the motion capture data corresponding to the part under the third image are both determined using a target detection algorithm, and the target detection algorithm is used to implement the face detection processing or the body detection processing, and the interpolation parameters corresponding to the part are determined based on the performance of the target detection algorithm and the variation amplitude of the motion capture data corresponding to the part under the image sequence.

[0184] In a possible implementation manner, the first image is a two-dimensional image captured in real time using a monocular camera.

[0185] Based on the relevant contents of the above-mentioned data processing device 300, it can be known that the working principle of the device 300 is: after obtaining a first image for describing some objects, firstly, face detection processing is performed on the first image to obtain a face detection result, so that the face detection result is used to indicate the face information of one of the objects, and body detection processing is performed on the first image to obtain a body detection result, so that the body detection result is used to indicate the body information of one of the objects, so that when it is detected that the overlapping area between the body area indicated by the body detection result and the face area indicated by the face detection result is not greater than the preset area threshold, it can be determined that the body described by the body detection result and the face described by the face detection result do not belong to the same object, so the body detection result is updated. The method comprises the following steps: obtaining a body detection result after the update and a face detection result after the update, so that the overlapping area between the body region indicated by the updated body detection result and the face region indicated by the face detection result is greater than a preset area threshold, so that the body described by the updated body detection result and the face described by the face detection result belong to the same object; then, determining the motion capture data corresponding to the first image based on the updated body detection result and the face detection result, so that the motion capture data can more accurately represent the action of an object presented by the first image, so that when the motion capture data is subsequently used to drive the virtual image, the virtual image can accurately restore the action presented by the object in the first image, so that the motion capture data can be determined from one image, thereby effectively reducing the difficulty of obtaining the motion capture data.

[0186] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the data processing method provided in the embodiment of the present application.

[0187] See also Figure 4 , which shows a schematic diagram of the structure of an electronic device 400 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0188] like Figure 4 As shown, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0189] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0190] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0191] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept, and the technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0192] The embodiment of the present application also provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the data processing method provided in the embodiment of the present application.

[0193] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0194] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0195] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0196] The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the method.

[0197] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0198] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0199] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit / module does not, in some cases, constitute a limitation on the unit itself.

[0200] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0201] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0202] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0203] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0204] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0205] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0206] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method, characterized in that: The method comprises: Acquire a first image, wherein the first image is used to describe at least one object; Performing face detection processing on the first image to obtain a face detection result, wherein the face detection result is used to indicate face information of the object, wherein the face information includes a face area, and performing body detection processing on the first image to obtain a body detection result, wherein the body detection result is used to indicate body information of the object, wherein the body information includes a body area, wherein the body includes a head; In response to the overlapping area between the body area indicated by the body detection result and the facial area indicated by the face detection result being not greater than a preset area threshold, the body detection result is updated so that the overlapping area between the body area indicated by the updated body detection result and the facial area indicated by the face detection result is greater than the preset area threshold, and the motion capture data corresponding to the first image is determined based on the updated body detection result and the facial detection result.

2. The method according to claim 1, characterized in that The method further comprises: In response to an overlapping area between a body region indicated by the body detection result and a face region indicated by the face detection result being greater than a preset area threshold, motion capture data corresponding to the first image is determined based on the body detection result and the face detection result.

3. The method according to claim 1, characterized in that The process of determining the physical examination result includes: Acquire a first area, where the first area is a body area indicated by the body detection result; According to the first region, performing skeleton point detection processing on the first image to obtain skeleton point information, wherein the skeleton point information is used to describe the state of each skeleton point in the first region; The body detection result is determined according to the first region and the skeleton point information.

4. The method according to claim 3, characterized in that If the multiplexing condition is met, the first region is obtained by multiplexing the body region recognition result of the second image, the time when the second image is photographed by the monocular camera is earlier than the time when the first image is photographed by the monocular camera, the process of determining the motion capture data corresponding to the second image is earlier than the process of determining the motion capture data corresponding to the first image, and in the process of determining the motion capture data corresponding to the second image, the body region recognition result of the second image is obtained by performing body region recognition processing on the second image; If the reuse condition is not met, the first area is obtained by performing body area recognition processing on the first image.

5. The method according to claim 4, characterized in that If the first region is obtained by reusing the body region recognition result of the second image, the process of determining the updated body detection result includes: Performing body region recognition processing on the first image to obtain a second region; According to the second region, performing skeleton point detection processing on the first image to obtain a processing result, wherein the processing result is used to describe the state of each skeleton point in the second region; The updated body detection result is determined according to the second area and the processing result.

6. The method according to claim 1, characterized in that After obtaining the face detection result, the method further includes: If the area of ​​the face region indicated by the face detection result does not exceed a first threshold, adjusting the face detection result using preset face information, the adjusted face detection result being the preset face information, and the motion capture data being determined based on the adjusted face detection result; and / or, After obtaining the physical examination result, the method further includes: If the area of ​​the body region indicated by the body detection result does not exceed the second threshold, the body detection result is adjusted using the preset body information, and the adjusted body detection result is the preset body information. The motion capture data is determined based on the adjusted body detection result.

7. The method according to claim 1, characterized in that The body detection result includes a hand detection result; After obtaining the body detection result, the method further comprises at least one of the following steps: If the hand detection result is used to indicate that the hand is in an invisible state, the hand detection result in the body detection result is replaced with the preset hand information, and the replaced body detection result includes the preset hand information; If the hand detection result is used to indicate that the hand is in a visible state, and there is a gesture matching the gesture indicated by the hand detection result among the pre-set multiple candidate gestures, the hand detection result in the body detection result is replaced by the hand information pre-configured for the matching gesture, and the replaced body detection result includes the configured hand information; The motion capture data is determined based on the replaced body detection result.

8. The method according to claim 1, characterized in that The method further comprises: According to a preset frame rate, interpolation processing is performed on the motion capture data corresponding to the first image and the motion capture data corresponding to the third image to obtain a motion capture data sequence, the frame rate of the motion capture data sequence is the preset frame rate, the image sequence captured by the monocular camera includes the first image and the third image, and the arrangement position of the third image in the image sequence is adjacent to the arrangement position of the first image in the image sequence; The motion capture data sequence is used to drive the virtual image.

9. The method according to claim 8, characterized in that The method further comprises: Determining a use time of each motion capture data in the motion capture data sequence according to a time when the first image is captured, a time when the third image is captured, and a preset delay time; The method of using the motion capture data sequence to drive the virtual image includes: For any motion capture data in the motion capture data sequence, the virtual image is driven by using the motion capture data at the time when the motion capture data is used.

10. The method according to claim 9, characterized in that The sum of the time consumed in the process of determining the motion capture data corresponding to the first image and the time consumed in the process of determining the motion capture data corresponding to the third image is less than the preset delay time.

11. The method according to claim 8, characterized in that The motion capture data corresponding to the first image includes the motion capture data corresponding to each part under the first image, the motion capture data corresponding to the third image includes the motion capture data corresponding to each part under the third image, and the motion capture data sequence is used to indicate the interpolation processing result of each part; For any of the parts, the interpolation processing result of the part is obtained by interpolating the motion capture data corresponding to the part in the first image and the motion capture data corresponding to the part in the third image based on the interpolation parameters corresponding to the part, and the interpolation parameters are used to indicate the smoothness of the interpolation processing result of the part.

12. The method according to claim 11, characterized in that For any of the parts, the motion capture data corresponding to the part under the first image and the motion capture data corresponding to the part under the third image are both determined using a target detection algorithm, and the target detection algorithm is used to implement the face detection processing or the body detection processing. The interpolation parameters corresponding to the part are determined based on the performance of the target detection algorithm and the variation amplitude of the motion capture data corresponding to the part under the image sequence.

13. The method according to any one of claims 1 to 12, characterized in that: The first image is a two-dimensional image captured in real time using a monocular camera.

14. A data processing device, characterized in that: include: An acquisition unit, configured to acquire a first image, wherein the first image is used to describe at least one object; a detection unit, configured to perform face detection processing on the first image to obtain a face detection result, wherein the face detection result is used to indicate face information of the object, wherein the face information includes a face area, and to perform body detection processing on the first image to obtain a body detection result, wherein the body detection result is used to indicate body information of the object, wherein the body information includes a body area, and wherein the body includes a head; An updating unit is used to update the body detection result in response to the overlapping area between the body area indicated by the body detection result and the facial area indicated by the facial detection result being not greater than a preset area threshold, so that the overlapping area between the body area indicated by the updated body detection result and the facial area indicated by the facial detection result is greater than the preset area threshold, and determine the motion capture data corresponding to the first image based on the updated body detection result and the facial detection result.

15. An electronic device, characterized in that: The device comprises: a processor and a memory; The memory is used to store instructions or computer programs; The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the method according to any one of claims 1 to 13.

16. A computer readable medium, characterized in that The computer-readable medium stores instructions or computer programs, and when the instructions or computer programs are executed on a device, the device executes the method according to any one of claims 1 to 13.

17. A computer program product, characterized in that It comprises a computer program carried on a non-transitory computer-readable medium, the computer program comprising a program code for executing the method according to any one of claims 1 to 13.