A data processing method and related device

By identifying the contact between the object's feet and the motion platform, the motion data is corrected, solving the drift problem of virtual avatars on mobile terminals and achieving a realistic and stable presentation of virtual avatars.

CN117037262BActive Publication Date: 2025-12-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210896291.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-12-09
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

Virtual avatars exhibit drifting during their movements on mobile devices, resulting in an unrealistic and unstable presentation.

Method used

By acquiring images and performing recognition processing, the contact between the object's feet and the motion platform is determined. Based on the contact situation, the motion data is corrected to ensure that the virtual feet of the virtual avatar are in contact with the motion platform.

Benefits of technology

This solves the problem of virtual avatar drift and improves the realism and stability of virtual avatars.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117037262B_ABST
    Figure CN117037262B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method and related equipment, wherein the method comprises: acquiring an image to be processed, the image being obtained by photographing an object on a motion platform; performing identification processing on the image to obtain motion data of the object; determining a contact condition of a foot of the object and the motion platform according to the motion data; correcting the motion data based on the contact condition of the foot of the object and the motion platform; and displaying a virtual image matched with the object according to the corrected motion data, a virtual foot of the virtual image being in contact with the motion platform. Embodiments of the present application can solve the drift problem of the virtual image and ensure the presentation of the virtual image to be real and stable.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of Internet, in particular to a data processing method, a data processing device, a computer device, a computer readable storage medium and a computer program product. BACKGROUND

[0002] With the development of Internet technology, an object can shoot the action of the object through the camera of a mobile terminal device such as a mobile phone or a tablet, and drive a virtual image with the same action to be displayed on the screen according to the shot action of the object. For example, when the object performs a dance in the real world, the object can see the virtual image corresponding to the object and the object performing the same dance action on the screen of the mobile terminal. However, it is found in practice that the virtual image in the mobile terminal often has a drift phenomenon (i.e. the feet of the virtual image are not in contact with the ground) during the action, which makes the presentation of the virtual image not real and stable, and the visual effect is poor. SUMMARY

[0003] The embodiments of the present application provide a data processing method and related devices, which can solve the drift problem of the virtual image and ensure the real and stable presentation of the virtual image.

[0004] In one aspect, the embodiments of the present application provide a data processing method, which comprises:

[0005] obtaining an image to be processed, the image being obtained by shooting an object on a motion platform;

[0006] performing identification processing on the image to obtain action data of the object;

[0007] determining the contact condition of the feet of the object and the motion platform according to the action data;

[0008] correcting the action data based on the contact condition of the feet of the object and the motion platform;

[0009] displaying a virtual image matched with the object according to the corrected action data, and the virtual feet of the virtual image being in contact with the motion platform.

[0010] In one aspect, the embodiments of the present application provide a data processing device, which comprises:

[0011] an obtaining unit, configured to obtain an image to be processed, the image being obtained by shooting an object on a motion platform;

[0012] a processing unit, configured to perform identification processing on the image to obtain action data of the object;

[0013] the processing unit is further configured to determine the contact condition of the feet of the object and the motion platform according to the action data.

[0014] The processing unit is further configured to correct the motion data based on the contact condition of the feet of the object with the motion platform.

[0015] The processing unit is further configured to display the virtual image matching the object according to the corrected motion data, and a virtual foot of the virtual image is in contact with the motion platform.

[0016] In an embodiment, the motion data comprises at least one of: position data and posture data; the position data comprises coordinates of a center of gravity of the object in a three-dimensional space; and the posture data comprises a rotation matrix formed by rotation angles of respective rotation joints of the object.

[0017] The object has N feet; and the contact condition of the feet of the object with the motion platform comprises at least one of: all the N feet of the object are in contact with the motion platform; M feet of the N feet of the object are in contact with the motion platform, and the remaining N-M feet of the N feet are not in contact with the motion platform, where M and N are positive integers and M is less than or equal to N.

[0018] In an embodiment, when determining the contact condition of the feet of the object with the motion platform according to the motion data, the processing unit is specifically configured to:

[0019] determine two-dimensional coordinates of a foot key point of each foot according to the motion data;

[0020] predict a contact probability of the foot key point of each foot with the motion platform according to the two-dimensional coordinates of the foot key point of each foot;

[0021] determine the contact condition of the feet of the object with the motion platform according to the contact probability of the foot key point of each foot with the motion platform.

[0022] In an embodiment, when correcting the motion data based on the contact condition of the feet of the object with the motion platform, the processing unit is specifically configured to:

[0023] if the contact condition comprises that all the N feet of the object are in contact with the motion platform, the posture data of the object is corrected.

[0024] In an embodiment, the number of the foot key points of each foot is multiple, and when correcting the posture data of the object, the processing unit is specifically configured to:

[0025] maintain a height of a lowest point of the multiple foot key points of each foot in a three-dimensional space unchanged, and adjust rotation angles of knee joints and thigh joints of the object.

[0026] In one embodiment, the processing unit, when correcting the motion data based on the contact condition of the feet of the object with the motion platform, can be specifically used for:

[0027] If the contact condition includes M feet of N feet of the object being in contact with the motion platform, but the remaining N-M feet are not in contact with the motion platform, the position data of the object is corrected.

[0028] In one embodiment, the processing unit, when correcting the position data of the object, can be specifically used for:

[0029] Keeping the coordinates of the key points of the M feet in the three-dimensional space unchanged, and adjusting the coordinates of the center of gravity of the object in the three-dimensional space.

[0030] In one embodiment, the processing unit, when determining the two-dimensional coordinates of the key points of the feet of each foot according to the motion data, can be specifically used for:

[0031] Obtaining object skeleton data;

[0032] According to the posture data and the object skeleton data, the coordinates of each rotation joint of the object in the three-dimensional space are calculated;

[0033] According to the position data and the fixed focal length, the coordinates of each rotation joint in the three-dimensional space are respectively projected to obtain the two-dimensional coordinates of each rotation joint;

[0034] According to the two-dimensional coordinates of each rotation joint, the two-dimensional coordinates of the key points of the feet of each foot are determined.

[0035] In one embodiment, the processing unit, when determining the two-dimensional coordinates of the key points of the feet of each foot according to the two-dimensional coordinates of each rotation joint, can be specifically used for:

[0036] According to the posture data, a bounding box of a target region in the object is calculated;

[0037] According to the two-dimensional coordinates of the rotation joints included in the target region, key point recognition processing is performed on the target region in the bounding box to obtain the two-dimensional coordinates of the key points of the feet of each foot.

[0038] In one embodiment, the processing unit, when predicting the contact probability of the key points of the feet of each foot with the motion platform according to the two-dimensional coordinates of the key points of the feet of each foot, can be specifically used for:

[0039] Each of the N feet in the image is framed to obtain a frame image corresponding to each foot;

[0040] According to the two-dimensional coordinates of the foot key points of each foot, the trained first neural network is called to respectively predict the foot key points in the N frame images, to obtain the contact probability of the foot key points of each foot with the sports platform.

[0041] In an embodiment, the action data of the object is obtained by calling the trained second neural network to perform recognition processing on the image; the processing unit is further configured to:

[0042] calling the trained second neural network to perform recognition processing on the image to obtain the heat Figure Two dimensional key points of the object;

[0043] According to the heat Figure Two dimensional key points of the object, it is judged whether the object is located in the image;

[0044] If the object is located in the image, the step of determining the contact between the feet of the object and the sports platform according to the action data is performed.

[0045] In an embodiment, the obtaining unit is further configured to obtain a first sample set, the first sample set comprising two-dimensional coordinates of foot key points of each sample foot of a sample object in a first sample image and a contact label of the foot key points of each sample foot with the sports platform;

[0046] The processing unit is further configured to perform frame processing on the sample feet in the first sample image to obtain a frame image corresponding to each sample foot; according to the two-dimensional coordinates of the foot key points of each sample foot, the first neural network is called to respectively predict the foot key points in the frame image corresponding to each sample foot, to obtain the contact probability of the foot key points of each sample foot with the sports platform; and the first neural network is optimized according to the contact probability of the foot key points of each sample foot with the sports platform and the corresponding contact label, to obtain the trained first neural network.

[0047] In an embodiment, the obtaining unit is further configured to obtain a second sample set, the second sample set comprising a second sample image and annotation information corresponding to the second sample image, the annotation information comprising three-dimensional annotation information and two-dimensional annotation information of a sample object in the second sample image;

[0048] The processing unit is further configured to train the second neural network according to the second sample image and the corresponding annotation information, to obtain an intermediate neural network;

[0049] The obtaining unit is further configured to obtain a third sample set, the third sample set comprising a third sample image, the third sample image being at least one of the following: a second sample image in the second sample set, a sample image obtained by cropping a randomly selected second sample image, and a sample image obtained by performing color space processing on a randomly selected second sample image;

[0050] The processing unit is further configured to train the intermediate neural network according to a third sample image included in a third sample set and corresponding label information, to obtain the trained second neural network.

[0051] In an embodiment, when the processing unit trains the second neural network according to the second sample image and the corresponding label information to obtain the intermediate neural network, the processing unit can be specifically configured to:

[0052] invoke the second neural network to perform recognition processing on the second sample image to obtain action data of a sample object in the second sample image and two-dimensional key point coordinates of the sample object;

[0053] perform model loss calculation according to the action data of the sample object and the three-dimensional label information to obtain a first loss value;

[0054] perform model loss calculation according to the two-dimensional label information of the sample object and the two-dimensional key point coordinates to obtain a second loss value;

[0055] adjust network parameters of the second neural network according to the first loss value and the second loss value to obtain the intermediate neural network.

[0056] In an embodiment, the two-dimensional coordinates of the foot key point of each foot are obtained by invoking the trained third neural network to perform key point recognition processing on the target region in the label frame, and the obtaining unit is further configured to obtain a fourth sample set, the fourth sample set including a fourth sample image, a sample label frame of a target region of a sample object in the fourth sample image, and a two-dimensional coordinate label of a foot key point of a sample foot in the sample object;

[0057] The processing unit is further configured to invoke the third neural network to perform key point recognition processing on the target region of the sample object in the sample label frame to obtain the two-dimensional coordinates of the foot key point of the sample foot, and adjust network parameters of the third neural network according to the two-dimensional coordinates of the foot key point of the sample foot and the corresponding two-dimensional coordinate label to obtain the trained third neural network.

[0058] In an embodiment, when the obtaining unit obtains the second sample set, the obtaining unit can be specifically configured to:

[0059] obtain a sample image and a sample video with three-dimensional label information, and take the obtained sample image with three-dimensional label information and an image obtained by processing the sample video with three-dimensional label information as the second sample image; or

[0060] The sample image with two-dimensional label information is acquired, and a pre-trained three-dimensional construction model is used to fit the acquired sample image with two-dimensional label information to obtain three-dimensional label information corresponding to the sample image.

[0061] In the embodiment of the present application, the image to be processed can be identified to obtain action data of the object, and the contact between the feet of the object and the movement platform (such as the ground) is determined according to the action data of the object, and the action data of the object is corrected based on the contact between the feet of the object and the movement platform, and the virtual image matched with the object is displayed according to the corrected action data. The drift problem of the virtual image can be solved through the correction processing, and the presentation of the virtual image is ensured to be real and stable. BRIEF DESCRIPTION OF DRAWINGS

[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0063] Figure One is a schematic diagram of the architecture of a data processing system provided by an exemplary embodiment of the present application;

[0064] Figure Two is a flowchart of a data processing method provided by an exemplary embodiment of the present application;

[0065] Figure Three a is a structure diagram of a trained second neural network provided by an exemplary embodiment of the present application;

[0066] Figure Three b is a structure diagram of an encoding module provided by an exemplary embodiment of the present application;

[0067] Figure Three c is a schematic diagram of the foot key points of each foot provided by an exemplary embodiment of the present application;

[0068] Figure Three d is a schematic diagram of displaying an object and a virtual image provided by an exemplary embodiment of the present application;

[0069] Figure Four is a flowchart of a data processing method provided by another exemplary embodiment of the present application;

[0070] Figure Fiveis a flowchart of a model training method provided by an example embodiment of the present application.

[0071] Figure Six is a structural diagram of a pre-trained three-dimensional construction model provided by an example embodiment of the present application.

[0072] Figure Seven is a structural diagram of a data processing apparatus provided by an example embodiment of the present application.

[0073] Figure Eight is a structural diagram of a computer device provided by an example embodiment of the present application. DETAILED DESCRIPTION

[0074] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0075] First, the related technologies involved in the embodiments of the present application are briefly introduced:

[0076] I. Artificial Intelligence (AI)

[0077] Artificial intelligence is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to design and implement principles and methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0078] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.

[0079] II. Computer Vision (CV)

[0080] Computer vision is a science that studies how to make machines "see". More specifically, it refers to the use of cameras and computers to replace human eyes to identify and measure targets, and further perform image processing to make the computer processing more suitable for human observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common face recognition, fingerprint recognition, and other biometric identification technologies.

[0081] III. Machine Learning (ML)

[0082] Machine learning is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and other disciplines. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. It is applied in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.

[0083] IV. Virtual Avatar

[0084] Virtual avatar refers to a virtual person, animal, plant, etc. with a digitalized appearance, which exists in dependence on a display device and has the appearance, behavior, and thoughts of the corresponding living being. For example, a virtual person has the appearance, behavior (e.g., can speak and perform actions), and thoughts (e.g., can talk to people) of a human being. Virtual avatars have a wide range of applications, such as in the field of social communication (e.g., virtual avatar dialogue), the field of games (e.g., game characters), and the field of live streaming (e.g., virtual anchors, AI idols), etc. Users can create virtual avatars such as 3D cartoon avatars or game characters, and realize their imagination in the virtual world.

[0085] V. Virtual Avatar Driving Technology

[0086] Virtual image driving technology is a cutting-edge emerging research field. Currently, there are two main implementation methods. One is to implement through professional collection equipment for 3D scanning, motion capture, CG animation, etc., which is commonly used in the film industry, has a high cost, requires high professionalism, and has a high use threshold. The other is to use the camera components of mobile terminals such as mobile phones and tablet computers to recognize object actions through image shooting, and then drive the virtual image to perform the same action, which uses the chip computing power of mobile terminals, has a relatively low cost, requires low professionalism, has a relatively low use threshold, and is therefore more widely used, but has a relatively poor visual experience, for example, the virtual image is prone to drift when presented in a mobile terminal.

[0087] To solve the problem of drift of virtual images in the presentation process of mobile terminals and make the presentation of virtual images more realistic and stable, the embodiments of the present application provide a data processing scheme. The general principle of the data processing scheme is that in the process of driving virtual images to present according to object actions, the contact between the object's feet and the movement platform (such as the ground) is first judged (such as whether it is in contact, single-foot contact, multi-foot contact, etc.), and the object's action data is corrected based on the contact, and the corresponding virtual image is driven to display according to the corrected action data. The purpose of the correction is to optimize the foot contact, so that the virtual feet of the driven virtual image are always in contact with the movement platform, solving the problem of drift of virtual images in the presentation process, making the displayed virtual image more realistic and more stable.

[0088] Secondly, the data processing scheme provided by the embodiments of the present application can be applied to the following application scenarios:

[0089] I. Live broadcast scenario

[0090] When the live broadcast object wants to live broadcast dance, it can enter the live broadcast interface. At this time, the camera of the mobile terminal is opened. When the live broadcast object makes various dance actions on the ground (i.e. the movement platform), the camera of the mobile terminal can be called to shoot the live broadcast object on the ground in real time, and the live broadcast object and the virtual image making the same dance action with the object will be displayed in the mobile terminal. For example, when the live broadcast object stands on one foot on the ground, the virtual image displayed in the mobile terminal also stands on one foot on the ground; for example, when the live broadcast object makes a jumping action, the virtual image displayed in the mobile terminal also makes a jumping action; for example, when the object stands on both feet on the ground, the virtual image displayed in the mobile terminal also stands on both feet on the ground.

[0091] II. Game scenario

[0092] When the game object wants to interact with the game, the game interface can be entered, at which time the camera of the mobile terminal is opened. When the game object makes various game actions on the ground, the camera of the mobile terminal can be called to take real-time photos of the game object, and the game object and the virtual image making the same game action as the game object will be displayed in the game interface. For example, when the game object makes a tree cutting action on the ground, the game object making the tree cutting action on the ground will be displayed in the game interface, and the virtual image displayed in the game interface will also make the tree cutting action on the ground. For another example, when the game object makes a boxing action with both feet standing on the ground, the game object making the boxing action with both feet standing on the ground will be displayed in the game interface, and the virtual image displayed at the same time will also make the boxing action with both feet standing on the ground.

[0093] III. Vehicle monitoring scenario

[0094] During the driving of the vehicle, the driving process of the vehicle can be photographed to obtain an image to be processed; the image to be processed is recognized to obtain action data of the vehicle; whether the feet (i.e., the tires) of the vehicle are in contact with the ground during the driving process can be determined according to the action data, and when it is determined that there is contact with the ground, the action data is corrected, and then a virtual image (i.e., a virtual vehicle) matched with the vehicle can be displayed according to the corrected action data. In this way, it can be quickly checked whether there is a risk of the vehicle taking off during the driving process.

[0095] The data processing scheme of the embodiments of the present application can be combined with artificial intelligence. For example, during the process of recognizing the photographed image to obtain the action data of the object, the image can be recognized by means of computer vision technology in artificial intelligence. For another example, when it is determined whether the feet of the object are in contact with the movement platform according to the action data, the contact probability of the feet and the movement platform can be determined by machine learning, and whether the feet of the object are in contact with the movement platform can be determined based on the contact probability of the feet and the movement platform.

[0096] Next, the data processing system provided by the present application is introduced. Please refer to Figure One , Figure One A data processing system is provided for the embodiments of the present application. As shown in Figure One , the data processing system can include a computer device 101 and a server 102. The number of computer devices is not limited by the present application, and of course, the number of servers can also be multiple, and the number of servers is still not limited by the present application. The computer device 101 and the server 102 in the data processing system can be directly or indirectly connected through wired or wireless communication. Among them:

[0097] The computer device 101 runs various clients, which can be a social client, a client dedicated to live broadcast, and the like. The computer device 101 is configured with a camera component (such as a camera), which can be invoked to capture an object on a motion platform; the object can have a virtual image in the client that can make the same actions, expressions, and the like as the object. The computer device can be a terminal device, which can include but is not limited to a mobile terminal (such as a smartphone), a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, a smart wearable device, a flying device, a smart home appliance, and the like.

[0098] The server 102, which can correspond to the client, provides technical support for the services provided by the client. The server 102 can store login data (such as a login account and a nickname) of an object logging into the client, image data (such as appearance, clothing, and hair accessories) of a virtual image, and the like. The server 102 can be a standalone physical server, a server cluster or a distributed system formed by multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.

[0099] In one embodiment, the data processing flow is described by taking the computer device 101 and the server 102 as an example.

[0100] When the object performs some actions on the motion platform, for example, when the object wants to live broadcast or video chat, the object can log into the corresponding client through a login account and live broadcast or video chat in the client; at this time, the camera component of the computer device can be invoked to capture the object on the motion platform to obtain an image to be processed. The object can be a person, an animal, an object, or the like performing actions. The actions performed by the object can include but are not limited to dancing, singing, conversing, live broadcasting, playing games, sports, and the like. The motion platform refers to a platform or a plane that can carry the object to perform actions, which can include but is not limited to the ground, a stage, or other platforms that can carry the object to perform actions, and the present application does not limit the motion platform.

[0101] ②The computer device performs recognition processing on the photographed image to obtain action data of the object, and the action data can include at least one of position data and posture data. The position data can be used to describe the position of the object in a three-dimensional space, and the posture data can be used to describe the posture assumed by the object in the movement process (for example, the posture data can be used to describe that the action posture performed by the object is to spread the hands).

[0102] ③The computer device determines the contact of the feet of the object with the ground according to the action data of the object. If it is determined that the feet of the object are in contact with the movement platform, an optimization algorithm can be used to correct the action data of the object in the three-dimensional space, for example, to correct the position data or to correct the posture data. The number of feet of the object can be N, and the contact of the feet of the object with the ground includes the following: the feet of the N feet are in contact with the movement platform; only the feet of M feet of the N feet are in contact with the movement platform, and the feet of the other N-M feet are not in contact with the movement platform. Correspondingly, the correction of the action data of the object using the optimization algorithm can include the following two cases: when the feet of the N feet of the object are in contact with the movement platform, the feet of each foot are kept in contact with the movement platform, and the posture data of the object is corrected; when the feet of the M feet of the N feet of the object are in contact with the movement platform, the feet of the M feet are kept in contact with the movement platform, and the position data of the object is corrected. M and N are both positive integers and M is less than or equal to N.

[0103] ④After the action data is corrected, the computer device 101 can obtain the image data of the virtual image matched with the object from the server 102, the image data of the virtual image can be used to generate the appearance of the virtual image, and then the computer device 101 can map the image data and the corrected action data to generate a virtual image matched with the object, and render the virtual image through a rendering engine to display the virtual image in the client. The virtual feet of the displayed virtual image are in contact with the movement platform, which solves the problem of foot drift of the virtual image, and the virtual image can make the same action as the object according to the corrected action data.

[0104] ⑤If it is determined that the feet of the object are not in contact with the movement platform, for example, the feet of the N feet of the object are not in contact with the ground, the object can be in a running or jumping state, and at this time, the optimization algorithm is not needed to correct the action data of the object. The computer device 101 can obtain the image data of the virtual image matched with the object from the server 102, and then the computer device 101 can map the image data and the action data to generate a virtual image matched with the object, and render the virtual image through a rendering engine to display the virtual image in the client.

[0105] By means of the data processing system, the image to be processed can be recognized to obtain the action data of the object on the motion platform. In the process of driving the virtual image to present according to the action data of the object, it is first determined whether the feet of the object are in contact with the motion platform (such as the ground) (such as whether the feet are in contact, one-foot contact, multi-foot contact, etc.), and the action data of the object is corrected based on the contact condition, and the corresponding virtual image is driven to display according to the corrected action data. The purpose of the correction is to optimize the foot stepping, so that the virtual feet of the driven virtual image are always in contact with the motion platform, solving the problem of drift of the virtual image in the presentation process, so that the displayed virtual image is more realistic and more stable.

[0106] Next, the data processing method is described in detail. Please refer to Figure Two , Figure Two A flowchart of a data processing method provided by an embodiment of the present application. The data processing method can be executed by a computer device; the data processing method described in this embodiment can include the following steps S201-S205.

[0107] S201, obtaining an image to be processed, the image being obtained by photographing an object on a motion platform.

[0108] The motion platform can be the ground, a stage, or any platform (such as a table) capable of carrying the object to move. The object can be a person (such as an anchor, an ordinary user, a dancer), an animal (such as a cat or a dog), a vehicle, etc. The image to be processed can include all features of the object or part of the features of the object, such as an image to be processed including only the upper body of a person or an image to be processed including all features of a person (i.e. including the upper body and the lower body of the person). The number of images to be processed can be one or more, which is not limited by the embodiment of the present application.

[0109] In some optional embodiments, the computer device can obtain the image to be processed by real-time image acquisition of the object on the motion platform through the camera component configured by the computer device (such as a camera carried by the computer device, an externally connected camera device, etc.). Or, the image or video obtained by photographing the object on the motion platform is stored in the local space. When the image obtained by photographing the object on the motion platform is stored in the local space, the computer device can directly obtain the image of the object on the motion platform from the local space as the image to be processed. When the video obtained by photographing the object on the motion platform is stored in the local space, the computer device can obtain the video of the object on the motion platform from the local space, and obtain any frame of image from the video as the image to be processed.

[0110] S202, performing recognition processing on the image to obtain action data of the object.

[0111] The motion data may include at least one of the following: position data and posture data. Position data may include the coordinates of the object's center of gravity in three-dimensional space; posture data may include a rotation matrix formed by the rotation angles of the object's various rotational joints; rotational joints may include, but are not limited to: pelvis, left and right thighs, left and right knees, left and right ankles, left and right feet, spine 123 (three points in total), left and right scapulae, left and right shoulders, left and right elbows, left and right wrists, left and right palms, neck, head, etc.

[0112] In some alternative embodiments, image recognition processing can be implemented by invoking a neural network. In one feasible implementation, this neural network can be trained using the PyTorch framework and deployed on a computer device using the TNN framework. The trained second neural network can perform relatively accurate image recognition processing. The computer device can invoke the trained second neural network to process the image and obtain the object's motion data. The trained neural network follows the camera assumption of weak perspective projection, and the image to be processed is assigned a fixed virtual focal length (e.g., 5000, 6000, etc.). Under this camera assumption, the object's position data can be represented by a three-dimensional array (x, y, z), where (x, y, z) represents the coordinates of the object's center of gravity in three-dimensional space, or the object's position data. The pose data is represented by the rotation angles of each rotation joint, where the rotation angle refers to the rotation angle in three-dimensional space. Three-dimensional space refers to the camera coordinate system.

[0113] The structure of the trained second neural network can be as follows: Figure Three a As shown, the trained second neural network may include a fully connected network, a regression module, and a discriminator. The fully connected network can be used to identify and process the features of the object, the regression module can process the identified object features to obtain action data, and the discriminator can be used to determine whether the object is within the image. This fully connected network can be configured as follows: Figure Three b As shown, this fully connected network can include convolutional layers, pooling layers, non-linear activation functions, upsampling layers, and concatenation layers. Convolutional layers are used to extract features of objects from the image to be processed; pooling layers (also known as downsampling) are used to compress the extracted features; upsampling layers are used to restore the compressed features; and concatenation layers combine the compressed and restored features. Figure Three a In This means that (θ, β) can be mapped to the three-dimensional shape and pose of an object through SMPL (Skinned Multi-Person Linear). Figure Three a (3D shape and pose data of object 30).

[0114] In one embodiment, the output of the trained second neural network can further include heat map of the object in addition to the motion data of the object. Figure Two dimensional key points according to the heat map of the object. Figure Two The three-dimensional key points can be used to determine whether the object is in the image. The computer device can call the trained second neural network to perform recognition processing on the image, and obtain the three-dimensional key points included in the object. The three-dimensional key points included in the object are projected to obtain the heat map of the object. Figure Two dimensional key points. Then, the computer device can determine whether the object is in the image according to the heat map of the object. Figure Two dimensional key points. Then, the computer device can determine whether the object is in the image according to the heat map of the object.

[0115] S203, determining the contact between the feet of the object and the motion platform according to the motion data.

[0116] The number of feet of the object is N; for example, the object is a human being, and the number of feet of the object is 2; for another example, the object is an animal (such as a cat), and the number of feet of the object is 4. The contact between the feet of the object and the motion platform includes any one of the following: the feet of the N feet of the object are in contact with the motion platform; the feet of M feet of the N feet of the object are in contact with the motion platform, but the feet of the remaining N-M feet of the object are not in contact with the motion platform, M and N are positive integers and M is less than or equal to N; the feet of the N feet of the object are not in contact with the motion platform. For example, the object is a human being, and the contact between the feet of the object and the motion includes any one of the following: the feet of the two feet of the object are in contact with the motion platform (such as the ground); the feet of one foot of the object are in contact with the motion platform (such as the ground), and the feet of the other foot are not in contact with the motion platform; the two feet of the object are not in contact with the motion platform (i.e., the object is in a jumping state).

[0117] In some optional embodiments, the computer device can directly determine the contact status between the object's feet and the motion platform based on motion data. The computer device can acquire the object's skeletal data, calculate the coordinates of each rotational joint in three-dimensional space based on posture data and the object's skeletal data; then, based on position data and a fixed focal length, project the coordinates of each rotational joint in three-dimensional space to obtain the two-dimensional coordinates of each rotational joint; the computer device can obtain the two-dimensional coordinates of the object from the waist to the feet from the two-dimensional coordinates of each rotational joint, calculate the relative position of the foot's two-dimensional coordinates with the ground, and determine whether the relative position is less than or equal to a position threshold; if the relative position is less than or equal to the position threshold, it can be determined that the foot is in contact with the motion platform; if the relative position is greater than the position threshold, it can be determined that the foot is not in contact with the motion platform. The position threshold can be set according to requirements.

[0118] In some alternative embodiments, determining whether a foot is in contact with the motion platform using foot coordinates may be inaccurate. Therefore, in this embodiment, more precise foot key points can be introduced to more accurately determine whether the foot is in contact with the motion platform. In this case, the computer device determines the contact status of the object's feet with the motion platform based on motion data as follows: based on the motion data, first determine the two-dimensional coordinates of the foot key points of each foot; then, based on the two-dimensional coordinates of the foot key points of each foot, predict the contact probability of each foot key point with the motion platform; and based on the contact probability of each foot key point with the motion platform, determine the contact status of each foot with the motion platform. The number of foot key points for each foot can be one or more; for example, if the object is a person, the number of foot key points for each foot is two, such as... Figure Three c As shown, the key points of the right foot can include: heel key point A and toe key point C; the key points of the left foot can include heel key point B and toe key point D. Of course, the key points of each foot can be customized as needed. For example, the key points of each foot can include: arch key point, heel key point, and toe key point.

[0119] The computer equipment can determine whether the contact probability between the key points of each foot and the motion platform is greater than a contact threshold. Taking any one of N feet (i.e., the target foot) as an example, if there is a key point on the foot with a contact probability greater than the contact threshold, then the contact status between the target foot and the motion platform is determined as: the target foot is in contact with the motion platform; if there is no key point on the foot with a contact probability greater than the contact threshold, then the contact status between the target foot and the motion platform is determined as: the foot is not in contact with the motion platform. This method can determine the contact status of each foot with the motion platform.

[0120] For example, N is 2, which respectively represents the left foot and the right foot, and the foot key points of each foot include: a heel key point and a toe key point; the computer device can respectively determine whether the contact probability of the heel key point in the right foot with the motion platform is greater than the contact threshold, and whether the contact probability of the toe key point in the right foot with the motion platform is greater than the contact threshold; if there is a foot key point corresponding to a contact probability greater than the contact threshold in the right foot, it is determined that the contact between the foot of the right foot and the motion platform is: the foot of the right foot is in contact with the motion platform. It should be understood that, as long as any one of the heel key point and the toe key point in the right foot is in contact with the ground, it can be determined that the foot of the right foot is in contact with the motion platform. Similarly, the contact between the foot of the left foot and the motion platform can also be determined according to the heel key point and the toe key point of the left foot, which will not be described here.

[0121] In the method, the two-dimensional coordinates of the foot key points of each foot are determined according to the action data, which can be: obtaining object skeleton data, calculating the coordinates of each rotation joint of the object in a three-dimensional space according to the posture data and the object skeleton data, and then respectively projecting the coordinates of each rotation joint in the three-dimensional space to obtain the two-dimensional coordinates of each rotation joint according to the position data and the fixed focal length (i.e. the virtual focal length described above), and then determining the two-dimensional coordinates of the foot key points of each foot according to the two-dimensional coordinates of each rotation joint.

[0122] As an implementation manner, the foot key points are included in each rotation joint, and the computer device can directly obtain the two-dimensional coordinates of the foot key points of each foot from the two-dimensional coordinates of each rotation joint. As another implementation manner, the foot key points are not included in the two-dimensional coordinates of each rotation joint, but only the approximate foot coordinates of each foot are included. In this case, key point recognition processing needs to be performed on each foot to obtain the two-dimensional coordinates of the foot key points of each foot. The specific implementation manner of the computer device for determining the two-dimensional coordinates of the foot key points of each foot according to the two-dimensional coordinates of each rotation joint can be: calculating a marking box of a target region of the object according to the posture data, and performing key point recognition processing on the target region in the marking box according to the two-dimensional coordinates of the rotation joints included in the target region to obtain the two-dimensional coordinates of the foot key points of each foot. The target region of the object refers to a region formed by the N feet of the object, for example, the object is a person, and the target region is the lower body of the person; the marking box can be a rectangular box, a circular box, an irregular box, etc., which can be used to completely frame the target region of the object. For example, the object is a person, the target region is the lower body of the person, and the computer device can determine a rectangular box (x, y, w, h) of the lower body of the person according to the posture data, wherein x and y represent the coordinates of the upper left corner of the rectangular box, w represents the width, and h represents the height.

[0123] The specific implementation manner of predicting the contact probability of each foot key point of each foot with the motion platform according to the two-dimensional coordinates of the foot key points of each foot can be: performing frame selection processing on the N feet in the image respectively to obtain a frame-selected image corresponding to each foot; and predicting the foot key points in the N frame-selected images respectively according to the two-dimensional coordinates of the foot key points of each foot to obtain the contact probability of each foot key point of each foot with the motion platform.

[0124] It should be noted that after determining that the object is not located in the image in step S202, the marking frame of the target region can be filled with a target value, which can be 0, 1, etc. The target value is filled in the marking frame to achieve the contact of the foot of the object with the motion platform without determining the object.

[0125] S204, correcting the action data based on the contact of the foot of the object with the motion platform.

[0126] Wherein, for different contact conditions of the foot of the object with the motion platform, the position data and / or posture data in the action data can be corrected.

[0127] As an implementation manner, if the contact condition includes that the feet of N feet of the object are in contact with the motion platform, the computer device can correct the posture data of the object. The computer device can keep the height of each foot key point in the three-dimensional space unchanged, and adjust the rotation angle of the knee joint and the thigh joint of the object. By keeping the height of each foot key point in the three-dimensional space unchanged, the object can be controlled to be in contact with the ground at all times, and by adjusting the rotation angle of the knee joint and the thigh joint of the object, the virtual image can be controlled to make the same posture as the object. When the number of foot key points of each foot is multiple, the computer device can keep the lowest point of the multiple foot key points of each foot in the three-dimensional space unchanged, and adjust the rotation angle of the knee joint and the thigh joint of the object, so that the body of the object is not moved and the leg is slightly corrected. The lowest point of the multiple foot key points of each foot can be determined according to the height of each foot key point in the three-dimensional space. If the three-dimensional space is defined as x-axis, y-axis and z-axis, then the height in the three-dimensional space can be determined by z. For example, as shown in Figure Three c The heel key point A and the toe key point C are in contact with the ground; the coordinates of the heel key point A in the three-dimensional space are (1, 2, 1), and the coordinates of the toe key point C in the three-dimensional space are (0.5, 1, 0.5). The computer device can determine that the lowest point of the multiple foot key points of the right foot is the toe key point C, and the computer device can keep the height of the toe key point C in the three-dimensional space unchanged.

[0128] As another implementation manner, if the contact condition includes that M feet of N feet of the object are in contact with the motion platform, but the remaining N-M feet are not in contact with the motion platform, the position data of the object is corrected. The computer device can keep the coordinates of the foot key points of the M feet in the three-dimensional space unchanged, and adjust the coordinates of the center of gravity of the object in the three-dimensional space. By keeping the coordinates of the foot key points in the three-dimensional space, the object can be controlled to be always in contact with the ground when moving, and the virtual image can be controlled to make the same gesture (such as movement) as the object.

[0129] In one embodiment, if the contact condition includes that the feet of N feet of the object are not in contact with the motion platform, it is determined that the object is in a jumping state or a running state, and the computer device can directly map the action data to the virtual image to obtain a virtual image matched with the object, so that the virtual image mapped by the action data can control the virtual image to make the same gesture as the object.

[0130] S205, display the virtual image matched with the object according to the corrected action data, and the virtual feet of the virtual image are in contact with the motion platform.

[0131] The virtual image can be a 3D character image, an animation character image, an animal, etc. In a specific implementation, the computer device can map the corrected action data to the virtual image by using a drawing tool to obtain a virtual image matched with the object, and use a rendering engine to render and process the virtual image matched with the object to display the virtual image matched with the object. The drawing tool can be an OpenGL tool. The virtual image mapped by the corrected action data can control the virtual feet of the virtual image to be in contact with the motion platform and make the same gesture as the object. Alternatively, the object can also be displayed in the display interface provided by the computer device while displaying the virtual image matched with the object. For example, as shown in FIG. 3, the object 31 and the virtual image 32 matched with the object are displayed in the display interface provided by the computer device, the displayed object 31 and the virtual image 32 matched with the object make the same gesture (i.e., one hand makes a waving gesture, and the other hand makes a waist-crossing gesture), and the displayed object 31 and the virtual image 32 matched with the object are both in contact with the motion platform 33. Figure Three d

[0132] ​In the embodiment of the present application, the computer device can obtain an image to be processed, the image being obtained by photographing an object on a motion platform; the image is identified to obtain motion data of the object; the contact between the feet of the object and the motion platform is determined according to the motion data; the motion data is corrected based on the contact between the feet of the object and the motion platform; and a virtual image matched with the object is displayed according to the corrected motion data, the virtual feet of the virtual image being in contact with the motion platform. The virtual feet of the virtual image can be in contact with the motion platform, the problem of the virtual image producing foot drift is solved, and the displayed virtual image is more realistic.

[0133] Please refer to Figure Four , Figure Four A flowchart of a data processing method provided in the embodiment of the present application is shown. The data processing method can be executed by a computer device; the data processing method described in the embodiment can include the following steps S401-S410.

[0134] S401, an image to be processed is obtained, the image being obtained by photographing an object on a motion platform.

[0135] S402, a trained second neural network is called to identify the image, and motion data of the object is obtained, the motion data including at least one of position data and posture data, the posture data including a rotation matrix formed by rotation angles of each rotation joint of the object.

[0136] S403, object skeleton data is obtained, and coordinates of each rotation joint of the object in a three-dimensional space are calculated according to the posture data and the object skeleton data.

[0137] S404, the coordinates of each rotation joint of the object in the three-dimensional space are projected respectively according to the position data and a fixed focal length, and two-dimensional coordinates of each rotation joint are obtained.

[0138] S405, a marking box of a target region in the object is calculated according to the posture data; the feet of the object include N feet, and the target region can include the N feet of the object.

[0139] S406, a trained third neural network is called to identify the target region in the marking box according to the two-dimensional coordinates of the rotation joints included in the target region, and two-dimensional coordinates of foot key points of each foot are obtained.

[0140] In the key point recognition processing of the target region in the marking box, a neural network can be introduced to accurately recognize the key points of the feet. Specifically, the computer device can call the trained third neural network to perform key point recognition processing on the target region of the object in the marking box to obtain the two-dimensional coordinates of the key points of the feet of each foot. The trained third neural network can include a convolutional layer, a pooling layer, a nonlinear activation function, an up-sampling layer, and a splicing layer. The trained third neural network forms a fully connected network structure through the assembly of different layers. The convolutional layer can be used to extract object features of the target region in the marking box, the pooling layer can be used to compress the object features extracted by the convolutional layer, the up-sampling layer can be used to restore the compressed object features, and the splicing layer can be used to splice the compressed object features and the restored object features. The nonlinear activation function is used to process the features in the splicing layer to obtain the two-dimensional coordinates of the key points of the feet of each foot.

[0141] Before calling the trained third neural network, the PyTorch framework can be used for training and deployed in the computer device in the TNN framework. The computer device can obtain a fourth sample set, which includes fourth sample images, sample marking boxes of target regions of sample objects in the fourth sample images, and two-dimensional coordinate labels of key points of sample feet in the sample objects. Then, the computer device calls the third neural network to perform key point recognition processing on the target region of the sample object in the sample marking box to obtain the two-dimensional coordinates of the key points of the sample feet. Then, according to the two-dimensional coordinates of the key points of the sample feet and the corresponding two-dimensional coordinate labels, the network parameters of the third neural network are adjusted to obtain the trained third neural network. The number of sample feet can be N. The process of adjusting the network parameters of the third neural network to obtain the trained third neural network according to the two-dimensional coordinates of the key points of the sample feet and the corresponding two-dimensional coordinate labels can be: using an L2 loss algorithm to calculate the model loss according to the two-dimensional coordinates of the key points of each sample foot and the corresponding two-dimensional coordinate labels to obtain a model loss value, and adjusting the network parameters of the third neural network according to the model loss value to obtain the trained third neural network. The L2 loss algorithm can also be referred to as L2 loss or MSE (Mean Square Error) loss.

[0142] S407, respectively, frame the N feet in the image to obtain a frame image corresponding to each foot.

[0143] In an embodiment, the computer device can perform bounding box processing on each foot according to the foot key points of each foot, to obtain a bounding box image corresponding to each foot. Taking a person as an object, the computer device can determine a square box corresponding to the left foot according to the foot key points of the left foot, and then perform bounding box processing on the left foot by using the square box, to obtain a bounding box image corresponding to the left foot. Similarly, the right foot is processed by using a square box corresponding to the right foot, to obtain a bounding box image corresponding to the right foot. The center of the square box corresponding to each foot is the center of the heel key point and the toe key point of each foot, and the side length is L times the length of the line connecting the heel key point and the toe key point, where L can be 1.5, 2, etc., which is not limited in the present application.

[0144] In another embodiment, for any one of the N feet, a bounding box can also be directly used to perform bounding box processing on the foot, to obtain a bounding box image corresponding to the foot. It should be noted that the bounding box needs to frame the entire foot, so that the foot key points of each foot can be predicted to a certain extent, which can help determine the contact between the foot of the object and the motion platform.

[0145] S408, according to the two-dimensional coordinates of the foot key points of each foot, calling the trained first neural network to predict the foot key points in the N bounding box images, to obtain the contact probability of the foot key points of each foot with the motion platform.

[0146] As an implementation manner, the computer device can call the trained first neural network to predict the foot key points in the N bounding box images according to the two-dimensional coordinates of the foot key points of each foot, to obtain the contact probability of the foot key points of each foot with the motion platform. For example, the object is a person, and the N feet include a left foot and a right foot. The number of foot key points of each foot is multiple, which are a heel key point and a toe key point. According to the two-dimensional coordinates of the foot key points of the left foot, the computer device can call the trained first neural network to predict the foot key points in the bounding box image corresponding to the left foot, to obtain the contact probability of the foot key points of the left foot with the motion platform. Then, the trained first neural network is called to predict the foot key points in the bounding box image corresponding to the right foot, to obtain the contact probability of the foot key points of the right foot with the motion platform. At this time, the contact probability of the foot key points of each foot with the motion platform can include the contact probability of the heel key point with the motion platform and the contact probability of the toe key point with the motion platform.

[0147] The trained first neural network can be a classification model, and whether each foot key point contacts the motion platform can be a binary classification state. The trained first neural network can include a convolutional layer, a pooling layer, a nonlinear activation function (such as softmax), and a fully connected layer. The convolutional layer can be used to extract features from the bounding image, the pooling layer can be used to compress the extracted features, the fully connected layer can be used to connect the features processed by the pooling layer, and the nonlinear activation function can output the contact probability of each foot key point with the motion platform. The trained first neural network forms a fully connected network structure through the assembly of different layers. The contact probability is compared with the contact threshold to determine whether the foot key point contacts the motion platform. It should be understood that if the number of the object's feet is N, the trained first neural network needs to be called N times, that is, each time the trained first neural network is called, only one foot of the N feet is predicted. Of course, for N feet, the computer device can call N trained first neural networks in parallel, and each trained first neural network predicts one foot.

[0148] When training the first neural network, PyTorch framework can be used for training, and TNN framework can be used for deployment in the computer device. Specifically, the computer device can obtain a first sample set, which includes the two-dimensional coordinates of the foot key points of each sample foot of a sample object in a first sample image and the contact labels (i.e., whether the foot key points contact the motion platform) of the foot key points of each sample foot with the motion platform. The sample feet in the first sample image are respectively processed to obtain the bounding images corresponding to each sample foot. The first neural network is called to predict the foot key points in the bounding images corresponding to each sample foot according to the two-dimensional coordinates of the foot key points of each sample foot, to obtain the contact probabilities of the foot key points of each sample foot with the motion platform. The first neural network is optimized according to the contact probabilities of the foot key points of each sample foot with the motion platform and the corresponding contact labels, to obtain the trained first neural network. The specific implementation of optimizing the first neural network according to the contact probabilities of the foot key points of each sample foot with the motion platform and the corresponding contact labels to obtain the trained first neural network can be: using an L2 loss algorithm to calculate the model loss according to the contact probabilities of the foot key points of each sample foot with the motion platform and the corresponding contact labels, to obtain a model loss value, adjusting the network parameters of the first neural network according to the model loss value, and obtaining the trained first neural network.

[0149] S409、According to the contact probability of each foot key point with the motion platform, the contact between the foot of the object and the motion platform is determined.

[0150] S410, based on the contact situation of the feet of the object with the motion platform, correcting the action data, and displaying a virtual image matched with the object according to the corrected action data.

[0151] The specific implementation of steps S409-S410 can refer to the specific implementation of steps S203-S205 in the above Figure Two The specific implementation of steps S409-S410 can refer to the specific implementation of steps S203-S205 in the above

[0152] In the embodiment of the application, the trained second neural network is called to perform recognition processing on the obtained to-be-processed image, and action data of the object is obtained; then, based on the posture data and the object skeleton data in the action data, coordinates of each rotation joint of the object in a three-dimensional space are calculated, and based on the position data in the posture data and the fixed focal length, the coordinates of each rotation joint in the three-dimensional space are respectively projected to obtain two-dimensional coordinates of each rotation joint. Then, based on the two-dimensional coordinates of the rotation joints included in the target region, the trained third neural network is called to perform key point recognition processing on the target region in the marking box, and two-dimensional coordinates of foot key points of each foot are obtained; then, based on the two-dimensional coordinates of the foot key points of each foot, the trained first neural network is called to respectively predict the foot key points in the N frame images, to obtain a contact probability of the foot key points of each foot with the motion platform, and based on the contact probability of the foot key points of each foot with the motion platform, a contact situation of the feet of the object with the motion platform is determined. Then, based on the contact situation of the feet of the object with the motion platform, the action data is corrected, and a virtual image matched with the object is displayed according to the corrected action data. Adopting different neural networks to respectively perform recognition processing on the image, to identify the foot key points of each foot in the image, and to predict the contact probability of the foot key points of each foot with the motion platform can improve the accuracy of determining the contact situation of the foot key points of each foot with the motion platform, and thus can improve the correction accuracy of the action data, so that the virtual image is in contact with the motion platform, the problem of foot drift of the virtual image is solved, and the displayed virtual image is more realistic.

[0153] The embodiment of the application further provides a model training method, which can be executed by a computer device, and can include steps S501-S502.

[0154] S501, acquire a second sample set, the second sample set comprising a second sample image and annotation information corresponding to the second sample image, the annotation information comprising three-dimensional annotation information and two-dimensional annotation information of a sample object in the second sample image. The three-dimensional annotation information can comprise at least one of position data labels and pose data labels of the sample object in a three-dimensional space. The two-dimensional annotation information can comprise position data (such as two-dimensional coordinates) of a rotation joint of the sample object in a two-dimensional space. The number of second sample images can be one or more.

[0155] In one embodiment, the computer device can acquire sample images with three-dimensional annotation information and sample videos, and take the acquired sample images with three-dimensional annotation information as second sample images and take images obtained by processing the sample videos with three-dimensional annotation information as second sample images. The images obtained by processing the sample videos with three-dimensional annotation information can be any one or more frames of the sample videos.

[0156] The computer device can acquire sample images with three-dimensional annotation information and sample videos from a target position. The target position can be a storage device dedicated to storing sample images with three-dimensional annotation information, or the target position can be a three-dimensional dataset such as human3.6m or ochuman. In addition, the computer device can also collect sample images with three-dimensional annotation information and sample videos through a motion capture device.

[0157] In one embodiment, acquiring the second sample set can also be: acquiring sample images with two-dimensional annotation information, and using a pre-trained three-dimensional construction model to perform fitting processing on the acquired sample images with two-dimensional annotation information to obtain three-dimensional annotation information corresponding to the sample images, and taking the sample images corresponding to the obtained three-dimensional annotation information as second sample images. The computer device can acquire sample images with two-dimensional annotation information from a two-dimensional dataset and sample images downloaded from a network.

[0158] The pre-trained three-dimensional construction model can be obtained by using an EFT (Exemplary Fine-Tuning) algorithm. The structure of the pre-trained three-dimensional construction model can be as follows: Figure SixAs shown, when training the three-dimensional construction model, a large number of images with two-dimensional annotation information and three-dimensional annotation information (such as pose data) can be input into the three-dimensional construction model, and then the computer device can call the three-dimensional construction model to perform three-dimensional modeling processing on the sample object in the image with two-dimensional annotation information, to obtain the pose data of the sample object in the three-dimensional space, and determine the three-dimensional key points of the sample object in the three-dimensional space according to the pose data in the three-dimensional space, and perform two-dimensional projection on the three-dimensional key points of the sample object to obtain the two-dimensional key points of the sample object; then loss calculation is performed according to the pose data and the three-dimensional annotation information of the sample object to obtain a pose loss value; and loss calculation is performed according to the two-dimensional key points and the two-dimensional annotation information to obtain a projection loss value, and a model loss value is obtained according to the pose loss value and the projection loss value, that is, L EFT = L 投影 + L 姿态 ; then the three-dimensional construction model is exemplarily fine-tuned according to the model loss value to obtain an intermediate three-dimensional construction model; through multiple (such as 100 times, 1000 times) iteration training on the intermediate three-dimensional construction model, a pre-trained three-dimensional construction model can be obtained. Wherein, L EFT represents the model loss value, L 投影 represents the projection loss value, and L 姿态 represents the pose loss value.

[0159] It should be understood that the second sample set can include one or more of the following: sample images with three-dimensional annotation information, images obtained by processing sample videos with three-dimensional annotation information, and sample images corresponding to three-dimensional annotation information obtained by the pre-trained three-dimensional construction model.

[0160] S502, according to the second sample image and the corresponding annotation information, the second neural network is iteratively trained N times to obtain a trained second neural network. Wherein, the value of N can be 150000, 20000, etc. The second neural network is trained using the PyTorch framework and deployed in the computer device in the TNN framework. The architecture of the second neural network can be as shown in Figure Three a .

[0161] In a specific implementation, the computer device can train the second neural network according to the second image and the corresponding annotation information to obtain an intermediate neural network; obtain a third sample set, the third sample set including third sample images, the third sample images being at least one of the following: a second sample image in the second sample set, a sample image obtained by cropping a randomly selected second sample image, and a sample image obtained by performing color space processing on a randomly selected second sample image; for example, the computer device can randomly select a region containing an object according to the object two-dimensional key points from the second sample set, crop the region, and add the cropped sample image to the third sample set, in this way, the number of sample images can be increased, and data augmentation can be realized; or, the computer device can randomly use color space to perform data enhancement on the second sample image (for example, add a deep color to the sample object in the second sample image), and through data enhancement and data augmentation, a more accurate neural network model can be trained.

[0162] It should be understood that in each iteration training, data augmentation is performed by randomly selecting a region containing an object according to the human two-dimensional key points from the second sample set and cropping the region, or data enhancement is randomly performed by using color space, a new third sample set is obtained, and the intermediate neural network is trained by using the new third sample set to obtain the final trained second neural network.

[0163] In the training of the second neural network, the model is supervised training. The specific implementation process of training the second neural network according to the second sample image and the corresponding annotation information to obtain the intermediate neural network can be: calling the second neural network to perform recognition processing on the second sample image to obtain action data of a sample object in the second sample image and two-dimensional key point coordinates of the sample object obtained by projecting three-dimensional key coordinates of the sample object; performing model loss calculation according to the action data of the sample object and the three-dimensional annotation information to obtain a first loss value; performing model loss calculation according to the two-dimensional annotation information of the sample object and the two-dimensional key point coordinates to obtain a second loss value; adjusting network parameters of the second neural network according to the first loss value and the second loss value to obtain the intermediate neural network. In the model training process, the first loss value and the second loss value are both L2 loss, and the network parameters of the second neural network are adjusted according to the first loss value and the second loss value to obtain the intermediate neural network. In a specific implementation, the computer device can calculate a total loss value of the model according to the first loss value and the second loss value, and adjust the network parameters of the second neural network according to the total loss value to obtain the intermediate neural network. The calculation formula of the total loss value of the model is as follows:

[0164] Loss=L旋转 +L 投影

[0165] wherein, Loss represents a total loss value of the model, L 旋转 represents a first loss value, L 投影 represents a second loss value.

[0166] In the embodiments of the present application, a second sample set is obtained, the second sample set including second sample images and label information corresponding to the second sample images, and the second neural network is iteratively trained N times according to the second sample images and the corresponding label information to obtain a trained second neural network, by which the action data of the object in the image can be accurately identified in real time.

[0167] Please refer to Figure Seven , Figure Seven is a structural schematic diagram of a data processing apparatus provided by the embodiments of the present application; the data processing apparatus can be a computer program (including program code) running in a computer device, for example, the data processing apparatus can be an application software in the computer device; the data processing apparatus can be used to execute part or all of the steps in the method embodiments shown in Figure Two and Figure Four . Please refer to Figure Seven , the data processing apparatus includes the following units:

[0168] The acquisition unit 701 is configured to acquire an image to be processed, the image being obtained by photographing an object on a motion platform;

[0169] The processing unit 702 is configured to perform identification processing on the image to obtain action data of the object;

[0170] The processing unit 702 is further configured to determine a contact condition of a foot of the object with the motion platform according to the action data;

[0171] The processing unit 702 is further configured to correct the action data based on the contact condition of the foot of the object with the motion platform;

[0172] The processing unit 702 is further configured to display a virtual image matched with the object according to the corrected action data, a virtual foot of the virtual image being in contact with the motion platform.

[0173] In one embodiment, the action data includes at least one of the following: position data and posture data; the position data includes coordinates of the center of gravity of the object in a three-dimensional space; the posture data includes a rotation matrix formed by rotation angles of each rotation joint of the object.

[0174] The number of the feet of the object is N; the contact of the feet of the object with the motion platform includes any one of the following: the feet of the N feet of the object are in contact with the motion platform; the feet of M feet of the N feet of the object are in contact with the motion platform, but the feet of the remaining N-M feet are not in contact with the motion platform, M and N are positive integers and M is less than or equal to N.

[0175] In an embodiment, the processing unit 702, when determining the contact of the feet of the object with the motion platform according to the action data, can be specifically configured to:

[0176] determining the two-dimensional coordinates of the foot key points of each foot according to the action data;

[0177] predicting the contact probability of the foot key points of each foot with the motion platform according to the two-dimensional coordinates of the foot key points of each foot;

[0178] determining the contact of the feet of the object with the motion platform according to the contact probability of the foot key points of each foot with the motion platform.

[0179] In an embodiment, the processing unit 702, when correcting the action data based on the contact of the feet of the object with the motion platform, can be specifically configured to:

[0180] if the contact includes that the feet of the N feet of the object are in contact with the motion platform, then the posture data of the object is corrected.

[0181] In an embodiment, the number of the foot key points of each foot is multiple, and the processing unit 702, when correcting the posture data of the object, can be specifically configured to:

[0182] keeping the height of the lowest point of the multiple foot key points of each foot in the three-dimensional space unchanged, and adjusting the rotation angles of the knee joint and the thigh joint of the object.

[0183] In an embodiment, the processing unit 702, when correcting the action data based on the contact of the feet of the object with the motion platform, can be specifically configured to:

[0184] if the contact includes that the feet of M feet of the N feet of the object are in contact with the motion platform, but the feet of the remaining N-M feet are not in contact with the motion platform, then the position data of the object is corrected.

[0185] In an embodiment, the processing unit 702, when correcting the position data of the object, can be specifically configured to:

[0186] keeping the coordinates of the foot key points of the M feet in the three-dimensional space unchanged, and adjusting the coordinates of the center of gravity of the object in the three-dimensional space.

[0187] In an embodiment, the processing unit 702, when determining the two-dimensional coordinates of the foot key points of each foot according to the action data, can be specifically configured to:

[0188] obtain object skeleton data;

[0189] calculate coordinates of each rotation joint of the object in three-dimensional space according to the posture data and the object skeleton data;

[0190] project the coordinates of each rotation joint in three-dimensional space respectively according to the position data and the fixed focal length to obtain two-dimensional coordinates of each rotation joint;

[0191] determine the two-dimensional coordinates of the foot key points of each foot according to the two-dimensional coordinates of each rotation joint.

[0192] In an embodiment, the processing unit 702, when determining the two-dimensional coordinates of the foot key points of each foot according to the two-dimensional coordinates of each rotation joint, can be specifically configured to:

[0193] calculate a mark box of a target region in the object according to the posture data;

[0194] perform key point recognition processing on the target region in the mark box according to the two-dimensional coordinates of the rotation joints included in the target region to obtain the two-dimensional coordinates of the foot key points of each foot.

[0195] In an embodiment, the processing unit 702, when predicting the contact probability of the foot key points of each foot with the motion platform according to the two-dimensional coordinates of the foot key points of each foot, can be specifically configured to:

[0196] perform frame selection processing on N feet in the image respectively to obtain a frame selection image corresponding to each foot;

[0197] call the trained first neural network to predict the foot key points in the N frame selection images respectively according to the two-dimensional coordinates of the foot key points of each foot to obtain the contact probability of the foot key points of each foot with the motion platform.

[0198] In an embodiment, the action data of the object is obtained by calling the trained second neural network to perform recognition processing on the image; the processing unit 702 is further configured to:

[0199] call the trained second neural network to perform recognition processing on the image to obtain the heat Figure Two key points of the object;

[0200] determine whether the object is located in the image according to the heat Figure Two key points of the object;

[0201] If the object is located in the image, a step of determining, according to the action data, a contact condition of a foot of the object with the motion platform is performed.

[0202] In one embodiment, the acquisition unit 701 is further configured to acquire a first sample set, the first sample set including two-dimensional coordinates of foot key points of each sample foot of a sample object in a first sample image and a contact label of the foot key points of each sample foot with the motion platform.

[0203] The processing unit 702 is further configured to perform frame selection processing on each sample foot in the first sample image to obtain a frame-selected image corresponding to each sample foot, call the first neural network to respectively predict the foot key points in the frame-selected image corresponding to each sample foot according to the two-dimensional coordinates of the foot key points of each sample foot, and obtain a contact probability of the foot key points of each sample foot with the motion platform, and perform optimization processing on the first neural network according to the contact probability of the foot key points of each sample foot with the motion platform and the corresponding contact label, to obtain a trained first neural network.

[0204] In one embodiment, the acquisition unit 701 is further configured to acquire a second sample set, the second sample set including a second sample image and annotation information corresponding to the second sample image, the annotation information including three-dimensional annotation information and two-dimensional annotation information of a sample object in the second sample image.

[0205] The processing unit 702 is further configured to train the second neural network according to the second sample image and the corresponding annotation information, to obtain an intermediate neural network.

[0206] The acquisition unit 701 is further configured to acquire a third sample set, the third sample set including a third sample image, the third sample image being at least one of the following: a second sample image in the second sample set, a sample image obtained by cropping a randomly selected second sample image, and a sample image obtained by performing color space processing on a randomly selected second sample image.

[0207] The processing unit 702 is further configured to train the intermediate neural network according to the third sample image included in the third sample set and the corresponding annotation information, to obtain a trained second neural network.

[0208] In one embodiment, when the processing unit 702 trains the second neural network according to the second sample image and the corresponding annotation information to obtain the intermediate neural network, the processing unit 702 can be specifically configured to:

[0209] Call the second neural network to perform identification processing on the second sample image, to obtain action data of the sample object in the second sample image and two-dimensional key point coordinates of the sample object.

[0210] According to the action data and the three-dimensional annotation information of the sample object, model loss calculation is performed to obtain a first loss value.

[0211] According to the two-dimensional annotation information and the two-dimensional key point coordinates of the sample object, model loss calculation is performed to obtain a second loss value.

[0212] According to the first loss value and the second loss value, the network parameters of the second neural network are adjusted to obtain an intermediate neural network.

[0213] In an embodiment, the two-dimensional coordinates of the foot key points of each foot are obtained by calling the trained third neural network to perform key point recognition processing on the target region in the marking box, and the obtaining unit 701 is further configured to obtain a fourth sample set, the fourth sample set including a fourth sample image, a sample marking box of a target region of a sample object in the fourth sample image, and a two-dimensional coordinate label of a foot key point of a sample foot in the sample object.

[0214] The processing unit 702 is further configured to call the third neural network to perform key point recognition processing on the target region of the sample object in the sample marking box to obtain the two-dimensional coordinates of the foot key points of the sample foot, and adjust the network parameters of the third neural network according to the two-dimensional coordinates of the foot key points of the sample foot and the corresponding two-dimensional coordinate label to obtain the trained third neural network.

[0215] In an embodiment, when the obtaining unit 701 obtains the second sample set, it can be specifically configured to:

[0216] Obtain sample images and sample videos with three-dimensional annotation information, and use the obtained sample images with three-dimensional annotation information and the images obtained by processing the sample videos with three-dimensional annotation information as the second sample images; or

[0217] Obtain sample images with two-dimensional annotation information, and use a pre-trained three-dimensional construction model to perform fitting processing on the obtained sample images with two-dimensional annotation information to obtain three-dimensional annotation information corresponding to the sample images, and use the sample images corresponding to the obtained three-dimensional annotation information as the second sample images.

[0218] In the embodiments of the present application, the image to be processed can be recognized to obtain the action data of the object, and the contact between the foot of the object and the motion platform (such as the ground) is determined according to the action data of the object, and the action data of the object is corrected based on the contact between the foot of the object and the motion platform, and the virtual image matching the object is displayed according to the corrected action data. The drift problem of the virtual image can be solved through the correction processing, and the presentation of the virtual image is ensured to be real and stable.

[0219] Further, the embodiment of the present application further provides a structural schematic diagram of a computer device, which can be seen from Figure Eight The computer device can include a processor 801, an input device 802, an output device 803 and a memory 804. The processor 801, the input device 802, the output device 803 and the memory 804 are connected through a bus. The memory 804 is used for storing a computer program, and the computer program includes program instructions. The processor 801 is used for executing the program instructions stored in the memory 804.

[0220] In the embodiment of the present application, the processor 801 executes the following operations by running the executable program code in the memory 804:

[0221] An image to be processed is acquired, the image being obtained by photographing an object on a motion platform;

[0222] The image is recognized to obtain action data of the object;

[0223] According to the action data, a contact condition of a foot of the object and the motion platform is determined;

[0224] The action data is corrected based on the contact condition of the foot of the object and the motion platform;

[0225] A virtual image matched with the object is displayed according to the corrected action data, and a virtual foot of the virtual image is in contact with the motion platform.

[0226] In one embodiment, the action data includes at least one of the following: position data and posture data; the position data includes coordinates of a barycenter of the object in a three-dimensional space; and the posture data includes a rotation matrix formed by rotation angles of each rotation joint of the object.

[0227] The number of feet of the object is N; and the contact condition of the foot of the object and the motion platform includes any one of the following: the feet of the N feet of the object are all in contact with the motion platform; the feet of M feet of the N feet of the object are in contact with the motion platform, but the feet of the remaining N-M feet of the object are not in contact with the motion platform, M and N are positive integers and M is less than or equal to N.

[0228] In one embodiment, when the processor 801 determines the contact condition of the foot of the object and the motion platform according to the action data, the processor 801 can be specifically used for:

[0229] According to the action data, two-dimensional coordinates of a foot key point of each foot are determined;

[0230] According to the two-dimensional coordinates of the foot key point of each foot, a contact probability of the foot key point of each foot and the motion platform is predicted;

[0231] According to the contact probability of the foot key points of each foot and the motion platform, the contact of the feet of each foot of the object with the motion platform is determined.

[0232] In an embodiment, the processor 801 can be specifically configured to:

[0233] If the contact condition includes that the feet of N feet of the object are in contact with the motion platform, the posture data of the object is corrected.

[0234] In an embodiment, the number of foot key points of each foot is a plurality, and the processor 801 can be specifically configured to:

[0235] The height of the lowest point of the plurality of foot key points of each foot in the three-dimensional space is kept unchanged, and the rotation angles of the knee joint and the thigh joint of the object are adjusted.

[0236] In an embodiment, the processor 801 can be specifically configured to:

[0237] If the contact condition includes that the feet of M feet of N feet of the object are in contact with the motion platform, but the feet of the remaining N-M feet are not in contact with the motion platform, the position data of the object is corrected.

[0238] In an embodiment, the processor 801 can be specifically configured to:

[0239] The coordinates of the foot key points of the M feet in the three-dimensional space are kept unchanged, and the coordinates of the center of gravity of the object in the three-dimensional space are adjusted.

[0240] In an embodiment, the processor 801 can be specifically configured to:

[0241] Obtain object skeleton data;

[0242] According to the posture data and the object skeleton data, the coordinates of each rotation joint of the object in the three-dimensional space are calculated;

[0243] According to the position data and the fixed focal length, the coordinates of each rotation joint in the three-dimensional space are respectively projected to obtain the two-dimensional coordinates of each rotation joint;

[0244] According to the two-dimensional coordinates of each rotation joint, the two-dimensional coordinates of each foot key point are determined.

[0245] In an embodiment, the processor 801, when determining the two-dimensional coordinates of the foot key points of each foot according to the two-dimensional coordinates of the rotation joints of the respective rotation joint, can be specifically used for:

[0246] According to the posture data, the marking box of the target region in the object is calculated;

[0247] According to the two-dimensional coordinates of the rotation joints included in the target region, the key point recognition processing is performed on the target region in the marking box to obtain the two-dimensional coordinates of the foot key points of each foot.

[0248] In an embodiment, the processor 801, when predicting the contact probability of the foot key points of each foot with the motion platform according to the two-dimensional coordinates of the foot key points of each foot, can be specifically used for:

[0249] Frame selection processing is performed on N feet in the image respectively to obtain the frame selection image corresponding to each foot;

[0250] According to the two-dimensional coordinates of the foot key points of each foot, the trained first neural network is called to predict the foot key points in the N frame selection images respectively to obtain the contact probability of the foot key points of each foot with the motion platform.

[0251] In an embodiment, the action data of the object is obtained by calling the trained second neural network to perform recognition processing on the image; the processor 801 is further used for:

[0252] Calling the trained second neural network to perform recognition processing on the image to obtain the heat Figure Two key points of the object;

[0253] According to the heat Figure Two key points of the object, it is judged whether the object is located in the image;

[0254] If the object is located in the image, the step of determining the contact condition of the foot of the object with the motion platform according to the action data is executed.

[0255] In an embodiment, the processor 801 is further used for

[0256] Obtaining a first sample set, the first sample set including the two-dimensional coordinates of the foot key points of each sample foot of a sample object in a first sample image and the contact label of the foot key points of each sample foot with the motion platform;

[0257] Frame selection processing is performed on the sample feet in the first sample image respectively to obtain the frame selection image corresponding to each sample foot;

[0258] According to the two-dimensional coordinates of the foot key points of each sample foot, a first neural network is called to predict the foot key points in the framed image corresponding to each sample foot to obtain the contact probability of the foot key points of each sample foot with the motion platform;

[0259] According to the contact probability of the foot key points of each sample foot with the motion platform and the corresponding contact label, the first neural network is optimized to obtain the trained first neural network.

[0260] In an embodiment, the processor 801 is further configured to

[0261] A second sample set is obtained, the second sample set including second sample images and annotation information corresponding to the second sample images, the annotation information including three-dimensional annotation information and two-dimensional annotation information of sample objects in the second sample images;

[0262] The second neural network is trained according to the second sample images and the corresponding annotation information to obtain an intermediate neural network;

[0263] A third sample set is obtained, the third sample set including third sample images, the third sample images being at least one of the following: the second sample images in the second sample set, sample images obtained by cropping randomly selected second sample images, and sample images obtained by performing color space processing on randomly selected second sample images;

[0264] The intermediate neural network is trained according to the third sample images included in the third sample set and the corresponding annotation information to obtain the trained second neural network.

[0265] In an embodiment, when the processor 801 trains the second neural network according to the second sample images and the corresponding annotation information to obtain the intermediate neural network, the processor 801 can be specifically configured to:

[0266] The second neural network is called to perform identification processing on the second sample images to obtain action data of the sample objects in the second sample images and two-dimensional key point coordinates of the sample objects;

[0267] Model loss calculation is performed according to the action data of the sample objects and the three-dimensional annotation information to obtain a first loss value;

[0268] Model loss calculation is performed according to the two-dimensional annotation information of the sample objects and the two-dimensional key point coordinates to obtain a second loss value;

[0269] The network parameters of the second neural network are adjusted according to the first loss value and the second loss value to obtain the intermediate neural network.

[0270] In an embodiment, the two-dimensional coordinates of the foot key points of each foot are obtained by calling the trained third neural network to perform key point recognition processing on the target region in the labeled frame.

[0271] obtaining a fourth sample set, the fourth sample set including fourth sample images, sample labeled frames of target regions of sample objects in the fourth sample images, and two-dimensional coordinate labels of foot key points of sample feet in the sample objects;

[0272] calling the third neural network to perform key point recognition processing on the target region of the sample object in the sample labeled frame to obtain the two-dimensional coordinates of the foot key points of the sample feet;

[0273] adjusting network parameters of the third neural network according to the two-dimensional coordinates of the foot key points of the sample feet and the corresponding two-dimensional coordinate labels to obtain the trained third neural network.

[0274] In an embodiment, when obtaining the second sample set, the processor 801 can be specifically configured to:

[0275] obtaining sample images and sample videos with three-dimensional annotation information, and taking the obtained sample images with three-dimensional annotation information and images obtained by processing the sample videos with three-dimensional annotation information as the second sample images; or

[0276] obtaining sample images with two-dimensional annotation information, and using a pre-trained three-dimensional construction model to perform fitting processing on the obtained sample images with two-dimensional annotation information to obtain three-dimensional annotation information corresponding to the sample images, and taking the sample images corresponding to the obtained three-dimensional annotation information as the second sample images.

[0277] In the embodiments of the present application, the image to be processed can be identified to obtain action data of the object, and the contact between the foot of the object and the movement platform (such as the ground) can be determined according to the action data of the object, and the action data of the object can be corrected based on the contact between the foot of the object and the movement platform, and the virtual image matched with the object can be displayed according to the corrected action data. The drift problem of the virtual image can be solved through the correction processing, and the presentation of the virtual image can be ensured to be real and stable.

[0278] In addition, it should be pointed out here that the embodiments of the present application also provide a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program includes program instructions, and when the processor executes the above program instructions, the above Figure Two and Figure FourThe method in the corresponding embodiment will not be described here again. For technical details not disclosed in the computer-readable storage medium embodiments involved in the present application, please refer to the description of the method embodiments of the present application. As an example, the program instructions can be deployed on one computer device, or executed on multiple computer devices located in one place, or executed on multiple computer devices distributed in multiple places and interconnected through a communication network.

[0279] According to an aspect of the present application, a computer program product is provided, which includes a computer program stored in a computer readable storage medium. A processor of a computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program, so that the computer device can perform the method as described above. Figure Two and Figure Four The method in the corresponding embodiment will not be described here again.

[0280] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.

[0281] The above disclosure is only a preferred embodiment of the present application, and of course cannot limit the scope of the rights of the present application. A person of ordinary skill in the art can understand that all or part of the above-mentioned embodiments can be implemented, and equivalent changes made according to the claims of the present application still fall within the scope of the present application.

Claims

1. A data processing method, characterized by, The method comprises the following steps: acquiring an image to be processed, the image being obtained by photographing an object on a motion platform; performing identification processing on the image to obtain motion data of the object, the motion data comprising at least one of position data and posture data, the position data comprising coordinates of the center of gravity of the object in a three-dimensional space, and the posture data comprising a rotation matrix formed by rotation angles of each rotation joint of the object; the object has N feet, N being a positive integer; acquiring object skeleton data, and calculating coordinates of each rotation joint of the object in the three-dimensional space according to the posture data and the object skeleton data; projecting the coordinates of each rotation joint of the object in the three-dimensional space according to the position data and a fixed focal length to obtain two-dimensional coordinates of each rotation joint; determining two-dimensional coordinates of a foot key point of each foot according to the two-dimensional coordinates of each rotation joint, and predicting a contact probability of the foot key point of each foot with the motion platform according to the two-dimensional coordinates of the foot key point of each foot; determining a contact condition of the feet of the object with the motion platform according to the contact probability of the foot key point of each foot with the motion platform; correcting the motion data based on the contact condition of the feet of the object with the motion platform; displaying a virtual image matching the object according to the corrected motion data, the virtual image being in contact with the motion platform at a virtual foot portion.

2. The method of claim 1, wherein, The contact condition of the feet of the object with the motion platform comprises any one of the following: the feet of the N feet of the object are all in contact with the motion platform; the feet of M feet of the N feet of the object are in contact with the motion platform, but the feet of the remaining N-M feet are not in contact with the motion platform, M being a positive integer and M being less than or equal to N.

3. The method of claim 2, wherein, The correction of the motion data based on the contact condition of the feet of the object with the motion platform comprises: if the contact condition comprises that the feet of the N feet of the object are all in contact with the motion platform, the posture data of the object is corrected.

4. The method of claim 3, wherein, The number of foot key points of each foot is a plurality, and the correction of the posture data of the object comprises: keeping the height of the lowest point of the plurality of foot key points of each foot in the three-dimensional space unchanged, and adjusting the rotation angles of the knee joint and the thigh joint of the object.

5. The method of claim 1 or 2, wherein, The correction of the motion data based on the contact condition of the feet of the object with the motion platform comprises: if the contact condition comprises that the feet of M feet of the N feet of the object are in contact with the motion platform, but the feet of the remaining N-M feet are not in contact with the motion platform, the position data of the object is corrected.

6. The method of claim 5, wherein, The correction of the position data of the object comprises: keeping the coordinates of the foot key points of the M feet in the three-dimensional space unchanged, and adjusting the coordinates of the center of gravity of the object in the three-dimensional space.

7. The method of claim 1, wherein, The determination of the two-dimensional coordinates of the foot key point of each foot according to the two-dimensional coordinates of each rotation joint comprises: According to the posture data, a mark box of a target region in the object is calculated; According to the two-dimensional coordinates of the rotating joints included in the target region, key point recognition processing is performed on the target region in the mark box to obtain two-dimensional coordinates of foot key points of each foot.

8. The method of claim 1, wherein, The prediction of the contact probability of the foot key points of each foot with the motion platform according to the two-dimensional coordinates of the foot key points of each foot includes: Respectively, the N feet in the image are framed to obtain a corresponding framed image for each foot; According to the two-dimensional coordinates of the foot key points of each foot, a trained first neural network is called to predict the foot key points in the N framed images to obtain the contact probability of the foot key points of each foot with the motion platform.

9. The method of claim 1, wherein, The action data of the object is obtained by calling a trained second neural network to recognize the image; the method further includes: Calling the trained second neural network to recognize the image to obtain a heat map two-dimensional key point of the object; According to the heat map two-dimensional key point of the object, it is judged whether the object is located in the image or not; If the object is located in the image, the step of determining the contact between the feet of the object and the motion platform according to the action data is executed.

10. The method of claim 8, wherein, The method further includes: Obtaining a first sample set, the first sample set including two-dimensional coordinates of foot key points of each sample foot of a sample object in a first sample image and a contact label of the foot key points of each sample foot with a motion platform; Respectively, the sample feet in the first sample image are framed to obtain a corresponding framed image for each sample foot; According to the two-dimensional coordinates of the foot key points of each sample foot, a first neural network is called to predict the foot key points in the framed image corresponding to each sample foot to obtain the contact probability of the foot key points of each sample foot with the motion platform. According to the contact probability of the foot key points of each sample foot with the motion platform and the corresponding contact label, the first neural network is optimized to obtain a trained first neural network.

11. The method of claim 9, wherein, The method further includes: Obtaining a second sample set, the second sample set including a second sample image and corresponding annotation information of the second sample image, the annotation information including three-dimensional annotation information and two-dimensional annotation information of a sample object in the second sample image; According to the second sample image and the corresponding annotation information, a second neural network is trained to obtain an intermediate neural network; Obtaining a third sample set, the third sample set including a third sample image, the third sample image being at least one of the following: a second sample image in the second sample set, a sample image obtained by cropping a randomly selected second sample image, and a sample image obtained by color space processing on a randomly selected second sample image; According to the third sample image included in the third sample set and the corresponding annotation information, the intermediate neural network is trained to obtain a trained second neural network.

12. The method of claim 11, wherein, The second neural network is called to recognize and process the second sample image, to obtain action data of a sample object in the second sample image and two-dimensional key point coordinates of the sample object. According to the action data of the sample object and the three-dimensional label information, model loss calculation is performed to obtain a first loss value. According to the two-dimensional label information of the sample object and the two-dimensional key point coordinates, model loss calculation is performed to obtain a second loss value. According to the first loss value and the second loss value, the network parameters of the second neural network are adjusted to obtain an intermediate neural network. The two-dimensional coordinates of the foot key points of each foot are obtained by calling the trained third neural network to perform key point recognition processing on the target region in the marking box, and the method further comprises:

13. The method of claim 7, wherein, Obtain a fourth sample set, which includes a fourth sample image, a sample marking box of a target region of a sample object in the fourth sample image, and two-dimensional coordinate labels of foot key points of a sample foot in the sample object. The third neural network is called to perform key point recognition processing on the target region of the sample object in the sample marking box, to obtain the two-dimensional coordinates of the foot key points of the sample foot. According to the two-dimensional coordinates of the foot key points of the sample foot and the corresponding two-dimensional coordinate labels, the network parameters of the third neural network are adjusted to obtain the trained third neural network. The second sample set is obtained, including:

14. The method of claim 11, wherein, Obtain a sample image with three-dimensional label information and a sample video, and use the obtained sample image with three-dimensional label information and the image obtained by processing the sample video with three-dimensional label information as the second sample image; or Obtain a sample image with two-dimensional label information, and use a pre-trained three-dimensional construction model to perform fitting processing on the obtained sample image with two-dimensional label information to obtain the corresponding three-dimensional label information of the sample image, and use the sample image corresponding to the obtained three-dimensional label information as the second sample image. It includes:

15. A data processing apparatus, characterized by An acquisition unit is configured to acquire an image to be processed, the image being obtained by photographing an object on a motion platform; A processing unit is configured to perform recognition processing on the image to obtain action data of the object; the action data includes at least one of position data and posture data; the position data includes coordinates of a center of gravity of the object in a three-dimensional space; the posture data includes a rotation matrix formed by rotation angles of each rotation joint of the object; the number of feet of the object is N, and N is a positive integer; The processing unit is further configured to acquire object skeleton data of the object, and calculate coordinates of each rotation joint of the object in a three-dimensional space according to the posture data and the object skeleton data; and project the coordinates of each rotation joint in the three-dimensional space respectively according to the position data and a fixed focal length to obtain two-dimensional coordinates of each rotation joint. ​ According to the two-dimensional coordinates of the respective rotation joints, two-dimensional coordinates of foot key points of each foot are determined, and according to the two-dimensional coordinates of the foot key points of each foot, a contact probability of the foot key points of each foot with the motion platform is predicted, and according to the contact probability of the foot key points of each foot with the motion platform, a contact condition of a foot of each foot of the object with the motion platform is determined; The processing unit is further configured to correct the action data based on the contact condition of the foot of the object with the motion platform; The processing unit is further configured to display a virtual image matched with the object according to the corrected action data, and a virtual foot of the virtual image is in contact with the motion platform.

16. A computer device, comprising: Comprising: A processor adapted to execute a computer program; A computer readable storage medium having a computer program stored therein, the computer program being executed by the processor to perform the data processing method according to any one of claims 1-14.

17. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to perform the data processing method according to any one of claims 1-14.

18. A computer program product, characterised in that, The computer program product comprises a computer program, and the computer program is executed by the processor to implement the data processing method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Virtual human walking motion synthesis method

    CN106504307A