A method and device for processing motion data used in games
By using a neural network model in the game to process user video data, generate three-dimensional posture data and synchronize it to the virtual character, the problem of users being unable to make their own personalized movements is solved, and the display of unique movements and the improvement of user experience are achieved.
Patent Information
- Application Number
- CN201811614253.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-12-27
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2038-12-27
AI Technical Summary
In existing games, users cannot independently create personalized virtual character actions, resulting in insufficient user experience.
The server receives the user's video data, uses a neural network model to extract and convert the posture and action data in the video, generates three-dimensional posture data and synchronizes it to the virtual character in the game device, realizing the personalized display of the user's self-made actions.
Users can display their own unique virtual character actions in the game to meet personalized needs and enhance user experience.
Smart Images

Figure CN109529350B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of game personalization, and in particular to a method and device for processing action data used in games. Background Art
[0002] Currently, the actions of virtual characters in various online games or stand-alone games are all produced by game providers, that is, the actions of virtual characters in games that are professionally produced content (PGC).
[0003] More and more game users want unique avatar animations to meet their personalized needs. However, existing games that offer personalized avatar animations are considered PGC, and the number of personalized avatar animations provided by game providers is very limited. To meet users' personalized needs, the best solution is for users to create their own avatar animations, which is considered user-generated content (UGC). Obviously, most users do not have the ability to create avatar animations.
[0004] Therefore, it is particularly important to enable users to use their own virtual character actions in the game to meet the user's personalized needs for virtual character actions and improve user experience. Summary of the Invention
[0005] In order to solve the above problems, the present application proposes a motion data processing method and device for use in games. This method enables users to use their own virtual character actions, meet users' needs for personalized virtual character actions in games, and enhance user experience.
[0006] A first aspect of the present application provides a method for processing motion data applied to a game, the method comprising: a server receiving video data sent by a game device, the video data including gesture motion data of a motion of an object to be captured;
[0007] The server processes the video data to obtain a file in a first predetermined format, wherein the file in the first predetermined format stores three-dimensional posture data of the object to be captured performing a corresponding posture action in the video data;
[0008] The server sends the file in the first predetermined format.
[0009] It should be noted that the three-dimensional posture data in the first predetermined format file is continuous and consistent with the corresponding posture action of the object to be captured in the video data.
[0010] In one possible implementation, the server sends the file in the first predetermined format, specifically: the server sends the file in the first predetermined format to the game server corresponding to the game device, so that the game server synchronizes the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the game device.
[0011] In a possible implementation, the server sends the file in the first predetermined format, specifically by synchronizing the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the game device.
[0012] In a possible implementation, the server processes the video data by: decomposing the video data into at least one picture; wherein each picture corresponds to a frame of the video data;
[0013] Extracting, based on the first neural network model, two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture obtained by decomposing the video data;
[0014] Determining, based on the second neural network model, three-dimensional coordinate data corresponding to each marker point on the object to be captured contained in each image obtained by decomposing the video data, wherein the local coordinate system is a coordinate system determined by the center of mass of the object to be captured;
[0015] Based on the position data of each marking point on the object to be captured in the at least one picture, and according to the three-dimensional coordinate data corresponding to each marking point in the local coordinate system, the three-dimensional coordinate data corresponding to each marking point on the object to be captured in the preset three-dimensional space is determined.
[0016] In a possible implementation, extracting, based on the first neural network model, the two-dimensional coordinate data of at least one marker point on the object to be captured contained in each image obtained by decomposing the video data is specifically:
[0017] Inputting each image obtained by decomposing the video data into a first neural network model, and outputting at least one confidence map corresponding to each image, wherein the coordinates of the pixel with the maximum brightness in each confidence map correspond to the coordinates of a marker point on the object to be captured;
[0018] According to the at least one confidence map corresponding to each picture, two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture is determined.
[0019] In a possible implementation, before inputting each image obtained by decomposing the video data into the first neural network model and outputting at least one confidence map corresponding to each image, the method further includes:
[0020] Obtain a picture sample library; wherein the picture sample library contains multiple sample pictures;
[0021] Marking at least one marker point on the object to be captured contained in each sample image in the image sample library;
[0022] The first neural network model is obtained through machine learning training using multiple sample images in the image sample library and at least one marked point marked on the sample images.
[0023] In one possible implementation, before determining, based on the second neural network model, the three-dimensional coordinate data corresponding to each marker point on the subject to be captured contained in each image obtained by decomposing the video data, in a local coordinate system, the method further includes:
[0024] Obtain a picture sample library; wherein the picture sample library contains multiple sample pictures;
[0025] Marking at least one marker point on the object to be captured contained in each sample image in the image sample library;
[0026] Based on the two-dimensional coordinates of each marked point marked on each sample image, obtaining three-dimensional coordinate data of each marked point on the object to be captured in the local coordinate system;
[0027] The second neural network model is obtained through machine learning training using the two-dimensional coordinate data of at least one marked point on each sample image in the image sample library and the corresponding three-dimensional coordinate data in the local coordinate system.
[0028] A second aspect of the present application provides a motion data processing device for use in a game, the device comprising a receiving unit, a processing unit, and a sending unit; wherein the receiving unit is configured to receive video data sent by a gaming device, the video data comprising gesture motion data of a motion of an object to be captured;
[0029] The processing unit is configured to process the video data to obtain a file in a first predetermined format, wherein the file in the first predetermined format stores three-dimensional posture data of the object to be captured performing a corresponding posture action in the video data;
[0030] The sending unit is used to send the file in the first predetermined format.
[0031] In a possible implementation, the sending unit sends the file in the first predetermined format, specifically: the sending unit sends the file in the first predetermined format to the game server corresponding to the game device, so that the game server synchronizes the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the game device.
[0032] In a possible implementation, the sending unit sends the file in the first predetermined format, specifically: the sending unit synchronizes the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the game device.
[0033] In a possible implementation, the processing unit processes the video data, specifically: the processing unit decomposes the video data into at least one picture; wherein each picture corresponds to a frame of the video data;
[0034] Extracting, based on the first neural network model, two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture obtained by decomposing the video data;
[0035] Determining, based on the second neural network model, three-dimensional coordinate data corresponding to each marker point on the object to be captured contained in each image obtained by decomposing the video data, wherein the local coordinate system is a coordinate system determined by the center of mass of the object to be captured;
[0036] Based on the position data of each marking point on the object to be captured in the at least one picture, and according to the three-dimensional coordinate data corresponding to each marking point in the local coordinate system, the three-dimensional coordinate data corresponding to each marking point on the object to be captured in the preset three-dimensional space is determined.
[0037] In one possible implementation, the processing unit extracts, based on the first neural network model, two-dimensional coordinate data of at least one marker point on the object to be captured contained in each image obtained by decomposing the video data. Specifically, each image obtained by decomposing the video data is input into the first neural network model, and at least one confidence map corresponding to each image is output, wherein the coordinates of the pixel with the maximum brightness in each confidence map correspond to the coordinates of a marker point on the object to be captured.
[0038] According to the at least one confidence map corresponding to each picture, two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture is determined.
[0039] In a possible implementation, the processing unit is further configured to obtain a picture sample library; wherein the picture sample library contains a plurality of sample pictures;
[0040] The processing unit marks at least one marker point on the object to be captured contained in each sample image in the image sample library;
[0041] The first neural network model is obtained through machine learning training using multiple sample images in the image sample library and at least one marked point marked on the sample images.
[0042] In a possible implementation, the processing unit is further configured to obtain a sample image library; wherein the sample image library includes a plurality of sample images; and the processing unit labels at least one marker point on the object to be captured contained in each sample image in the sample image library;
[0043] Based on the two-dimensional coordinates of each marked point marked on each sample image, obtaining three-dimensional coordinate data of each marked point on the object to be captured in the local coordinate system;
[0044] The second neural network model is obtained through machine learning training using the two-dimensional coordinate data of at least one marked point on each sample image in the image sample library and the corresponding three-dimensional coordinate data in the local coordinate system.
[0045] The method of the present application enables users to use their own virtual character actions, meeting the user's demand for personalized virtual character actions in the game, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0047] Figure 1 A schematic diagram of a method for processing motion data applied to a game provided by an embodiment of the present application;
[0048] Figure 2 A flowchart of another method for processing action data in a game provided by an embodiment of the present application;
[0049] Figure 3 A schematic diagram of a video-based posture data capture system provided in an embodiment of the present application;
[0050] Figure 4 A schematic diagram of a video-based posture data capture system provided in an embodiment of the present application;
[0051] Figure 5a A schematic diagram of a video screenshot provided in an embodiment of the present application;
[0052] Figure 5b A schematic diagram of another video screenshot provided in an embodiment of the present application;
[0053] Figure 6a A schematic diagram of extracting human body joints from an image provided by an embodiment of the present application;
[0054] Figure 6b A schematic diagram of another method for extracting human body joints from an image provided in an embodiment of the present application;
[0055] Figure 7 A schematic diagram of the structure of an action data processing device for use in games provided in an embodiment of the present application. DETAILED DESCRIPTION
[0056] In order to more clearly illustrate the overall concept of the present application, a detailed description is given below in an illustrative manner in conjunction with the accompanying drawings.
[0057] In order to solve the problem that users' personalized needs for virtual character actions in existing games are not met, the method of the present application enables users to use their own virtual character actions, meet users' personalized needs for virtual character actions in games, and improve user experience.
[0058] The games referred to in the embodiments of the present invention include action role playing games (ARPGs), third-person shooting games (TPSs), first-person shooting games (FPSs), massively multiplayer online role playing games (MMORPGs), etc. MMORPGs are a type of online game in which players play a virtual character and control the character to perform corresponding game actions.
[0059] These games feature avatars. When a user plays a game on a device, the selected avatar is displayed, along with the corresponding avatar actions. The game server assigns different avatar actions to each avatar, with each avatar assigned one or more actions.
[0060] In this case, when multiple users play the same game and select the same virtual character, the virtual character will display one or more inherent virtual character actions. This cannot meet the user's personalized requirements, such as making the virtual character actions displayed by the virtual character they select unique and distinctive, with the user's personal characteristics.
[0061] It should be noted that the virtual character described in the embodiment of the present invention can be a game character with the same appearance as the human body, or a game character virtualized with the human body as the main body. In the embodiment of the present invention, the virtual character is described by taking the game character with the same appearance as the human body as an example.
[0062] In addition, the gaming devices in the embodiments of the present invention include mobile terminals, tablet computers, and game consoles, and devices capable of installing game software programs, as well as devices capable of playing online games via the Internet. In the embodiments of the present invention, the gaming devices are described using mobile terminals as an example.
[0063] In the embodiment of the present invention, Figure 1 As shown, in step S101, when a user plays game A through a game device and selects a virtual character in the game, the user uploads the video data produced by the user to the server of game A through the game device; in step S102, the server of game A sends the video data to the motion data processing server; in step S103, the motion data processing server processes the posture and motion data of the object to be captured in the video data to obtain a file in a predetermined format; in step S104, the motion data processing server sends the file in the predetermined format to the server of game A; in step S105, the server of game A sends the file in the predetermined format to the game device, and the virtual character on the game device can then display the posture and motion of the object to be captured.
[0064] It should be pointed out that the above process is a general case of this embodiment. For example, the game device can also directly send video data to the above-mentioned motion data processing server, and then the motion data processing server processes the video data and sends the obtained file in a predetermined format directly to the game device. In the above process, the server of game A acts as a data transfer station; those skilled in the art will know that the game device and the game server can also be the same device; the game server and the motion data processing server can also be the same device. For the convenience of description, the following describes the embodiment of the present invention using the information interaction process between the game device and the motion data processing server as an example; hereinafter, the motion data server is also referred to as the server.
[0065] At this time, the video data can be shot by the user himself or other types of video data. The embodiment of the present invention does not limit the type of video data and the format of the video data. In the following description, for the sake of convenience, the video data is described by taking the video shot by the user himself as an example.
[0066] Therefore, in the embodiment of the present invention, the user sees the personalized actions displayed by the virtual character on the game interface of the gaming device, which satisfies the user's demand for personalized actions of the virtual character in the game, thereby improving the user experience.
[0067] The following is a detailed description of the above-mentioned action data processing applied to the game.
[0068] like Figure 2 As shown, a method for processing motion data applied in a game includes steps S201-S203.
[0069] S201: A server receives video data sent by a game device, where the video data includes gesture data of a movement of an object to be captured.
[0070] The video data may be a video recorded by the user in real time containing the gesture and motion data of the object to be captured, or a video recorded or produced by the user in advance containing the gesture and motion data of the object to be captured. The object to be captured may be a person or object moving in the video.
[0071] At this time, the posture data of the object to be captured can be the posture data of any moving target, which can be the posture data of the human body, the posture data of a moving object, or the posture data of a robot.
[0072] S202: The server processes the video data to obtain a file in a first predetermined format, where the file in the first predetermined format stores three-dimensional posture data of the object to be captured performing corresponding posture actions in the video data.
[0073] In one possible embodiment, the server processes the video data, specifically by the following steps: decomposing the video data into at least one picture; wherein each picture corresponds to a frame of the video data; extracting, based on a first neural network model, two-dimensional coordinate data of at least one marker point on the object to be captured contained in each picture obtained by decomposing the video data; determining, based on a second neural network model, three-dimensional coordinate data corresponding to each marker point on the object to be captured contained in each picture obtained by decomposing the video data, wherein the local coordinate system is a coordinate system determined by the center of mass of the object to be captured; determining, based on the position data of each marker point on the object to be captured in the at least one picture, three-dimensional coordinate data corresponding to each marker point on the object to be captured in a preset three-dimensional space according to the three-dimensional coordinate data corresponding to each marker point in the local coordinate system.
[0074] In one example, based on the first neural network model, the two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture obtained by decomposing the video data is extracted. Specifically, each picture obtained by decomposing the video data is input into the first neural network model, and at least one confidence map corresponding to each picture is output, wherein the coordinates of the pixel with the largest brightness in each confidence map correspond to the coordinates of a marking point on the object to be captured; based on the at least one confidence map corresponding to each picture, the two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture is determined.
[0075] In one example, before inputting each picture obtained by decomposing the video data into the first neural network model and outputting at least one confidence map corresponding to each picture, the method also includes: obtaining a picture sample library; wherein the picture sample library contains multiple sample pictures; marking at least one marker point on the object to be captured contained in each sample picture in the picture sample library; using the multiple sample pictures in the picture sample library and the at least one marker point marked on the sample pictures, obtaining the first neural network model through machine learning training.
[0076] In one example, before determining the three-dimensional coordinate data corresponding to each marker point in the local coordinate system based on the two-dimensional coordinate data of each marker point on the object to be captured contained in each picture obtained by decomposing the video data based on the second neural network model, the method also includes: obtaining a picture sample library; wherein the picture sample library contains multiple sample pictures; marking at least one marker point on the object to be captured contained in each sample picture in the picture sample library; based on the two-dimensional coordinates of each marker point marked on each sample picture, obtaining the three-dimensional coordinate data of each marker point on the object to be captured in the local coordinate system; using the two-dimensional coordinate data of at least one marker point marked on each sample picture in the picture sample library and the corresponding three-dimensional coordinate data in the local coordinate system, the second neural network model is obtained through machine learning training.
[0077] S203: The server sends the file in the first predetermined format.
[0078] In one example, the server sends the file in the first predetermined format, specifically: the server sends the file in the first predetermined format to the game server corresponding to the game device, so that the game server synchronizes the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the game device.
[0079] In one example, the server sends the file in the first predetermined format, specifically: the server synchronizes the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the game device.
[0080] The above S201-S203 utilizes the powerful learning and deduction capabilities of the neural network to convert the human body movements in a video into 3D motion data through the first neural network model and the second neural network model, and finally converts the 3D motion data into a first predetermined format file.
[0081] In one example, the file in the first predetermined format can be an FBX file, which converts the generated human joint 3D data into an FBX file for use by game artists; the FBX file can also be further converted into a BIP file required by 3D action producers, so that 3D action producers can use the captured human posture data in scenarios such as 3D animation and game production.
[0082] Below is Figure 3 - Figure 6b Taking as an example, the video-based posture data capture in the above process is described in detail.
[0083] The various embodiments of this application are described using the capture of human body posture data as an example. It is easy to note that by marking changes in at least one point on the subject to be captured, the posture data of the subject to be captured during motion can be determined. In the case of a human body, the marked point can be a joint point of the human body.
[0084] In one example, an embodiment of the present application provides a video-based gesture data capture system, such as Figure 3 As shown, the system includes: a camera device 1 and an image processing device 2. The image processing device 2 collects video data of the movement of the object to be captured (for example, a human body) through the camera device 1 connected thereto, and the collected video data includes data on multiple posture movements of the human body during movement. After obtaining the video data of the human body movement, the image processing device 2 performs frame decomposition processing on the video data to obtain multiple pictures, each of which corresponds to a frame image of the video data. Since the posture movements of the human body during movement are constantly changing, the joints of the human body in each picture (including but not limited to the head joint, neck joint, left shoulder joint, right shoulder joint, left elbow joint, right elbow joint, left elbow joint, right elbow joint, left pelvic joint, right pelvic joint, left knee joint, right knee joint, left ankle joint, right ankle joint, etc.) are in different positions on the picture.
[0085] It should be noted that the above-mentioned image processing device 2 and camera device 1 can be two components of the same device (for example, collecting the user's posture data and performing image processing indoors through a smart mobile device such as a mobile phone), or they can be two independent devices. In the case where the image processing device 2 and camera device 1 are two independent devices, the communication method between the two can be wired communication (for example, collecting the performer's posture data indoors and performing image processing through an image processing device such as a computer) or wireless communication. In the case of wireless communication, it can be a communication method based on a local area network (for example, collecting the user's posture data indoors through a mobile phone, etc., and sending it to a computer for image processing), or it can be a communication method based on the Internet (for example, collecting the user's posture data through a mobile phone, and then sending it to an application server for image processing based on a client application).
[0086] For example, Figure 3 This is a schematic diagram of an optional video-based posture data capture system provided in an embodiment of the present application, such as Figure 3 As shown, the video data obtained by the server 5 can be the motion posture data of the object to be captured collected in real time by a camera device such as a mobile phone, or it can be the video data containing the motion posture data of the object to be captured uploaded by the user through a device such as a mobile phone 3 or a computer 4.
[0087] Regardless of which of the above methods is used, when the image processing device 2 obtains the video data of human body movement captured by the camera device 1 and decomposes the video data into multiple pictures, the image processing device 2 can extract each joint point of the object to be captured contained in each picture obtained by decomposing the video data based on the first neural network model (for example, Figure 5a or Figure 5b and based on the second neural network model, determining the three-dimensional coordinate data of each joint point on the object to be captured contained in each picture obtained by decomposing the video data, according to the two-dimensional coordinate data of the left wrist joint point of the human body shown in the figure; and based on the second neural network model, determining the three-dimensional coordinate data of each joint point in a preset three-dimensional space according to the two-dimensional coordinate data of each joint point on the object to be captured contained in each picture obtained by decomposing the video data; finally, generating a file in a first predetermined format according to the three-dimensional coordinate data of each joint point in the preset three-dimensional space on the object to be captured contained in each picture obtained by decomposing the video data, wherein the file in the first predetermined format stores the three-dimensional posture data of the object to be captured performing the corresponding posture action in the video data.
[0088] It is easy to notice that the above-mentioned first neural network model and second neural network model are models pre-trained using artificial intelligence algorithms, wherein the first neural network model can be used to estimate the joint positions of the human body in the picture to obtain 2D coordinate data; the second neural network model can be used to convert the 2D joint coordinate data into corresponding 3D coordinate data. The powerful learning and deduction capabilities of neural networks are utilized to convert the human body movements in the video into 3D motion data. The neural network trained using a large number of real videos and pictures can effectively identify human postures in various environments. The embodiment of the present invention also has good processing of occlusion and self-occlusion of the human body, and has no restrictions on the use environment and the camera position for shooting the video, so it can process various video materials.
[0089] In one example, the first neural network model may be a convolutional neural network model based on a residual network, and the second neural network model may be a deep neural network model.
[0090] It is easy to notice that before using the first neural network model to extract the two-dimensional coordinate data of at least one joint point on the body of the object to be captured contained in each picture, the first neural network model needs to be trained first. Specifically, a picture sample library containing multiple sample pictures is collected, and at least one joint point on the body of the object to be captured contained in each sample picture in the picture sample library is marked (for example, head joint, neck joint, left shoulder joint, right shoulder joint, left elbow joint, right elbow joint, left wrist joint, right wrist joint, left pelvic joint, right pelvic joint, left knee joint, right knee joint, left ankle joint, right ankle joint), and the first neural network model is obtained through machine learning training using multiple sample pictures in the picture sample library and at least one joint point marked on the sample pictures.
[0091] After training the first neural network model, the first neural network model can be used to extract the two-dimensional coordinate data of at least one joint point on the subject to be captured contained in each image. The image processing device 2 inputs each image obtained by decomposing the video data into the first neural network model to output at least one confidence map corresponding to each image. Based on the at least one confidence map corresponding to each image, the two-dimensional coordinate data of at least one joint point on the subject to be captured contained in each image is determined, wherein the coordinates of the pixel with the highest brightness in each confidence map correspond to the coordinates of a joint point on the subject to be captured.
[0092] In one example, the first neural network model can be trained based on the following objective function:
[0093] Among them, E represents the objective function, H' j (x, y) is the coordinate of each joint point on the predicted sample image, H j (x, y) is the coordinate of the joint point annotated in the sample image, N represents the number of training samples, and j is a natural number.
[0094] In one example, when determining the two-dimensional coordinate data of at least one joint point on the subject to be captured contained in each image based on at least one confidence map corresponding to each image, the two-dimensional coordinates of the joint point corresponding to each confidence map can be determined using the following Gaussian response map formula:
[0095] Among them, G(x,y) represents the Gaussian distribution of pixels on each confidence map; σ represents the standard deviation of the Gaussian distribution, and (x,y) represents the coordinates of each pixel on each confidence map.
[0096] In addition, it should be noted that before using the second neural network model to determine the three-dimensional coordinate data of each joint point on the subject to be captured contained in each image obtained by decomposing the video data, the second neural network model needs to be trained first. The training process is as follows: collecting a picture sample library containing multiple sample images, annotating at least one joint point on the subject to be captured contained in each sample image in the picture sample library (for example: head joint, neck joint, left shoulder joint, right shoulder joint, left elbow joint, right elbow joint, left wrist joint, right wrist joint, left pelvic joint, right pelvic joint, left knee joint, right knee joint, left ankle joint, right ankle joint); and based on the two-dimensional coordinates of each joint point annotated on each sample image, obtaining the three-dimensional coordinate data of each joint point on the subject to be captured in a local coordinate system, wherein the local coordinate system is a coordinate system determined by the center of mass of the subject to be captured; using the two-dimensional coordinate data of at least one joint point annotated on each sample image in the picture sample library and the corresponding three-dimensional coordinate data in the local coordinate system, the second neural network model is obtained through machine learning training.
[0097] In one example, the second neural network model can be trained based on the following objective function:
[0098] Among them, E represents the objective function, H' j (x, y, z) is the three-dimensional coordinate of each joint point on the predicted sample image, H j (x, y, z) is the three-dimensional coordinate of the joint point annotated in the sample image, N represents the number of training samples, and j is a natural number.
[0099] Below, the video-based posture data capture solution provided by the present application is specifically described in conjunction with Figures 5(a) and 5(b). Figure 5(a) shows a single image decomposed from a high jump video of a high jumper. A convolutional neural network (CNN) can be used to infer the image shown in Figure 5(a) to obtain the joint positions of the human body in the image. As shown in Figure 5(b), the head joint, neck joint, left shoulder joint, right shoulder joint, left elbow joint, right elbow joint, left wrist joint, right wrist joint, left pelvic joint, right pelvic joint, left knee joint, right knee joint, left ankle joint, and right ankle joint.
[0100] The 2D human joint data corresponding to Figure 5(b) is shown in Figure 6(a). Using the deep neural network (DNN) to infer the 2D human joint data, the 3D coordinate data of the human joints in the preset three-dimensional space can be obtained, as shown in Figure 6(b).
[0101] In order to train the above convolutional neural network, we can collect pictures of human bodies in various environments as samples and mark the human joints in the pictures. The marked information includes: head joint, neck joint, left shoulder joint, right shoulder joint, left elbow joint, right elbow joint, left wrist joint, right wrist joint, left pelvic joint, right pelvic joint, left knee joint, right knee joint, left ankle joint, right ankle joint, a total of 14 joints, and refer to the convolutional neural network built with reference to the standard residual network. The convolutional network inputs the picture sample and outputs the confidence map of the human joints in the sample. The confidence map Figure 1 There are 14 images in total, each of which contains a Gaussian response map of the corresponding position of a joint point in the sample image.
[0102] It's important to note that the confidence map is a 2D image, the same size as the input image or proportionally smaller. Each pixel in the confidence map is compared to the input image, indicating the probability that the pixel's location in the input image is a human joint.
[0103] In one example, 100,000 images can be used as samples and trained 200 times. The coordinates of the points with the maximum brightness in the confidence map are extracted, that is, the 2D positions of the human joints are obtained as training input for the deep neural network.
[0104] To train the neural network described above, traditional optical equipment can be used in a studio to capture human poses using professional actors. This data can be used to obtain the position of human joints in 3D space. This data is continuous motion data. The number and position of joints are consistent with 2D data. In one example, 10 cameras can be used to surround the human body, with the lenses pointed at the actor in the center.
[0105] It should be noted that the input of the deep neural network is 2D joint data (X, Y), and the output is based on the three-dimensional joint position (X, Y, Z) of the corresponding human body. Since the deep neural network has a powerful feature extraction capability, the input 2D data is first expanded and then input into 3D coordinates after passing through the hidden layer. The L2Loss function is still used for training, and the Relu function is used as the activation function. Because the normal human body is bilaterally symmetrical, a constraint on the length of the bones is also added to make the lengths of the left and right bones as consistent as possible. The 3D posture of the human body can be obtained using this deep network, but there is no displacement data of the human body in 3D space, and the continuity of the movement is insufficient, and there is jitter in the playback movement.
[0106] In one example, when determining the three-dimensional coordinate data corresponding to each joint of the object to be captured in a preset three-dimensional space based on the position data of each joint of the object to be captured in at least one image and the three-dimensional coordinate data corresponding to each joint in the local coordinate system, the three-dimensional coordinate data corresponding to each joint of the object to be captured in the preset three-dimensional space can be determined based on the position data of each joint of the object to be captured in at least one image and the three-dimensional coordinate data corresponding to each joint in the local coordinate system:
[0107]
[0108] Among them, z is the approximate depth of the object to be captured in the preset three-dimensional space, is the joint information output by the second neural network model; is the average value of all joint information output by the second neural network model; K i is the joint information output by the first neural network model; is the average value of all joint information output by the first neural network model.
[0109] In one example, in the post-processing stage, an iterative method is used to process a 3D human body to make its movements smoother, eliminate the shaking of the ankles, and keep them fixed to the ground.
[0110] In one example, the generated human joint 3D data can also be converted into an FBX file (i.e., a file in FilmBox software format) commonly used by game artists. Since the video is a continuous human movement, the information saved in FBX includes frame-by-frame movement information.
[0111] In one example, the FBX file may also be converted into a human motion file BIP file (BIP stands for Bipedal, and the BIP file is a file in a format unique to 3dsmax cs, used for making animations and 3D).
[0112] The FBX file format is a free, cross-platform 3D creation and exchange format developed by Autodesk. Through FBX, users can access 3D files from most 3D vendors. The FBX file format supports all major 3D data elements as well as 2D, audio, and video media elements.
[0113] BIP files are commonly used for footstep controllers and are commonly used in animation and 3D production. BIP is a format unique to 3dsMaxCS and can be opened with Natural Motion Endorphin (a motion capture simulator) or other software such as MotionBuilder (a 3D character animation software). BIP files are commonly used by game artists for character animation.
[0114] Based on the appearance of standard bones in 3D Max, the rotation direction and length of human joints are calculated. Importing this data into 3D Max can obtain the correct BIP bones and animation effects.
[0115] In one example, the development process of this application is completed using TensorFlow, and the operating environment is a PC with a Linux operating system installed.
[0116] The embodiment of the present application also provides an action data processing device for use in games, such as Figure 7 As shown, it includes a receiving unit, a processing unit and a sending unit.
[0117] The receiving unit is used to receive video data sent by the game device, wherein the video data includes gesture data of the object to be captured.
[0118] a processing unit, configured to process the video data to obtain a file in a first predetermined format, wherein the file in the first predetermined format stores three-dimensional posture data of the object to be captured performing a corresponding posture action in the video data;
[0119] A sending unit is used to send the file in the first predetermined format.
[0120] In one example, the sending unit sends the file in the first predetermined format, specifically: the sending unit sends the file in the first predetermined format to the game server corresponding to the game device, so that the game server synchronizes the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the game device.
[0121] In one example, the sending unit sends the file in the first predetermined format, specifically: the sending unit synchronizes the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the game device.
[0122] In one example, the processing unit processes the video data, specifically as follows: the processing unit decomposes the video data into at least one picture; wherein each picture corresponds to a frame of the video data; based on a first neural network model, extracts two-dimensional coordinate data of at least one marker point on the object to be captured contained in each picture obtained by decomposing the video data; based on a second neural network model, according to the two-dimensional coordinate data of each marker point on the object to be captured contained in each picture obtained by decomposing the video data, determines the three-dimensional coordinate data corresponding to each marker point in a local coordinate system, wherein the local coordinate system is a coordinate system determined by the center of mass of the object to be captured; based on the position data of each marker point on the object to be captured in the at least one picture, according to the three-dimensional coordinate data corresponding to each marker point in the local coordinate system, determines the three-dimensional coordinate data corresponding to each marker point on the object to be captured in a preset three-dimensional space.
[0123] In one example, the processing unit determines, based on at least one confidence map corresponding to each image, two-dimensional coordinate data of at least one joint point on the object to be captured contained in each image, including:
[0124] The two-dimensional coordinates of the joint points corresponding to each confidence map are determined using the following Gaussian response map formula:
[0125]
[0126] Among them, G(x,y) represents the Gaussian distribution of pixels on each confidence map; σ represents the standard deviation of the Gaussian distribution, and (x,y) represents the coordinates of each pixel on each confidence map.
[0127] In one example, the processing unit uses multiple sample images in the image sample library and at least one joint point marked on the sample images to obtain the first neural network model through machine learning training, including:
[0128] The first neural network model is trained based on the following objective function:
[0129]
[0130] Among them, E represents the objective function, H' j (x, y) is the coordinate of each joint point on the predicted sample image, H j (x, y) is the coordinate of the joint point annotated in the sample image, N represents the number of training samples, and j is a natural number.
[0131] In one example, the processing unit extracts the two-dimensional coordinate data of at least one marker point on the object to be captured contained in each picture obtained by decomposing the video data based on the first neural network model. Specifically, each picture obtained by decomposing the video data is input into the first neural network model, and at least one confidence map corresponding to each picture is output, wherein the coordinates of the pixel with the largest brightness in each confidence map correspond to the coordinates of a marker point on the object to be captured; and according to the at least one confidence map corresponding to each picture, the two-dimensional coordinate data of at least one marker point on the object to be captured contained in each picture is determined.
[0132] In one example, the processing unit is further used to obtain a picture sample library; wherein the picture sample library contains multiple sample pictures; the processing unit marks at least one marker point on the object to be captured contained in each sample picture in the picture sample library; and using the multiple sample pictures in the picture sample library and at least one marker point marked on the sample pictures, the first neural network model is obtained through machine learning training.
[0133] In one example, the processing unit is also used to obtain a picture sample library; wherein the picture sample library contains multiple sample pictures; the processing unit marks at least one marker point on the object to be captured contained in each sample picture in the picture sample library; based on the two-dimensional coordinates of each marker point marked on each sample picture, obtain the three-dimensional coordinate data of each marker point on the object to be captured in the local coordinate system; use the two-dimensional coordinate data of at least one marker point marked on each sample picture in the picture sample library and the corresponding three-dimensional coordinate data in the local coordinate system to obtain the second neural network model through machine learning training.
[0134] In one example, the processing unit obtains the three-dimensional coordinates corresponding to the two-dimensional coordinates of each joint point marked on each sample image in the preset three-dimensional space: the three-dimensional coordinates corresponding to the two-dimensional coordinates of each joint point in the preset three-dimensional space are obtained by the following formula:
[0135]
[0136] Among them, z is the approximate depth of the object to be captured in the preset space, is the joint information output by the second neural network model; is the average value of all joint information output by the second neural network model; K i is the joint information output by the first neural network model; is the average value of all joint information output by the first neural network model.
[0137] In one example, the processing unit uses the two-dimensional coordinates and corresponding three-dimensional coordinates of at least one joint point marked on each sample image in the image sample library to obtain the second neural network model through machine learning training, including:
[0138] The second neural network model is trained based on the following objective function:
[0139]
[0140] Among them, E represents the objective function, H' j (x, y, z) is the three-dimensional coordinate of each joint point on the predicted sample image, H j (x, y, z) is the actual three-dimensional coordinate of each joint point on the collected sample image, N represents the number of training samples, and j is a natural number.
[0141] At this time, the file in the first predetermined format may also be converted into a file in a second predetermined format, wherein the file in the second predetermined format is used for producing a three-dimensional animation.
[0142] Specifically, the file in the second predetermined format may be a BIP file, and the file in the first predetermined format may be an FBX file; converting the FBX file into a BIP file can be used for animation and 3D production.
[0143] For the convenience of description, where the parts related to the above device embodiments are not explained, please refer to the description of the parts related to the method embodiments, and no further details will be given here.
[0144] In an embodiment of the present application, the powerful learning and inference capabilities of neural networks are utilized to convert human body movements in a video into 3D motion data through a first neural network model and a second neural network model, and finally the 3D motion data is converted into an FBX file. Optionally, the FBX file can be further converted into a BIP file required by 3D motion producers, so that 3D motion producers can use the captured human body posture data in scenarios such as 3D animation and game production.
[0145] The method of the present application enables users to use their own virtual character actions, meeting users' needs for personalized virtual character actions in games and improving user experience.
[0146] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0147] Professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0148] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for processing motion data used in games, characterized in that: The method comprises: The gaming device transmits user-generated video data, based on the virtual character selected by the user while playing the game, to a server corresponding to the gaming device. The video data includes gesture and motion data of the object to be captured. The gaming device includes a mobile phone terminal, a tablet computer, a game console, and a device that can connect to the internet and play online games. The games specifically include games with virtual characters, including at least any one of the following: action role-playing games, third-person shooter games, first-person shooter games, and massively multiplayer online role-playing games; After receiving the video data sent by the game device, the server corresponding to the game device sends the video data to the action data processing server; the server corresponding to the game device acts as a data transfer station; The motion data processing server processes the video data to obtain a file in a first predetermined format, wherein the file in the first predetermined format stores three-dimensional posture data of the object to be captured performing a corresponding posture action in the video data; The action data processing server sends the file in the first predetermined format to the game server corresponding to the game device; The server corresponding to the gaming device synchronizes the three-dimensional posture data in the file in the first predetermined format to the corresponding virtual game character in the gaming device; Wherein, the virtual game character is a game character corresponding to the appearance of a human body, or a game character virtualized based on the appearance of a human body; The virtual game character on the game device displays the gesture actions corresponding to the three-dimensional gesture data, so that the user can see the virtual character actions made by the user displayed by the virtual game character on the game interface of the game device.
2. The method according to claim 1, characterized in that The server processes the video data, specifically: Decomposing the video data into at least one picture; wherein each picture corresponds to a frame of the video data; Extracting, based on the first neural network model, two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture obtained by decomposing the video data; Determining, based on the second neural network model, three-dimensional coordinate data corresponding to each marker point on the object to be captured contained in each image obtained by decomposing the video data, wherein the local coordinate system is a coordinate system determined by the center of mass of the object to be captured; Based on the position data of each marking point on the object to be captured in the at least one picture, and according to the three-dimensional coordinate data corresponding to each marking point in the local coordinate system, the three-dimensional coordinate data corresponding to each marking point on the object to be captured in the preset three-dimensional space is determined.
3. The method according to claim 2, characterized in that The extracting, based on the first neural network model, the two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture obtained by decomposing the video data is specifically: Inputting each image obtained by decomposing the video data into a first neural network model, and outputting at least one confidence map corresponding to each image, wherein the coordinates of the pixel with the maximum brightness in each confidence map correspond to the coordinates of a marker point on the object to be captured; According to the at least one confidence map corresponding to each picture, two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture is determined.
4. The method according to claim 3, characterized in that Before inputting each picture obtained by decomposing the video data into the first neural network model and outputting at least one confidence map corresponding to each picture, the method further includes: Obtain a picture sample library; wherein the picture sample library contains multiple sample pictures; Marking at least one marker point on the object to be captured contained in each sample image in the image sample library; The first neural network model is obtained through machine learning training using multiple sample images in the image sample library and at least one marked point marked on the sample images.
5. The method according to claim 2, characterized in that Prior to determining, based on the second neural network model, the three-dimensional coordinate data corresponding to each marker point on the object to be captured contained in each image obtained by decomposing the video data, the three-dimensional coordinate data of each marker point in the local coordinate system; the method further includes: Obtain a picture sample library; wherein the picture sample library contains multiple sample pictures; Marking at least one marker point on the object to be captured contained in each sample image in the image sample library; Based on the two-dimensional coordinates of each marked point marked on each sample image, obtaining three-dimensional coordinate data of each marked point on the object to be captured in the local coordinate system; The second neural network model is obtained through machine learning training using the two-dimensional coordinate data of at least one marked point on each sample image in the image sample library and the corresponding three-dimensional coordinate data in the local coordinate system.
6. A motion data processing device used in a game, applying the method according to any one of claims 1 to 5, characterized in that: The device includes a receiving unit, a processing unit, a display unit and a sending unit; wherein, The receiving unit is configured to receive video data transmitted by a gaming device, the video data including gesture data of an object to be captured; wherein the gaming device includes a mobile phone terminal, a tablet computer, a gaming console, and a device capable of playing online games via an internet connection; the game specifically includes a game with a virtual character, including at least any one of the following: an action role-playing game, a third-person shooter game, a first-person shooter game, and a massively multiplayer online role-playing game; The processing unit is configured to process the video data to obtain a file in a first predetermined format, wherein the file in the first predetermined format stores three-dimensional posture data of the object to be captured performing a corresponding posture action in the video data; The sending unit is configured to send the file in the first predetermined format to a game server corresponding to the game device; the server corresponding to the game device synchronizes the three-dimensional posture data in the file in the first predetermined format to a corresponding virtual game character in the game device; wherein the virtual game character is a game character corresponding to the human body shape, or a game character virtualized based on the human body shape; the sending unit is specifically a motion data processing server, and the server corresponding to the game device acts as a data transfer station; The display unit is used to display the gesture actions corresponding to the three-dimensional gesture data on the virtual game character on the game device, so that the user can see the virtual character actions made by the user displayed by the virtual game character on the game interface of the game device.
7. The device according to claim 6, characterized in that The processing unit processes the video data, specifically: The processing unit decomposes the video data into at least one picture; wherein each picture corresponds to a frame of the video data; Extracting, based on the first neural network model, two-dimensional coordinate data of at least one marking point on the object to be captured contained in each picture obtained by decomposing the video data; Determining, based on the second neural network model, three-dimensional coordinate data corresponding to each marker point on the object to be captured contained in each image obtained by decomposing the video data, wherein the local coordinate system is a coordinate system determined by the center of mass of the object to be captured; Based on the position data of each marking point on the object to be captured in the at least one picture, and according to the three-dimensional coordinate data corresponding to each marking point in the local coordinate system, the three-dimensional coordinate data corresponding to each marking point on the object to be captured in the preset three-dimensional space is determined.
8. The device according to claim 7, characterized in that The processing unit is further configured to obtain a sample image library; wherein the sample image library contains a plurality of sample images; The processing unit marks at least one marker point on the object to be captured contained in each sample image in the image sample library; Based on the two-dimensional coordinates of each marked point marked on each sample image, obtaining three-dimensional coordinate data of each marked point on the object to be captured in the local coordinate system; The second neural network model is obtained through machine learning training using the two-dimensional coordinate data of at least one marked point on each sample image in the image sample library and the corresponding three-dimensional coordinate data in the local coordinate system.
Citation Information
Patent Citations
Generating three dimensional models using single two dimensional images
US20180211438A1
Video-based attitude data capture method and system
CN109145788A