Human body 3D image generation method and device, electronic equipment and vehicle
By extracting node features from video frames and establishing a vector space, an image generator is used to generate 3D images that conform to the laws of human kinematics. This solves the problem of discontinuity caused by 2D image generation errors in existing technologies and achieves accurate prediction results.
Patent Information
- Application Number
- CN202410509781.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-10-28
AI Technical Summary
In the existing technology, errors are inevitable in the process of extracting human body parameters through a single 2D image, resulting in errors in the prediction results. The generated 3D human body movement is incoherent and does not conform to the laws of human kinematics, posing a safety hazard.
Node features are extracted from video frames, a vector space is established, target vector data after a preset time is calculated, and a 3D image of the human body is generated using an image generator. The image generator is generated through adversarial training between an image initial generator and an image discriminator to ensure that the generated 3D image conforms to the laws of human kinematics.
The generated 3D image is consistent with the human body movement in the video frame, conforms to the laws of human kinematics, and can perform accurate predictions in areas such as assisted driving.
Smart Images

Figure CN120852631A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition, and in particular to a method, apparatus, electronic device, and vehicle for generating 3D human body images. Background Art
[0002] In assisted driving, human movements can be predicted to a certain extent based on changes in the human body in existing videos, and corresponding processing methods can be implemented based on the predicted human movements. Current technology typically breaks down the video into multiple 2D images when predicting human movements, and predicts the 3D human movements in future frames based on the human motion parameters of each 2D image. However, errors inevitably occur in extracting human parameters from a single 2D image, leading to inaccurate prediction results. Furthermore, generating human movements solely from 2D images may result in disjointed 3D human motion that does not conform to human kinematics, leading to misjudgments of human behavior in assisted driving and posing certain safety hazards. Summary of the Invention
[0003] In view of this, this application provides a method, apparatus, electronic device and vehicle for generating human 3D images. The main purpose is to solve the problem in the prior art that the specific human movements in the predicted 2D images cannot be used as accurate reference data to participate in the subsequent technical methods for generating human movements.
[0004] To achieve the above objectives, the first aspect of this application discloses a method for generating a 3D human body image, the method comprising:
[0005] Extract node features from video frames in the video to be processed. The node features are used to represent the joint features of the human body in the video frames.
[0006] Establish the vector space of the node features according to the temporal arrangement of video frames;
[0007] Based on the vector space, calculate the target vector data after a preset time.
[0008] A control image generator is used to generate a 3D human body image using the target vector data. The image generator is generated by adversarial training of the image initial generator and the image discriminator.
[0009] A second aspect of this application provides a human body 3D image generation apparatus, the apparatus comprising:
[0010] An extraction module is used to extract node features from video frames in the video to be processed, wherein the node features are used to represent the joint features of the human body in the video frame;
[0011] A module is established to create a vector space for the node features according to the temporal arrangement of video frames;
[0012] The calculation module is used to calculate the target vector data after a preset time based on the vector space;
[0013] A generation module is used to control an image generator to generate a 3D human body image using the target vector data. The image generator is generated by adversarial training of an image initial generator and an image discriminator.
[0014] A third aspect of this application provides an electronic device, comprising:
[0015] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform any of the methods disclosed in the first aspect.
[0016] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0017] A fifth aspect of this application provides a vehicle in which the device as described in the second aspect or the electronic device as described in the third aspect is mounted.
[0018] In summary, according to the technical solution disclosed in this application, this application first extracts node features from the video frames in the video to be processed. These node features represent the joint features of the human body in the video frames. Then, a vector space of node features is established according to the temporal arrangement of the video frames. Next, target vector data is calculated based on the vector space after a preset time. Finally, an image generator is controlled to generate a 3D human image using the target vector data. The image generator is generated by adversarially training an image initializer and an image discriminator. This application can collect node features of a person in the video to be processed and establish a vector space containing changes in node features according to the temporal arrangement of the video frames. Based on the changes in features in the vector space, the vector data corresponding to the node features after a preset time can be predicted. Through the correspondence between node features and the human body, a 3D human image can be predicted and generated. The technical solution of this application can determine the human body situation after a preset time by utilizing the changes in node features of the human image in the video frame, and simultaneously represent the human body in the form of a 3D image. This allows the human body image in the 2D image to be generated into a 3D human body image through prediction. In addition, during the process of generating the 3D human body image, this application further uses an image generator to generate the 3D human body image. The generator is generated through adversarial training between an image initial generator and an image discriminator. The 3D human body image generated by the image generator can be coherent with the 3D human body movement in the video frame and conforms to the laws of human kinematics. In some fields where human body movement or position is highly applied, advance prediction can be performed.
[0019] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating the human body 3D image generation process provided in an embodiment of this application is shown.
[0023] Figure 2 The diagram shows a structural diagram of a human body 3D image generation device provided in an embodiment of this application. Detailed Implementation
[0024] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0025] To address the problem in existing technologies where predicted human movements in generated 2D images cannot serve as accurate reference data for subsequent methods of generating human movements, this application provides the following embodiments to solve the above problem:
[0026] This embodiment provides a method for generating 3D human body images, such as... Figure 1 The diagram shown is a flowchart of the method in this embodiment. The method in this embodiment may specifically include the following steps:
[0027] Step 101: Extract node features from the video frames in the video to be processed. The node features are used to represent the joint features of the human body in the video frames.
[0028] The implementation of this embodiment mainly involves predicting the corresponding actions of a person appearing in a video after a preset time. In this embodiment, each video frame in the video to be processed contains a person. Before predicting the person, it is necessary to capture as many specific actions of the person as possible as the video plays. For example, when a pedestrian is running in the video, by observing the regular swinging of the pedestrian's arms and the regular forward movement of the legs, it can be predicted that the pedestrian's running action 5 seconds later will be: the left foot lands and the right foot lifts up, the left arm swings backward and the right arm lifts up.
[0029] To improve recognition efficiency when identifying specific movements of a person in a video, the movement of certain points on the human body can be identified. These selected points are used as nodes in the human movement process. These nodes can be rotational joints, such as the ankle, knee, shoulder, or elbow. These joints participate in various human movements and have strong correlations when identifying different movements. Furthermore, the type of human movement can be quickly identified by observing the activity of parts connected to these joints. In this embodiment, relevant data on human joint movements in the video to be processed can be used as node features, such as the position and direction of movement of joints in different video frames. Additionally, in this embodiment, node feature extraction involves extracting nodes from the human figure within each video frame of the video to be processed.
[0030] In one feasible embodiment, extracting node features from video frames in the video to be processed includes:
[0031] Identify human image data in video frames of the video to be processed; aggregate the human image data to determine the human image in the video frame; extract node features from the human image data based on the human body movements represented by the human image.
[0032] In this embodiment, further extraction of node features is performed. Node features are the extraction of the movement process of specific points on the human body. In other words, node feature extraction is predicated on the recognition of a human image within the video frame. Therefore, before performing node feature extraction, it is necessary to first identify the specific location of the human image in the video frame. Identifying the human image in the video frame involves determining the location where human image data exists. This can be understood as representing identifiable visual information within the video frame, such as color information (skin or clothing color), contour information (human outline), or texture information. Specifically, the process of recognizing human image data can be performed using Convolutional Neural Networks (CNNs) to extract human image features from the video frame.
[0033] After identifying the aforementioned human image data in the video, the data is synthesized to obtain overall data for one or more human bodies. This data is then combined with all the human body data to determine the human image within the video frames. Based on this determination, the positions of nodes within the human image can be further analyzed to further determine the corresponding human actions based on those node features.
[0034] This embodiment further explains the specific process of node feature extraction. After determining the corresponding human image data in the image, node features are further extracted from the accurate human image determined by the human image data. By adopting the technical solution of this embodiment, node feature extraction is further performed using the accurate human image determined in the 2D image, ensuring the reliability of the extracted nodes, thereby ensuring the reliability of the predicted human actions.
[0035] Step 102: Establish the vector space of node features according to the temporal arrangement of video frames.
[0036] Since node features are extracted from each video frame, the node features corresponding to each video frame are temporally correlated. By sorting the node features temporally, the spatial correlation represented by the changes in node features over time can be further determined. For example, when filming a person running, taking the shoulder joint as a node, the upper arm connected to the shoulder joint rotates periodically around the shoulder joint as a fulcrum. By observing the upper arm connected to the shoulder joint, the motion vector of the node corresponding to the shoulder joint can be determined. This motion vector contains information such as the node's position, direction of motion, and intensity of motion.
[0037] Therefore, by recording the motion data represented by the node features of each human joint in each video frame, a complete vector space can be constructed. This vector space stores the motion vector data of the joints during human movement, as represented by the node features in different video frames. By determining the vector data of a node, the specific action execution process of that node can be reconstructed.
[0038] Step 103: Calculate the target vector data after a preset time based on the vector space.
[0039] After constructing the vector space, the vector data stored therein, representing node features, can be used to determine the vector data after a preset time. For example, when the vector space of the video to be processed shows that the elbow of a person changes sequentially from the chest to the head, it can be considered that the person is performing a shooting preparation motion. Based on the speed of the elbow's rise, it can be predicted that the elbow will perform a forward thrust, i.e., the shooting motion, after 1.3 seconds. The vector of the elbow position performing the shooting motion is used as the target vector data.
[0040] In one possible embodiment, the target vector data after a preset time is calculated based on the vector space, including:
[0041] In the vector space, target node features are determined, which are used to represent feature data of at least one human joint; target node features are connected in temporal order to generate temporal node trajectories; based on the temporal node trajectories, target node feature values after a preset time are calculated, and target node feature values after the preset time are used as target vector data.
[0042] This embodiment provides a detailed explanation of the calculation of target vector data. After establishing the vector space, a target node feature corresponding to a node is selected within the vector space. This target node feature is expressed in the form of vector data. Since the vector data corresponding to target node features exist spatially connected within video frames, the vector data corresponding to target node features at the same temporal sequence can be expressed as temporal node trajectories. Based on the determination of the temporal node trajectories, the target vector data corresponding to the target node feature values after a preset time can be accurately determined. For example, if the temporal node trajectory is regular or involves a specific motion, subsequent node features can be determined based on existing node features.
[0043] This embodiment proposes a specific scheme for determining node feature values. It establishes corresponding temporal node trajectories based on node features, and the vector values between adjacent temporal nodes exhibit correlation in their changes. The node feature values after a preset time can be determined through the temporal node trajectories, and the determined node feature values can also be expressed in the form of target vector data.
[0044] In one possible embodiment, based on the time-series node trajectory, the target node feature value after a preset time is calculated, and the target node feature value after the preset time is used as target vector data, including:
[0045] Based on the time-series node trajectories, an autoregressive model is generated; based on the model parameters of the autoregressive model, the recursive parameters between the time-series node trajectories and the target node feature values after a preset time are calculated; combining the time-series node trajectories and the recursive parameters, the target node feature values are calculated.
[0046] This embodiment proposes a specific method for generating target vector data, using an autoregressive model to calculate the feature values of target nodes. An autoregressive model (AR model) is a statistical method for processing time series data. It uses the previous periods of the same variable, such as x (x1 to xt-1), to predict the performance of the current period xt, assuming a linear relationship between them. Because it evolved from linear regression in regression analysis, except that it uses x to predict x (itself) instead of x predicting y, it is called autoregressive. The autoregressive model emphasizes the linear correlation between data points, and its internal recursive function can effectively calculate and generate the feature values of target nodes based on the time series node trajectory.
[0047] Step 104: Control the image generator to generate a 3D human body image using the target vector data. The image generator is generated by the image initial generator and the image discriminator adversarially trained image initial generator.
[0048] In the shooting examples listed above, since the shooting preparation motion involves not only the elbow, but also the wrist, waist, and knee, the target vector data of the corresponding joints can be predicted simultaneously based on the vector changes of each node in the vector space. For example, by establishing the vector space of a jogging human body, the running vector data of each joint can be calculated, for example, after 3 seconds, and the human running motion can be further generated.
[0049] The generated target vector data is represented in a 2D image. However, there are limitations to further recognition of human figures displayed in 2D images. For example, when driver assistance functions adjust vehicle speed based on human movements, it is difficult to further determine the distance difference between the vehicle and the human figure based on the human figure in the 2D image. Therefore, this embodiment further proposes to generate a 3D human figure from the target vector data.
[0050] Since the positions of nodes corresponding to node features within the human body are fixed, the location of the target node features and the data of the executed actions can be accurately obtained by reading the target vector data corresponding to the target node features. For example, based on the vector data corresponding to the twisting of the waist and the position of the waist within the human body, a 3D image of the waist can be generated relatively accurately. Through the target node features of each node, a 3D image of the entire human body can be generated.
[0051] However, in the actual process of generating 3D human images, the target vector data used to generate the 3D human image is calculated and predicted from the features of other nodes in the vector space. Given that the node features themselves have a certain recognition error, the predicted target vector data also has a certain possibility of error. Therefore, the 3D human image directly generated from the target vector data will also have a certain degree of error. To address this problem, this embodiment uses a specific image generator to generate the 3D human image. This image generator is generated through adversarial training between an initial image generator and an image discriminator. It can adjust the 3D human image generated from the target vector data based on its own generator parameters, ensuring that the final generated 3D human image has a coherent motion trajectory with the motion parameters corresponding to the human node features in the video frame. This ensures that the human image in the video frame and the generated 3D human image conform to the laws of human kinematics.
[0052] This embodiment provides an image generator that can combine the positional correspondence of node features between human bodies and generate a 3D human body image corresponding to the target vector data with the action information corresponding to the target vector data. Using this image generator, after obtaining the target vector data, a 3D human body image can be generated quickly, avoiding the generation process being too long and affecting the execution of subsequent functions. For example, subsequent functions can assist in the generation of vehicle control strategies for driving functions.
[0053] In one feasible embodiment, the training process of the image generator includes:
[0054] Establish an image initialization generator;
[0055] Acquire training vector data, which includes human motion parameters;
[0056] Using training vector data, control the image initial generator to generate initial 3D human images;
[0057] Extract node feature vectors from the initial 3D human body image;
[0058] Input the node feature vectors into the image discriminator to generate action recognition results;
[0059] Based on the action recognition results, the parameters of the initial image generator are modified until the target node feature vector in the initial human body 3D image generated by the modified initial image generator is less than the action recognition result generated by the input image discriminator. Then, the modified initial image generator is used as the image generator.
[0060] Furthermore, the input node feature vectors are fed into the image discriminator to generate action recognition results, including:
[0061] Obtain the comparison feature vector set, which is used to record human motion data;
[0062] In the image discriminator, a gated recurrent unit model is used to compare the feature vector set with the input node feature vectors to obtain the vector comparison results.
[0063] Read the vector comparison results and generate action recognition results.
[0064] This embodiment further explains the training process of the image generator. The training process is primarily based on training data, which mainly represents the vector data corresponding to each node feature of the human body in various action states. Node features can represent the movements of human joints during human movement. Given that the training data is represented as vector data, it can be further represented as training vector data. During a set of human movements, for example, as the human body moves from position A to position B, the training vector data records the motion parameters of each joint (node feature) during walking, where the motion parameters of each joint represent human action parameters. In addition to training vector data, this embodiment further provides an image discriminator trained on AMASS (a large-scale motion capture dataset). The image discriminator contains a set of feature vectors representing human motion data, serving as a comparison feature vector set. When training the image generator, training vector data is input into the initial image generator. Since the human motion parameters represented in the training vector data are extracted from video frames, and video frames are divided into segments of the video to be processed at certain intervals, the additional training vector data extracted from the video frames will result in the loss of certain motion parameters. Specifically, the continuity between adjacent motion parameters will be lost. At the same time, since errors are inevitable in the process of extracting human parameters from a single image, the predicted parameters will inevitably have errors, resulting in the problem that the predicted 3D human motion does not conform to the laws of human kinematics.
[0065] Therefore, in practical use, image generators directly trained from training vector data may produce 3D human images with inaccurate motion prediction and / or the 3D images may not conform to the laws of human movement. Therefore, each embodiment incorporates an image discriminator, creating an adversarial relationship with the initial image generator. After the initial image generator generates an initial 3D human image, the image discriminator further discriminates this initial 3D image. Specifically, it extracts the node feature vectors from the initial 3D human image and, based on the motion recognition results generated after inputting the node feature vectors, further identifies whether the initial 3D human image generated by the initial image generator is accurate. The motion recognition results can be obtained using the softmax function. Specifically, after the image discriminator inputs the node feature vectors, the input results are fed into a single-layer GRU (Gate Recurrent Unit) model to resolve gradient explosion between output results, making the motion represented by the output results smoother. By identifying the output results, the discriminator uses the motion recognition results to determine whether there is coherence between the initial 3D human image generated by the initial image generator and the node features corresponding to the training vector data. If the action recognition result is considered to have poor continuity, the internal parameters of the initial image generator are modified until the target node feature vector in the initial human body 3D image generated by the initial image generator is less than the preset threshold when the action recognition result generated by the input image discriminator is less than the preset threshold. At this point, the output result of the initial image generator can be confirmed to have converged. The initial image generator, after continuous parameter adjustment, can be used as an image generator.
[0066] In the training process of the image generator in this embodiment, an image discriminator is added, and a compositional adversarial network is formed. The 3D human motion dataset is used for training. The image discriminator makes the 3D human motion generated by the generator obtained through adversarial training more in line with the laws of human kinematics, thus ensuring the reliability of the 3D image.
[0067] This embodiment provides a human body 3D image generation device, such as... Figure 2 The diagram shown is a structural diagram of the device in this embodiment, including:
[0068] Extraction module 21 is used to extract node features from video frames in the video to be processed, wherein the node features are used to represent the joint features of the human body in the video frame;
[0069] Module 22 is used to establish the vector space of the node features according to the temporal arrangement of video frames;
[0070] Calculation module 23 is used to calculate the target vector data after a preset time based on the vector space;
[0071] The generation module 24 is used as an image generator to generate a 3D human body image using the target vector data. The image generator is generated by adversarial training of the image initial generator and the image discriminator.
[0072] In one possible embodiment, the extraction module 21 is configured to:
[0073] Identify human image data in video frames of the video to be processed;
[0074] By combining the aforementioned portrait data, a human image is determined within the video frame;
[0075] Based on the human body movements represented by the portrait, node features are extracted from the portrait data.
[0076] In one possible embodiment, the computing module 23 is configured to:
[0077] In the vector space, target node features are determined, which are used to represent feature data of at least one human joint;
[0078] Connect the target node features in chronological order to generate a chronological node trajectory;
[0079] Based on the time-series node trajectory, the target node feature value is calculated after a preset time, and the target node feature value after the preset time is used as target vector data.
[0080] In one possible embodiment, the computing module 23 is configured to:
[0081] Based on the time-series node trajectories, an autoregressive model is generated;
[0082] Based on the model parameters of the autoregressive model, the recursive parameters between the time-series node trajectory and the feature value of the target node after the preset time are calculated.
[0083] The target node feature value is calculated by combining the time-series node trajectory and the recursive parameters.
[0084] In one possible embodiment, generation module 24 is configured to:
[0085] The image generator is controlled to generate a 3D image of the human body by combining the node features with the correspondence between the node features and the human body and using the target vector data.
[0086] In one possible embodiment, the human body 3D image generation apparatus further includes a training module 20, used for:
[0087] Establish an image initialization generator and an image discriminator;
[0088] Acquire training vector data, which includes human motion parameters;
[0089] Using the training vector data, the image initial generator is controlled to generate an initial 3D human image;
[0090] Extract the node feature vectors from the initial 3D human body image;
[0091] The node feature vector is input into the image discriminator to generate an action recognition result;
[0092] Based on the action recognition result, the parameters of the image initial generator are modified until the target node feature vector in the target initial human 3D image generated by the modified image initial generator is less than the action recognition result generated by the image discriminator. Then, the modified image initial generator is used as the image generator.
[0093] In one possible embodiment, training module 20 is used for:
[0094] Obtain a set of comparison feature vectors, which is used to record human motion data;
[0095] In the image discriminator, a gated recurrent unit model is used to compare the comparison feature vector set with the input node feature vector to obtain the vector comparison result;
[0096] Read the vector comparison results and generate action recognition results.
[0097] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0098] Based on the above Figure 1 The method shown, and Figure 2 To achieve the above objectives, this application also provides an electronic device, which can be configured on the end side of a vehicle (such as a new energy vehicle). This device includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor. The processor executes a computer program to implement the above-described virtual device embodiments. Figure 1 The method shown.
[0099] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0100] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0101] Based on the above Figure 1 The method illustrated in this application also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the method corresponding to any embodiment. The storage medium may further include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device and supports the operation of the information processing program and other software and / or programs. The network communication module is used to realize communication between the components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0102] Based on the aforementioned electronic device, this application embodiment also provides a vehicle, which may specifically include: such as Figure 2 The device shown or the electronic equipment described above. The vehicle may specifically be a new energy vehicle or a traditional vehicle, etc.
[0103] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the scheme of this embodiment, compared with the prior art, this embodiment first extracts node features from the video frames in the video to be processed. Node features are used to represent the features of the human image in the video frames. Then, according to the temporal arrangement of the video frames, a vector space of node features is established. Then, based on the vector space, target vector data after a preset time is calculated. Finally, an image generator is controlled to generate a 3D human image using the target vector data. The image generator is generated by adversarial training of the image initial generator and the image discriminator. This application can collect human node features in the video to be processed and establish a vector space containing changes in node features according to the temporal arrangement of the video frames. Based on the changes in features in the vector space, the vector data corresponding to the node features after a preset time can be predicted. Through the correspondence between node features and the human body, the generation of a 3D human image can be predicted. The technical solution of this application can determine the human body situation after a preset time by utilizing the changes in node features of the human image in the video frame, and simultaneously represent the human body in the form of a 3D image. This allows the human body image in the 2D image to be generated into a 3D human body image through prediction. In addition, during the process of generating the 3D human body image, this application further uses an image generator to generate the 3D human body image. The generator is generated through adversarial training between an image initial generator and an image discriminator. The 3D human body image generated by the image generator can be coherent with the 3D human body movement in the video frame and conforms to the laws of human kinematics. In some fields where human body movement or position is highly applied, advance prediction can be performed.
[0104] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0105] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for generating a 3D human body image, characterized in that, include: Extract node features from video frames in the video to be processed. The node features are used to represent the joint features of the human body in the video frames. Establish the vector space of the node features according to the temporal arrangement of video frames; Based on the vector space, calculate the target vector data after a preset time. A control image generator is used to generate a 3D human body image using the target vector data. The image generator is generated by adversarial training of the image initial generator and the image discriminator.
2. The method according to claim 1, characterized in that, The training process of the image generator includes: Establish an image initialization generator; Acquire training vector data, which includes human motion parameters; Using the training vector data, the image initial generator is controlled to generate an initial 3D human image; Extract the node feature vectors from the initial 3D human body image; Input the node feature vector into the image discriminator to generate action recognition results; Based on the action recognition result, the parameters of the image initial generator are modified until the target node feature vector in the target initial human 3D image generated by the modified image initial generator is less than the action recognition result generated by the input image discriminator. Then, the modified image initial generator is used as the image generator.
3. The method according to claim 2, characterized in that, The process of inputting the node feature vector into the image discriminator to generate action recognition results includes: Obtain a set of comparison feature vectors, which is used to record human motion data; In the image discriminator, a gated recurrent unit model is used to compare the comparison feature vector set with the input node feature vector to obtain the vector comparison result; Read the vector comparison results and generate action recognition results.
4. The method according to claim 1, characterized in that, Extracting node features from video frames in the video to be processed includes: Identify human image data in video frames of the video to be processed; By combining the aforementioned portrait data, a human image is determined within the video frame; Based on the human body movements represented by the portrait, node features are extracted from the portrait data.
5. The method according to claim 1, characterized in that, The step of calculating the target vector data after a preset time based on the vector space includes: In the vector space, target node features are determined, which are used to represent feature data of at least one human joint; Connect the target node features in chronological order to generate a chronological node trajectory; Based on the time-series node trajectory, the target node feature value is calculated after a preset time, and the target node feature value after the preset time is used as target vector data.
6. The method according to claim 5, characterized in that, The step of calculating the target node feature value after a preset time based on the time-series node trajectory includes: Based on the time-series node trajectories, an autoregressive model is generated; Based on the model parameters of the autoregressive model, the recursive parameters between the time-series node trajectory and the feature value of the target node after the preset time are calculated. The target node feature value is calculated by combining the time-series node trajectory and the recursive parameters.
7. A human body 3D image generation device, characterized in that, include: An extraction module is used to extract node features from video frames in the video to be processed, wherein the node features are used to represent the joint features of the human body in the video frame; A module is established to create a vector space for the node features according to the temporal arrangement of video frames; The calculation module is used to calculate the target vector data after a preset time based on the vector space; A generation module is used to control an image generator to generate a 3D human body image using the target vector data. The image generator is generated by adversarial training of an image initial generator and an image discriminator.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-6.
10. A vehicle, characterized in that, The vehicle is equipped with the device as described in claim 7, or the electronic device as described in claim 8.
Citation Information
Cited By
Prior for high-resolution image synthesis
US20250078397A1