Hair strand segment processing model training method and device and hair strand segment processing method
By training a hair segment processing model and using a multilayer perceptron and variational autoencoder to generate a latent vector space, the problems of low efficiency and poor effect in hair model generation are solved, achieving more uniform and neat hair generation and reducing the cost of 3D image construction.
Patent Information
- Application Number
- CN202310359361.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-06
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-04-06
AI Technical Summary
In existing technologies, the generation efficiency of hair models is low, the hair generation effect is poor, and the construction cost of 3D images is high. In particular, the lack of direction of hair line segments in multi-view hairstyle images leads to messiness.
By acquiring training data of sample hair segments, a hair segment processing model is trained based on standard orientation and position coordinates. A latent vector space is generated using a multilayer perceptron and variational autoencoder. The model is adjusted to improve the accuracy of the predicted orientation. The position and initial orientation of the hair segments are obtained by combining multi-view hairstyle images, and the predicted orientation is output.
It improves the production efficiency of hair models, resulting in more uniform and neat hair strands, and reduces the cost of 3D character construction.
Smart Images

Figure CN116432732B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of computer vision, augmented reality, virtual reality, and deep learning, and can be applied to scenarios such as metaverse and digital humans. Background Technology
[0002] In related technologies, 3D virtual avatars have wide application value in metaverse and digital human scenarios such as social media, live streaming, and games. Generating virtual avatars from multi-view hairstyle images can effectively meet users' personalized needs. However, hair modeling in virtual avatars is complex, with a large number of hair strands. Due to the lack of directionality in hair line segments, hair prediction often results in a messy appearance.
[0003] Therefore, improving the production efficiency of hair models, enhancing the hair generation effect, and reducing the cost of 3D character construction has become an important research direction. Summary of the Invention
[0004] This disclosure provides a training method, apparatus, and hair segment processing method for a hair segment processing model.
[0005] According to one aspect of this disclosure, a training method for a hair segment processing model is provided, the method comprising:
[0006] Obtain training data for sample hair segments, including the standard orientation and position coordinates of the sample hair segments in three-dimensional space;
[0007] The initial direction of the sample hair segment is obtained based on the standard direction, and the initial direction and position coordinates are input into the hair segment processing model;
[0008] The initial direction and position coordinates are encoded to generate the latent vector space of the sample hair line segment, and the latent vector space is decoded to obtain the predicted direction of the sample hair line segment in three-dimensional space.
[0009] The hair segment processing model is adjusted based on the standard direction and the predicted direction to obtain the target hair segment processing model.
[0010] This disclosure significantly enhances the performance of the hair segment processing model, and can strengthen the hair segment processing model based on data properties, thereby improving the production efficiency of hair models, enhancing the hair generation effect, and reducing the cost of 3D image construction.
[0011] According to one aspect of this disclosure, a method for processing hair segments is provided, the method comprising:
[0012] Acquire multi-view hairstyle images, and based on the multi-view hairstyle images, obtain the position coordinates and initial orientation of multiple hair strand segments in three-dimensional space;
[0013] For any hair segment among multiple hair segments, the position coordinates and initial direction of the hair segment are input into the target hair segment processing model, and the target hair segment processing model outputs the predicted direction of the hair segment in three-dimensional space.
[0014] The target hair segment processing model is a model trained using the training method described in the first aspect embodiment above.
[0015] In this embodiment of the disclosure, the directional direction of the hairstyle line segment can be obtained, which reduces the ambiguity in the hair prediction, makes the generated hair strands more uniform and neat, improves the effect of hair generation, and can also improve the production efficiency of the hair model.
[0016] According to another aspect of this disclosure, a training apparatus for a hair segment processing model is provided, comprising:
[0017] The acquisition module is used to acquire training data for sample hair segments. The training data includes the standard orientation and position coordinates of the sample hair segments in three-dimensional space.
[0018] The input module is used to obtain the initial direction of the sample hair segment based on the standard direction, and input the initial direction and position coordinates into the hair segment processing model;
[0019] The processing module is used to encode the initial direction and position coordinates, generate the latent vector space of the sample hair line segment, and decode the latent vector space to obtain the predicted direction of the sample hair line segment in three-dimensional space.
[0020] The adjustment module is used to adjust the hair segment processing model based on the standard direction and the predicted direction to obtain the target hair segment processing model.
[0021] According to another aspect of this disclosure, a hair segment processing apparatus is provided, comprising:
[0022] The first acquisition module is used to acquire multi-view hairstyle images and, based on the multi-view hairstyle images, acquire the position coordinates and initial direction of multiple hair strand segments in three-dimensional space;
[0023] The processing module is used to input the position coordinates and initial direction of any hair segment from multiple hair segments into the target hair segment processing model, and the target hair segment processing model outputs the predicted direction of the hair segment in three-dimensional space.
[0024] The target hair segment processing model is a model trained using the training device described in the third aspect embodiment.
[0025] According to another aspect of this disclosure, an electronic device is provided, including at least one processor, and
[0026] A memory that is communicatively connected to at least one processor; wherein,
[0027] The memory stores instructions that can be executed by at least one processor, which enables the at least one processor to perform the training method of the hair segment processing model of the first aspect of the present disclosure or to perform the hair segment processing method of the second aspect of the present disclosure.
[0028] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are configured to cause a computer to perform steps of a training method for a hair segment processing model according to an embodiment of a first aspect of this disclosure or to perform steps of a hair segment processing method according to an embodiment of a second aspect of this disclosure.
[0029] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the training method of the hair segment processing model of the first aspect of this disclosure or the steps of the hair segment processing method of the second aspect of this disclosure.
[0030] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0031] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0032] Figure 1 This is a flowchart of a training method for a hair segment processing model according to an embodiment of the present disclosure;
[0033] Figure 2 This is a flowchart of a training method for a hair segment processing model according to an embodiment of the present disclosure;
[0034] Figure 3 This is a schematic diagram of a training method for a hair segment processing model according to an embodiment of the present disclosure;
[0035] Figure 4 This is a flowchart of a hair segment processing method according to an embodiment of the present disclosure;
[0036] Figure 5 This is a schematic diagram of a hair segment processing method according to an embodiment of the present disclosure;
[0037] Figure 6 This is a flowchart of a hair segment processing method according to an embodiment of the present disclosure;
[0038] Figure 7 This is a structural diagram of a training device for a hair segment processing model according to an embodiment of the present disclosure;
[0039] Figure 8 This is a structural diagram of a hair segment processing apparatus according to an embodiment of the present disclosure;
[0040] Figure 9 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0041] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0042] The embodiments disclosed herein relate to the fields of artificial intelligence technology, such as computer vision and deep learning.
[0043] Artificial Intelligence (AI) is a new technological science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence.
[0044] Computer vision technology is a technique that simulates the human visual process, enabling computers to perceive their environment and perform human visual functions. It integrates technologies such as image processing, artificial intelligence, and pattern recognition.
[0045] Augmented Reality (AR) is a technology that calculates the position and angle of camera images in real time and adds corresponding images. It is a new technology that seamlessly integrates real-world information and virtual-world information. The goal of this technology is to overlay the virtual world onto the real world on the screen and allow for interaction.
[0046] Virtual Reality (VR), also known as virtual reality or virtual reality technology, is a new and practical technology that emerged in the 20th century. VR technology encompasses computer science, electronic information, and simulation technology. Its basic implementation relies primarily on computer technology, utilizing and integrating the latest advancements in 3D graphics, multimedia, simulation, display, and server technologies. With the help of computers and other equipment, it creates a realistic 3D virtual world that provides a multi-sensory experience, including visual, tactile, and olfactory sensations, giving the viewer a sense of immersion.
[0047] Deep learning (DL) is a new research direction in the field of machine learning (ML), bringing it closer to its original goal—artificial intelligence (AI). Deep learning learns the inherent patterns and hierarchical representations of sample data; the information gained during this learning process greatly aids in interpreting data such as text, images, and sound. Its ultimate goal is to enable machines to possess analytical and learning capabilities like humans, capable of recognizing data such as text, images, and sound. Deep learning is a complex machine learning algorithm that has achieved results in speech and image recognition far exceeding previous related technologies. Deep learning has also made significant progress in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech recognition, recommendation and personalization technologies, and other related fields. Deep learning enables machines to mimic human activities such as sight, hearing, and thought, solving many complex pattern recognition problems and significantly advancing artificial intelligence-related technologies.
[0048] The training method, apparatus, and hair segment processing method of this disclosure are described below with reference to the accompanying drawings.
[0049] Figure 1 This is a flowchart of a training method for a hair segment processing model according to an embodiment of the present disclosure, as shown below. Figure 1 As shown, the method includes the following steps:
[0050] S101, Obtain training data for sample hair segments. The training data includes the standard orientation and position coordinates of the sample hair segments in three-dimensional space.
[0051] In some implementations, directed hair segments can be randomly sampled based on existing open-source hair data as sample hair segments, thereby obtaining training data for the sample hair segments. The training sample data includes the standard orientation and position coordinates of the sample hair segments in three-dimensional space.
[0052] For example, training sample data can be (D″x,D″y,D″z) or (Px,Py,Pz), where (D″x,D″y,D″z) represents the three-dimensional direction, that is, the standard direction of the sample hair segment in three-dimensional space; (Px,Py,Pz) represents the three-dimensional coordinates, that is, the position coordinates of the sample hair segment in three-dimensional space.
[0053] S102: Obtain the initial direction of the sample hair segment based on the standard direction, and input the initial direction and position coordinates into the hair segment processing model.
[0054] In some implementations, the standard direction of the sample hair segment is reversed, and the reversed direction is used as the initial direction of the sample hair segment. In this case, the initial direction is opposite to the standard direction.
[0055] In some implementations, to increase the accuracy of model training, the standard direction of the sample hair segments can be randomly reversed, and the randomly reversed direction can be used as the initial direction for the sample hair segments. In this case, the initial direction and the standard direction may be opposite or the same. Since the hair segment processing model aims to correct reversed hair segments, it significantly enhances the performance of the hair segment processing model, and the hair segment processing model can be strengthened based on the properties of the data.
[0056] The initial direction and position coordinates are input into the hair segment processing model, which then interprets the initial direction and position coordinates to predict the direction of the sample hair segment in three-dimensional space.
[0057] S103 encodes the initial direction and position coordinates to generate the latent vector space of the sample hair segment, and decodes the latent vector space to obtain the predicted direction of the sample hair segment in three-dimensional space.
[0058] In some implementations, the hair segment processing model can be a variational autoencoder (VAE) network based on a multilayer perceptron (MLP). This VAE network can obtain a smooth latent vector space through self-supervised learning and has good generalization performance with less training data.
[0059] The initial direction and position coordinates are encoded by the hair segment processing model to generate the latent vector space of the sample hair segment. In other words, a smooth latent vector space is constructed based on the initial direction and position coordinates. Furthermore, the latent vector space is decoded to obtain the predicted direction of the sample hair segment in three-dimensional space.
[0060] S104, The hair segment processing model is adjusted based on the standard direction and the predicted direction to obtain the target hair segment processing model.
[0061] In some implementations, the hair segment processing model is adjusted based on loss functions of standard and predicted directions, and the adjusted hair segment processing model is trained using the next sample hair segment.
[0062] In some implementations, if the preset number of iterations is reached or the loss function converges to the preset loss threshold, the training of the hair segment processing model is considered complete, and the target hair segment processing model is obtained.
[0063] The target hair segment processing model can generate a continuous and smooth three-dimensional direction prediction result in three-dimensional space based on the hair segment, thereby obtaining the predicted direction of any hair segment.
[0064] In this embodiment, training data for sample hair segments is obtained. The training data includes the standard orientation and position coordinates of the sample hair segments in three-dimensional space. An initial orientation of the sample hair segments is obtained based on the standard orientation, and the initial orientation and position coordinates are input into the hair segment processing model. The initial orientation and position coordinates are encoded to generate a latent vector space for the sample hair segments, and the latent vector space is decoded to obtain the predicted orientation of the sample hair segments in three-dimensional space. This disclosure significantly enhances the performance of the hair segment processing model, allowing for model strengthening based on data properties. This improves hair model production efficiency, enhances hair generation effects, and reduces the cost of three-dimensional image construction.
[0065] Figure 2 This is a flowchart of a training method for a hair segment processing model according to an embodiment of the present disclosure, as shown below. Figure 2 As shown, the method includes the following steps:
[0066] S201, Obtain training data for the sample hair segments. The training data includes the standard orientation and position coordinates of the sample hair segments in three-dimensional space.
[0067] S202: Obtain the initial direction of the sample hair segment based on the standard direction, and input the initial direction and position coordinates into the hair segment processing model.
[0068] For a description of steps S201 to S202, please refer to the content in the above embodiments, which will not be repeated here.
[0069] It should be noted that the hair segment processing model includes an undirected segment encoder and a directed segment decoder.
[0070] S203, the first coding layer in the undirected line segment encoder, performs position encoding on the position coordinates to obtain position feature representation.
[0071] The initial direction and position coordinates are input into the undirected line segment encoder of the hair segment processing model. The first encoding layer in the undirected line segment encoder performs position encoding on the position coordinates. The position coordinates are enhanced by position encoding, and the obtained position feature representation identifies the coordinate information of the position to be sampled.
[0072] S204 uses a multilayer perceptron in the undirected line segment encoder to resolve the initial direction.
[0073] The initial direction is obtained by parsing the initial direction using a multilayer perceptron (MLP).
[0074] S205, the second coding layer in the undirected line segment encoder, jointly encodes the parsed initial orientation and position feature representations to generate the latent vector space.
[0075] Alternatively, the latent vector space can be obtained using the following formula:
[0076] Latent=Encoder(PE(Px,Py,Pz),(D′x,D′y,D′z))
[0077] Where Latent represents the latent vector space, PE(...) represents position encoding, Encoder(...) represents common encoding, and (D′x,D′y,D′z) represents the initial direction after parsing.
[0078] In this embodiment of the disclosure, an undirected line segment encoder jointly understands the parsed initial direction and position feature representations, thereby constructing a smooth latent vector space.
[0079] S206 uses a directed line segment decoder to perform multi-layer perception on the latent vector space to obtain the predicted direction of the hair line segment in three-dimensional space.
[0080] Alternatively, the predicted direction can be obtained using the following formula:
[0081] (Dx, Dy, Dz) = Decoder(Latent)
[0082] Where (Dx,Dy,Dz) represents the prediction direction, and Decoder(...) represents the decoding direction.
[0083] In this embodiment of the disclosure, the directed line segment decoder generates a directed direction (i.e., a predicted direction) based on a continuous and smooth latent vector space. This smoothness ensures that a continuous direction signal can be produced in three-dimensional space. Since the position coordinates are used as input, the hair segment processing model can effectively predict the direction from different three-dimensional positions.
[0084] S207, Determine the loss function of the hair segment processing model based on the standard direction and the predicted direction.
[0085] In some implementations, the absolute value loss function is calculated based on the standard direction and the predicted direction. In other words, the loss function of the hair segment processing model is determined to be the L1 loss function.
[0086] S208, the hair segment processing model is adjusted based on the loss function to obtain the target hair segment processing model.
[0087] In some implementations, the hair segment processing model is adjusted in reverse based on the L1 loss function to obtain the target hair segment processing model.
[0088] like Figure 3 As shown in this embodiment, the standard direction of the sample hair segment is randomly reversed, and the randomly reversed direction is used as the initial direction corresponding to the sample hair segment. At this time, the initial direction and the standard direction may be opposite or the same. The initial direction and position coordinates are encoded based on the undirected line segment encoder (LD Encoder) to generate the latent vector space of the sample hair segment. The latent vector space is decoded based on the directed line segment decoder (LD Decoder) to obtain the predicted direction of the sample hair segment in three-dimensional space.
[0089] In this embodiment, the first encoding layer of the undirected line segment encoder performs position encoding on the position coordinates to obtain a position feature representation. The multilayer perceptron in the undirected line segment encoder parses the initial direction. The second encoding layer of the undirected line segment encoder jointly encodes the parsed initial direction and position feature representation to generate a latent vector space. The directed line segment decoder performs multilayer perceptron processing on the latent vector space to obtain the predicted direction of the hair segment in three-dimensional space. This disclosure significantly enhances the performance of the hair segment processing model, enabling enhancement of the hair segment processing model based on data properties. It can improve the efficiency of hair model production, enhance the hair generation effect, and reduce the cost of three-dimensional image construction.
[0090] Figure 4 This is a flowchart of a hair segment processing method according to an embodiment of the present disclosure, as follows: Figure 4 As shown, the method includes the following steps:
[0091] S401, acquire multi-view hairstyle images, and based on the multi-view hairstyle images, acquire the position coordinates and initial direction of multiple hair strand segments in three-dimensional space.
[0092] In some implementations, multi-view hairstyle images can be obtained from a preset database; in others, multi-view hairstyle images of the target object can be acquired using an image acquisition device, such as a camera.
[0093] In this embodiment of the disclosure, pose estimation is performed based on multi-view hairstyle images to obtain head pose data, and then three-dimensional directed line segment extraction is performed to obtain the position coordinates and initial direction of multiple hair line segments in three-dimensional space.
[0094] S402: For any hair segment among multiple hair segments, input the position coordinates and initial direction of the hair segment into the target hair segment processing model, and the target hair segment processing model outputs the predicted direction of the hair segment in three-dimensional space.
[0095] like Figure 5 As shown, for any hair segment among multiple hair segments 510, the position coordinates and initial direction of the hair segment are input into the target hair segment processing model. The target hair segment processing model jointly understands the position coordinates and initial direction of the hair segment and outputs the predicted direction of the hair segment in three-dimensional space to obtain the directed hair segment 520.
[0096] The target hair segment processing model is a model trained using the training method described above.
[0097] In this embodiment, multi-view hairstyle images are acquired, and the position coordinates and initial directions of multiple hair segments in three-dimensional space are obtained based on these images. For any hair segment among the multiple hair segments, the position coordinates and initial direction are input into a target hair segment processing model, which then outputs the predicted direction of the hair segment in three-dimensional space. This embodiment can obtain the directional direction of hairstyle segments, reducing ambiguity in hair prediction, resulting in more uniform and neat hair strands in the generated hairstyle, improving the hair generation effect, and also increasing the production efficiency of the hair model.
[0098] Figure 6 This is a flowchart of a hair segment processing method according to an embodiment of the present disclosure, as follows: Figure 6 As shown, the method includes the following steps:
[0099] S601, acquire multi-view hairstyle images.
[0100] For a description of step S601, please refer to the description in the above embodiments, which will not be repeated here.
[0101] S602 calls the 3D reconstruction tool to perform inter-frame pose estimation on multi-view hairstyle images and obtain global pose 3D data.
[0102] In this embodiment of the disclosure, the 3D reconstruction tool Colmap can be invoked to perform inter-frame pose estimation on multi-view hairstyle images to obtain global pose 3D data. COLMAP is a general-purpose Structure of Motion (SfM) and Multi-View Stereo (MVS) pipeline with both graphical and command-line interfaces. It provides extensive functionality for the reconstruction of ordered and unordered image sets.
[0103] S603 calls the motion capture system to estimate the head pose from the global 3D pose data and obtains the head pose 3D data.
[0104] The motion capture system Mocap is invoked to estimate the head pose from the global 3D pose data, thus obtaining the 3D head pose data. Motion capture is a computer vision technique used to estimate the 3D pose (position and orientation) of a target object using an external localization mechanism.
[0105] S604 performs neural rendering reconstruction on the three-dimensional head pose data and extracts the position coordinates and initial orientation of multiple hair segments in three-dimensional space.
[0106] The three-dimensional head pose data is reconstructed using neural rendering, and then three-dimensional directed line segment extraction is performed to extract the position coordinates and initial direction of multiple hair line segments in three-dimensional space.
[0107] S605: For any hair segment among multiple hair segments, input the position coordinates and initial direction of the hair segment into the target hair segment processing model, and the target hair segment processing model outputs the predicted direction of the hair segment in three-dimensional space.
[0108] For a description of step S605, please refer to the description in the above embodiments, which will not be repeated here.
[0109] S606 performs differentiable hair fitting based on the position coordinates and predicted directions of multiple hair segments to obtain a hairstyle model corresponding to a multi-view hairstyle image.
[0110] In this embodiment of the disclosure, after obtaining the predicted direction of the hair strands in three-dimensional space, differentiable hair strand fitting can be performed based on the position coordinates and predicted direction of multiple hair strands to construct a hairstyle model, and a hairstyle model library can be generated based on the constructed hairstyle model to facilitate the customization of personalized virtual images for users.
[0111] In this embodiment, a 3D reconstruction tool is invoked to perform inter-frame pose estimation on multi-view hairstyle images to obtain global pose 3D data. A motion capture system is then invoked to perform head pose estimation on the global pose 3D data to obtain head pose 3D data. Neural rendering reconstruction is performed on the head pose 3D data, and the position coordinates and initial directions of multiple hair segments in 3D space are extracted. For any hair segment among these multiple hair segments, the position coordinates and initial direction are input into a target hair segment processing model, which then outputs the predicted direction of the hair segment in 3D space. This disclosure can obtain the directional direction of hairstyle segments, reducing ambiguity in hair prediction, resulting in more uniform and neat hair strands in the generated hairstyle, and improving the hair generation effect.
[0112] Figure 7 This is a structural diagram of a training device for a hair segment processing model according to an embodiment of the present disclosure, as shown below. Figure 7 As shown, the training device 700 for the hair segment processing model includes:
[0113] The acquisition module 710 is used to acquire training data of the sample hair segments. The training data includes the standard orientation and position coordinates of the sample hair segments in three-dimensional space.
[0114] The input module 720 is used to obtain the initial direction of the sample hair segment based on the standard direction, and input the initial direction and position coordinates into the hair segment processing model;
[0115] The processing module 730 is used to encode the initial direction and position coordinates, generate the latent vector space of the sample hair line segment, and decode the latent vector space to obtain the predicted direction of the sample hair line segment in three-dimensional space.
[0116] The adjustment module 740 is used to adjust the hair segment processing model based on the standard direction and the predicted direction to obtain the target hair segment processing model.
[0117] In some implementations, the hair segment processing model includes an undirected segment encoder, processing module 730, which is also used for:
[0118] The position coordinates are encoded by the first coding layer in the undirected line segment encoder to obtain the position feature representation;
[0119] The initial direction is analyzed by the multilayer perceptron in the undirected line segment encoder;
[0120] The second coding layer in the undirected line segment encoder co-encodes the parsed initial orientation and position feature representations to generate a latent vector space.
[0121] In some implementations, the hair segment processing model includes a directed segment decoder, and the processing module 730 is also used for:
[0122] A directed line segment decoder performs multi-layer perception on the latent vector space to obtain the predicted direction of the hair line segment in three-dimensional space.
[0123] In some implementations, module 740 is also used for:
[0124] Based on the standard direction and the predicted direction, determine the loss function of the hair segment processing model;
[0125] The hair segment processing model is adjusted based on the loss function to obtain the adjusted hair segment processing model.
[0126] In some implementations, the input module 720 is also used for:
[0127] The standard direction of the sample hair segments is randomly reversed, and the randomly reversed direction is used as the initial direction of the sample hair segments.
[0128] This disclosure significantly enhances the performance of the hair segment processing model, can strengthen the hair segment processing model based on data properties, and can also improve the production efficiency of hair models, improve the hair generation effect, and reduce the cost of 3D image construction.
[0129] Figure 8 This is a structural diagram of a hair segment processing apparatus according to an embodiment of the present disclosure, as shown below. Figure 8 As shown, the hair segment processing device 800 includes:
[0130] The first acquisition module 810 is used to acquire multi-view hairstyle images and acquire the position coordinates and initial direction of multiple hair strand segments in three-dimensional space based on the multi-view hairstyle images;
[0131] Processing module 820 is used to input the position coordinates and initial direction of any hair segment from multiple hair segments into the target hair segment processing model, and the target hair segment processing model outputs the predicted direction of the hair segment in three-dimensional space.
[0132] The target hair segment processing model is a model trained using the training device 700 described above.
[0133] In some implementations, the first acquisition module 810 is also used for:
[0134] Use 3D reconstruction tools to perform inter-frame pose estimation on multi-view hairstyle images to obtain global pose 3D data;
[0135] The motion capture system is invoked to estimate the head pose from the global 3D pose data, thereby obtaining the 3D head pose data.
[0136] The three-dimensional head pose data is reconstructed using neural rendering, and the position coordinates and initial orientation of multiple hair segments in three-dimensional space are extracted.
[0137] In some implementations, the device also includes a second acquisition module 830, for:
[0138] Differentiable hair strand fitting is performed based on the position coordinates and predicted directions of multiple hair strands to obtain hairstyle models corresponding to multi-view hairstyle images.
[0139] In this embodiment of the disclosure, the directional direction of the hairstyle line segment can be obtained, which reduces the ambiguity in the hair prediction, makes the generated hair strands more uniform and neat, improves the effect of hair generation, and can also improve the production efficiency of the hair model.
[0140] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0141] Figure 9 This is a block diagram of an electronic device used to implement embodiments of the present disclosure. The electronic device can implement the training method or hair segment processing method of the hair segment processing model of the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0142] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0143] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0144] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the training method for a hair segment processing model or a hair segment processing method. For example, in some embodiments, the training method for a hair segment processing model or a hair segment processing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the training method for a hair segment processing model or a hair segment processing method described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured by any other suitable means (e.g., by means of firmware) to perform a training method or a hair segment processing method for the hair segment processing model.
[0145] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0146] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0147] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0148] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0149] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0150] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0151] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0152] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method of training a hair strand segment processing model, wherein, The method comprises: obtaining training data of a sample hair strand segment, the training data comprising standard direction and position coordinates of the sample hair strand segment in a three-dimensional space; obtaining an initial direction of the sample hair strand segment based on the standard direction, and inputting the initial direction and the position coordinates into a hair strand segment processing model; encoding the initial direction and the position coordinates to generate a latent vector space of the sample hair strand segment, and decoding the latent vector space to obtain a predicted direction of the sample hair strand segment in the three-dimensional space; adjusting the hair strand segment processing model based on the standard direction and the predicted direction to obtain a target hair strand segment processing model; wherein the hair strand segment processing model comprises an undirected line segment encoder, and the encoding of the initial direction and the position coordinates to generate the latent vector space of the sample hair strand segment comprises: position encoding of the position coordinates by a first encoding layer in the undirected line segment encoder to obtain a position feature representation; analysis of the initial direction by a multi-layer perceptron in the undirected line segment encoder; and joint encoding of the analyzed initial direction and the position feature representation by a second encoding layer in the undirected line segment encoder to generate the latent vector space; wherein the hair strand segment processing model comprises a directed line segment decoder, and the decoding of the latent vector space to obtain the predicted direction of the sample hair strand segment in the three-dimensional space comprises: multi-layer perception of the latent vector space by the directed line segment decoder to obtain the predicted direction of the hair strand segment in the three-dimensional space.
2. The method of claim 1, wherein, The adjusting of the hair strand segment processing model based on the standard direction and the predicted direction comprises: determining a loss function of the hair strand segment processing model according to the standard direction and the predicted direction; adjusting the hair strand segment processing model based on the loss function to obtain an adjusted hair strand segment processing model.
3. The method of claim 2, wherein, The obtaining of the initial direction of the sample hair strand segment based on the standard direction comprises: randomly reversing the standard direction of the sample hair strand segment, and taking the randomly reversed direction as the initial direction corresponding to the sample hair strand segment.
4. A method of treating a hair strand segment, wherein, The method comprises: obtaining a multi-view hairstyle image, and obtaining position coordinates and initial directions of a plurality of hair strand segments in a three-dimensional space based on the multi-view hairstyle image; for any hair strand segment in the plurality of hair strand segments, inputting the position coordinates and the initial direction of the hair strand segment into a target hair strand segment processing model, and outputting a predicted direction of the hair strand segment in the three-dimensional space by the target hair strand segment processing model; wherein the target hair strand segment processing model is a model trained by the training method according to any one of claims 1-3.
5. The method of claim 4, wherein, The obtaining of the position coordinates and the initial directions of the plurality of hair strand segments in the three-dimensional space based on the multi-view hairstyle image comprises: calling a three-dimensional reconstruction tool to perform inter-frame pose estimation on the multi-view hairstyle image to obtain global pose three-dimensional data; Call the motion capture system to perform head pose estimation on the global pose three-dimensional data, and obtain head pose three-dimensional data; Perform neural rendering reconstruction on the head pose three-dimensional data, and extract the position coordinates and initial directions of the plurality of hair strand line segments in the three-dimensional space.
6. The method of claim 4, wherein, Also includes: Based on the position coordinates and predicted directions of the plurality of hair strand line segments, perform differentiable hair fitting to obtain a hairstyle model corresponding to the multi-view hairstyle image.
7. A device for training a hair strand segment processing model, wherein, Includes: An acquisition module is configured to acquire training data of a sample hair strand line segment, the training data including a standard direction and position coordinates of the sample hair strand line segment in a three-dimensional space; An input module is configured to acquire an initial direction of the sample hair strand line segment based on the standard direction, and input the initial direction and the position coordinates into a hair strand line segment processing model; A processing module is configured to encode the initial direction and the position coordinates, generate a latent vector space of the sample hair strand line segment, and decode the latent vector space to obtain a predicted direction of the sample hair strand line segment in the three-dimensional space; An adjustment module is configured to adjust the hair strand line segment processing model based on the standard direction and the predicted direction to obtain a target hair strand line segment processing model; The hair strand line segment processing model includes an undirected line segment encoder, and the processing module is further configured to: perform position encoding on the position coordinates by a first encoding layer in the undirected line segment encoder to obtain a position feature representation; analyze the initial direction by a multi-layer perception in the undirected line segment encoder; and jointly encode the analyzed initial direction and the position feature representation by a second encoding layer in the undirected line segment encoder to generate the latent vector space; The hair strand line segment processing model includes a directed line segment decoder, and the processing module is further configured to: perform multi-layer perception on the latent vector space by the directed line segment decoder to obtain the predicted direction of the hair strand line segment in the three-dimensional space.
8. The apparatus of claim 7, wherein, The adjustment module is further configured to: Determine a loss function of the hair strand line segment processing model according to the standard direction and the predicted direction; Adjust the hair strand line segment processing model based on the loss function to obtain an adjusted hair strand line segment processing model.
9. The apparatus of claim 8, wherein, The input module is further configured to: Randomly reverse the standard direction of the sample hair strand line segment, and use the randomly reversed direction as the initial direction corresponding to the sample hair strand line segment.
10. A hair strand section treatment device, wherein, Includes: A first acquisition module is configured to acquire a multi-view hairstyle image, and acquire position coordinates and initial directions of a plurality of hair strand line segments in a three-dimensional space based on the multi-view hairstyle image; A processing module is configured to input the position coordinates and the initial directions of any hair strand line segment in the plurality of hair strand line segments into a target hair strand line segment processing model, and output a predicted direction of the hair strand line segment in the three-dimensional space by the target hair strand line segment processing model; The target hair strand line segment processing model is a model trained by the training device of any one of claims 7-9.
11. The apparatus of claim 10, wherein, The first acquisition module is further configured to: Call a three-dimensional reconstruction tool to perform inter-frame pose estimation on the multi-view hairstyle image, and obtain global pose three-dimensional data; Call a motion capture system to perform head pose estimation on the global pose three-dimensional data, and obtain head pose three-dimensional data; Perform neural rendering reconstruction on the head pose three-dimensional data, and extract position coordinates and initial directions of the plurality of hair strand line segments in a three-dimensional space.
12. The apparatus of claim 10, wherein, Further comprising a second obtaining module configured to: Perform differentiable hair fitting based on the position coordinates and predicted directions of the plurality of hair strand line segments, to obtain a hairstyle model corresponding to the multi-view hairstyle image.
13. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-3 or perform the method of any one of claims 4-6.
14. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the steps of the method according to any one of claims 1-3 or perform the steps of the method according to any one of claims 4-6.
15. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-3 or implements the method according to any one of claims 4-6.
Citation Information
Patent Citations
Three-dimensional hair style generation method and device, electronic equipment and storage medium
CN115409922A
Three-dimensional hair style generation method and device, electronic equipment and storage medium
CN115661375A