Digital person business card generation method, electronic equipment, medium and program product

By preprocessing the input images of the target user and establishing the binding relationship of the 3D Gaussian point cloud, combined with the generation of digital human voice data, the problem of high cost and insufficient accuracy in digital human business card generation is solved, realizing low-cost and highly realistic digital human business card generation, and improving user experience and information interaction efficiency.

CN121564174APending Publication Date: 2026-02-24INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511690227.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing digital persona generation technologies are costly and lack accuracy, making it difficult to achieve efficient and low-cost dynamic interaction and emotional expression.

Method used

By acquiring and preprocessing the input images of the target user, a binding relationship between the 3D Gaussian point cloud and the skeletal hierarchical structure is established. Combined with digital human voice data, a realistic digital human business card is generated, including 3D spatial coordinate alignment and mapping of image feature data, reconstruction into a 3D Gaussian point cloud, image rendering of the binding relationship, and generation of digital human voice data.

Benefits of technology

It reduces the cost of generating digital human business cards, enhances the realism and dynamic interactivity of digital human images, achieves deep binding with the target user's image, and improves the efficiency of information interaction and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564174A_ABST
    Figure CN121564174A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of artificial intelligence, and discloses a digital person business card generation method, electronic equipment, a medium and a program product. The method comprises the following steps: acquiring an input image of a target user in digital person business card generation, and preprocessing the input image to obtain image feature data; performing three-dimensional space coordinate alignment and mapping on the image feature data to obtain image three-dimensional data; reconstructing the image three-dimensional data into a three-dimensional Gaussian point cloud; establishing a binding relationship between the three-dimensional Gaussian point cloud and a skeleton hierarchical structure in the three-dimensional human body modeling tool; performing image rendering according to the binding relation to obtain a digital human model; generating corresponding digital human voice data according to the reply data in the digital human interaction process; and the digital human model and the digital human voice data are combined and displayed to obtain the digital human business card, so that a single-picture realistic digital human technology can be realized, and the fidelity of a digital human image is improved and the digital human image is deeply bound with a user image on the basis of reducing the data volume required for making the digital human.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology and can be applied to the field of financial technology. In particular, it relates to a method, electronic device, medium, and program product for generating a digital personal business card. Background Technology

[0002] For information exchange, users typically exchange business cards offline to maintain contact. To address the problems of easily lost and difficult-to-organize information on paper business cards, electronic business cards emerged. However, electronic business cards only provide static text and image information, lacking dynamic interaction and emotional expression capabilities. To solve the problem of electronic business cards' inability to interact dynamically, digital human business card technology has become a hot topic.

[0003] Traditional digital human business card generation typically relies on network or voxel-based 3D modeling methods to create digital human models. This results in extensive data collection and complex modeling processes, leading to high costs and low efficiency. For example, creating a high-quality digital human model can cost tens of thousands of yuan or even more, limiting the application scenarios for digital human business cards. Furthermore, when dealing with complex geometric structures and subtle expressions in digital human models, 3D models often suffer from high computational costs but insufficient accuracy.

[0004] Therefore, there is an urgent need to provide a new method for generating digital human models, reduce the cost of generating digital human business cards, and improve accuracy, so as to broaden the application scenarios of digital human business cards. Summary of the Invention

[0005] This invention provides a method, electronic device, medium, and program product for generating digital human business cards, in order to reduce the amount of data required for digital human creation and improve the realism of the digital human image.

[0006] According to one aspect of the present invention, a method for generating a digital personal business card is provided, the method comprising:

[0007] The input image of the target user in the digital human business card generation is obtained, and the input image is preprocessed to obtain image feature data;

[0008] The image feature data is aligned and mapped in three-dimensional space to obtain three-dimensional image data; the three-dimensional image data is then reconstructed into a three-dimensional Gaussian point cloud.

[0009] Establish the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy in the 3D human body modeling tool; and perform image rendering based on the binding relationship to obtain a digital human model.

[0010] Based on the response data during the digital human interaction process, corresponding digital human voice data is generated; the digital human model and the digital human voice data are combined and displayed to obtain a digital human business card.

[0011] According to another aspect of the present invention, a digital personal business card generation apparatus is provided, the apparatus comprising:

[0012] The image feature data generation module is used to acquire the input image of the target user in the digital human business card generation process, and to preprocess the input image to obtain image feature data.

[0013] The Gaussian point cloud generation module is used to align and map the image feature data in three-dimensional space to obtain three-dimensional image data; and to reconstruct the three-dimensional image data into a three-dimensional Gaussian point cloud.

[0014] The digital human model generation module is used to establish the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy structure in the 3D human body modeling tool; and to perform image rendering based on the binding relationship to obtain a digital human model.

[0015] The digital human business card generation module is used to generate corresponding digital human voice data based on the response data during the digital human interaction process; the digital human model and the digital human voice data are combined and displayed to obtain the digital human business card.

[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the digital human business card generation method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the digital human business card generation method according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the digital human business card generation method described in any embodiment of the present invention.

[0022] The technical solution of this invention involves acquiring the input image of the target user in the digital human business card generation process and preprocessing the input image to obtain image feature data; aligning and mapping the image feature data in three-dimensional space coordinates to obtain three-dimensional image data; reconstructing the three-dimensional image data into a three-dimensional Gaussian point cloud; establishing a binding relationship between the three-dimensional Gaussian point cloud and the skeletal hierarchy structure in a three-dimensional human modeling tool; rendering the image based on the binding relationship to obtain a digital human model; generating corresponding digital human voice data based on the response data during the digital human interaction process; and combining and displaying the digital human model and digital human voice data to obtain a digital human business card. This solves the problems of high cost and insufficient accuracy in digital human model generation. By aligning the input image in three-dimensional space coordinates and then reconstructing it into a three-dimensional Gaussian point cloud for digital human model building, single-image realistic digital human technology can be realized. This reduces the amount of data required for digital human production, lowers the cost of digital human production, and improves the realism of the digital human image, deeply binding it with the target user's image.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a method for generating a digital human business card according to Embodiment 1 of the present invention;

[0026] Figure 2 This is a flowchart of a method for generating a digital human business card according to Embodiment 2 of the present invention;

[0027] Figure 3 This is a schematic diagram of the structure of a digital human business card generation device according to Embodiment 3 of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the digital human business card generation method of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] Example 1

[0032] Figure 1 This is a flowchart of a method for generating a digital human business card according to Embodiment 1 of the present invention. This embodiment is applicable to situations where products or resources are introduced through digital human business cards. The method can be executed by a digital human business card generating device, which can be implemented in hardware and / or software. This device can be configured in an electronic device, such as a mobile device like a mobile phone, tablet computer (PAD), or wearable device, or a personal computer (PC). Figure 1 As shown, the method includes:

[0033] Step 110: Obtain the input image of the target user in the digital human business card generation process, and preprocess the input image to obtain image feature data.

[0034] The target user can be anyone who wants to create a digital avatar business card featuring their own image. For example, the target user could be a client manager at an application organization who wants to use the digital avatar business card to introduce products. The input image can be a two-dimensional image. The number of input images can be small, such as a single image. In this invention, the input image of the target user is obtained only after user authorization, and the method of acquisition is legal and legitimate. The preprocessing of the input image can include various processes, such as background removal, image segmentation, feature extraction, normalization, and one or more of the following: feature enhancement.

[0035] Optionally, the input image is preprocessed to obtain image feature data, including: dividing the input image into multiple image blocks, extracting corresponding feature sub-vectors for each image block, and adding position encoding to the feature sub-vectors; performing layer normalization on the feature sub-vectors to obtain a set of feature tokens, and using the set of feature tokens as image feature data.

[0036] For example, the input image can be divided into multiple image patches of the same size. For instance, the input image can be divided into 16×16 pixel image patches. The number of image patches can be 64×64. Each image patch can undergo feature extraction through a linear layer to obtain a feature vector. .in, It is the i-th image patch. The dimension of the feature vector. This represents a linear layer. Positional encoding can be added to the feature vectors, meaning the feature vectors are updated to... . This represents the positional encoding of the i-th image patch. Layer normalization is applied to the feature sub-vectors with added positional encodings to obtain the feature token set. . Representation layer normalization processing. Image feature data can be obtained by summing the feature token sets of each image patch, i.e. . Dimensions of image feature data.

[0037] Image preprocessing can extract high-quality, high-resolution features from the input image. High-quality image feature extraction can ensure that the digital human model is more realistic in terms of detail.

[0038] Step 120: Align and map the image feature data in three-dimensional space to obtain three-dimensional image data; reconstruct the three-dimensional image data into a three-dimensional Gaussian point cloud.

[0039] Aligning and mapping image feature data to three-dimensional spatial coordinates facilitates the creation of three-dimensional digital human models. There are various ways to map image feature data to three-dimensional spatial coordinates. For example, deep learning models, depth sensors, or multi-view stereo matching can be used to align and map image feature data to three-dimensional spatial coordinates.

[0040] In an optional embodiment of the present invention, the image feature data is aligned and mapped in three-dimensional space coordinates to obtain three-dimensional image data, including: splicing the image feature data and UV tokens in the pre-learned texture coordinate system to obtain a splicing result; and using a multi-head self-attention mechanism and a feedforward network to perform feature fusion and enhancement on the splicing result to obtain three-dimensional image data.

[0041] In UV tokens, U and V refer to two dimensions in UV coordinates. UV coordinates are a two-dimensional texture coordinate system used to map textures onto the surface of a three-dimensional model. U represents the horizontal direction, and V represents the vertical direction, equivalent to the X and Y axes in a two-dimensional Cartesian coordinate system. Using UV coordinates, each point on a two-dimensional texture image can be mapped to a point on the surface of a three-dimensional model, thus ensuring the correct display of the texture map on the model. UV tokens can be obtained through pre-learning. By concatenating image feature data and UV tokens, image features can be aligned and mapped to UV coordinates in three-dimensional space.

[0042] Data augmentation can be performed on the stitched results. For example, feature fusion and enhancement can be performed using a multi-head self-attention mechanism and a feedforward network to obtain 3D image data. Specifically, the 3D image data is... .in, This represents the UV token, where F is the image feature data. This indicates a splicing operation. This is a multi-head self-attention mechanism. To enhance the 3D data of the processed image.

[0043] By enhancing image features, they can be made more suitable for the construction and animation of 3D digital human models, which helps to improve the stability and consistency of digital human models under different lighting conditions and viewpoints.

[0044] To generate a digital human model from a single 2D image, 3D Gaussian point clouds can be obtained by reconstructing the image's 3D data. Gaussian point clouds, based on a Gaussian distribution, describe the geometric shape and appearance features of the 3D digital human model. The method for reconstructing 3D image data into a 3D Gaussian point cloud involves determining a Gaussian parameter set, which may include: the center coordinates of the Gaussian points, rotation angle, rotation scale, and color attributes.

[0045] For example, the center coordinates (x, y, z) of each Gaussian point in space are determined using a 3D coordinate prediction network based on the 3D image data; the rotation angle of the Gaussian ellipsoid is represented by quaternions or Euler angles. The rotation scale *s* controlling the Gaussian ellipsoid is determined, which may include scaling factors (sx, sy, sz) along three axes. The color attributes of the Gaussian points are represented using RGB (red, green, and blue light imaging) or spherical harmonic coefficients. Specifically, the three-dimensional image data can be reconstructed into a Gaussian parameter quadruple of a three-dimensional Gaussian point cloud using a multilayer perceptron network. . This represents a multilayer perceptron network. The Gaussian parameters are normalized to ensure numerical stability. Gaussian parameterization enables Gaussian point clouds to accurately describe the geometry and appearance features of a 3D digital human model.

[0046] However, the facial geometric changes and expression details of a digital human model achieved with a fixed Gaussian point cloud distribution may be subpar. To better capture facial details and expression changes, in this embodiment of the invention, the distribution of the Gaussian point cloud can be adaptively and dynamically adjusted. Optionally, reconstructing the three-dimensional image data into a three-dimensional Gaussian point cloud includes: determining the Gaussian parameters of each Gaussian point based on the three-dimensional image data; determining the local curvature of each Gaussian point based on the Gaussian parameter set; and adjusting the point cloud density of the neighborhood of the Gaussian point based on the local curvature to obtain the three-dimensional Gaussian point cloud.

[0047] Local curvature can be determined by comparing the current Gaussian point with its neighboring Gaussian points. For example, the local curvature of the current Gaussian point is... ,in, Indicates the current Gaussian point The set of neighboring Gaussian points in the neighborhood of . and These are the current Gaussian points. and neighboring Gauss point The normal direction. Adjusting the point cloud density of the Gaussian point's neighborhood based on local curvature can be achieved by using local curvature as the adjustment weight. The adjustment weight and point cloud density adjustment can be positively or negatively correlated. For example, when the local curvature is larger (i.e., the adjustment weight is larger), Gaussian points can be added in the neighborhood to increase the point cloud density; conversely, when the local curvature is smaller (i.e., the adjustment weight is smaller), Gaussian points can be reduced in the neighborhood to decrease the point cloud density.

[0048] The inventors discovered in their actual research that regions with higher curvature usually correspond to more complex geometric structures, such as the corners of the eyes and mouth. In order to better capture facial details and expression changes, the point cloud density in the neighborhood of the Gaussian point can be increased as the local curvature increases, and the point cloud density in the neighborhood of the Gaussian point can be decreased as the local curvature decreases.

[0049] Optionally, the point cloud density of the Gaussian point neighborhood is adjusted according to the local curvature to obtain a three-dimensional Gaussian point cloud, including: determining a first curvature threshold and a second curvature threshold for point cloud density adjustment based on each local curvature; wherein the first curvature threshold is greater than the second curvature threshold; when the target local curvature is greater than the first curvature threshold, inserting a new Gaussian point in the neighborhood of the target Gaussian point corresponding to the target local curvature; when the target local curvature is less than the second curvature threshold, merging adjacent Gaussian points in the neighborhood of the target Gaussian point corresponding to the target local curvature.

[0050] For example, the first curvature threshold can be greater than the average of all local curvatures, and the second curvature threshold can be less than the average of all local curvatures. Alternatively, the first curvature threshold can be a value of local curvature greater than or equal to a first preset ratio, where the first preset ratio can be greater than or equal to 80%; the second curvature threshold can be a value of local curvature less than or equal to a second preset ratio, where the second preset ratio can be less than or equal to 20%.

[0051] Inserting new Gaussian points can be achieved through Gaussian point splitting operations, which ensures the capture of details; merging adjacent Gaussian points can be achieved through Gaussian point pruning operations, which can reduce redundant calculations.

[0052] After adjusting the point cloud density of the Gaussian point's neighborhood based on local curvature, the point cloud distribution can be further optimized, for example, by establishing an optimization objective function and performing gradient descent to optimize the point cloud distribution. For instance, the objective function could be... .in, This represents the original current Gaussian point position. The optimized current Gaussian point position. For the neighboring Gauss point, To balance the parameters, by controlling the proximity of the current Gaussian point to its neighboring Gaussian points, the optimized position of the current Gaussian point is constrained to not deviate too far from the original position. This makes the local structure of the Gaussian point cloud more compact, avoids excessive dispersion, and preserves the details of the digital human model.

[0053] By adaptively adjusting the point cloud density of Gaussian point clouds, the point density can be dynamically increased in areas where more detail is needed (such as when facial expressions change), while maintaining a lower density in relatively flat areas. This achieves a good balance between detail and efficiency, more accurately capturing and representing the facial details and expressions of digital humans, thus showcasing highly realistic visual effects from different perspectives.

[0054] Step 130: Establish the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy in the 3D human body modeling tool; and perform image rendering based on the binding relationship to obtain the digital human model.

[0055] By utilizing the various pose and shape parameters provided by 3D human body modeling tools, the binding relationship between 3D Gaussian point clouds and skeletal hierarchical structures can be established, which skeletal joints will affect the Gaussian points, as well as the strength ratio of each joint.

[0056] Optionally, a binding relationship is established between the 3D Gaussian point cloud and the skeletal hierarchy in the 3D human body modeling tool, including: determining the vertices corresponding to each Gaussian point in the 3D Gaussian point cloud in the 3D human body modeling tool, and determining the skin weights corresponding to the vertices; determining the transformation matrices of each joint step by step from the root node according to the skeletal hierarchy in the 3D human body modeling tool; and performing weighted calculations on the transformation matrices of each Gaussian point, the corresponding vertex, the skin weights, and the corresponding joints to obtain the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy.

[0057] For example, when inputting a new target pose, the binding relationship between each Gaussian point and the vertices of the skeletal hierarchy in the 3D human modeling tool can be determined, inheriting the skinning weight information of the vertices. Starting from the root node, the transformation matrix of each joint is determined level by level; the transformation matrix can include rotation, translation, and other transformation information of the joints. Then, based on the vertices bound to each Gaussian point and their corresponding skinning weights, the transformation matrix is ​​weighted and mixed. Specifically, the new position of the Gaussian point is calculated by weighting and combining its original position with the transformation matrices of all relevant joints.

[0058] In addition to position updates, the rotational properties of the Gaussian point are also updated simultaneously. By combining the initial rotation of the Gaussian point with the rotational transformations that primarily affect the joints, the orientation of the Gaussian point is ensured to naturally follow changes as the body performs various movements. For example, when the arm is raised, the Gaussian point on the arm not only moves to a new position, but its orientation also adjusts accordingly with the rotation of the bones.

[0059] By binding 3D Gaussian point clouds with skeletal hierarchical structures, it is compatible with the animation pipeline of 3D human body modeling tools, and can directly utilize the various pose and shape parameters provided by 3D human body modeling tools to achieve accurate and realistic modeling of digital human models.

[0060] After binding and modeling 3D Gaussian point clouds with skeletal hierarchical structures, image rendering can be performed to display the digital human model. For example, through an intelligent rendering engine, real-time rendering and dynamic lighting and shadow adjustments can be performed to achieve a realistic visual effect for the digital human model.

[0061] Step 140: Generate corresponding digital human voice data based on the response data during the digital human interaction process; combine and display the digital human model and digital human voice data to obtain a digital human business card.

[0062] The digital human's voice data can be pre-recorded based on response data. Furthermore, to broaden the application scope of digital human business cards, voice data can be generated in real time based on response data, allowing digital human business cards to be applied to different business scenarios, not limited to fixed businesses.

[0063] By combining digital human models with digital human voice data, digital human business cards can be created. For example, these business cards can be used in product or resource introductions across various industries such as financial services, telecommunications, retail, and education. This allows the introduction content to be deeply linked to the target user's image, making it more realistic, meeting diverse needs, effectively improving user reach efficiency and conversion rates, and providing organizations with a more competitive digital output tool.

[0064] The technical solution of this embodiment obtains the input image of the target user in the digital human business card generation process and preprocesses the input image to obtain image feature data; aligns and maps the image feature data in three-dimensional space coordinates to obtain three-dimensional image data; reconstructs the three-dimensional image data into a three-dimensional Gaussian point cloud; establishes a binding relationship between the three-dimensional Gaussian point cloud and the skeletal hierarchy structure in the three-dimensional human modeling tool; and renders the image according to the binding relationship to obtain a digital human model; generates corresponding digital human voice data based on the response data during the digital human interaction process; and combines and displays the digital human model and digital human voice data to obtain a digital human business card. This solves the problems of high cost and insufficient accuracy in digital human model generation. By aligning the input image in three-dimensional space coordinates and then reconstructing it into a three-dimensional Gaussian point cloud for digital human model building, single-image realistic digital human technology can be realized. This reduces the amount of data required for digital human production, lowers the cost of digital human production, and improves the realism of the digital human image, deeply binding it with the image of the target user.

[0065] Example 2

[0066] Figure 2 This is a flowchart of a method for generating a digital human business card according to Embodiment 2 of the present invention. This embodiment is a further addition and refinement of the above technical solution, and the technical solution in this embodiment can be combined with various optional solutions in one or more of the above embodiments. Figure 2 As shown, the method includes:

[0067] Step 210: Obtain the input image of the target user in the digital human business card generation process, and preprocess the input image to obtain image feature data.

[0068] Optionally, the input image is preprocessed to obtain image feature data, including: dividing the input image into multiple image blocks, extracting corresponding feature sub-vectors for each image block, and adding position encoding to the feature sub-vectors; performing layer normalization on the feature sub-vectors to obtain a set of feature tokens, and using the set of feature tokens as image feature data.

[0069] Step 220: Align and map the image feature data in three-dimensional space to obtain the three-dimensional image data.

[0070] Optionally, the image feature data is aligned and mapped in three-dimensional space to obtain three-dimensional image data, including: stitching the image feature data and UV tokens in the pre-learned texture coordinate system to obtain the stitching result; and using a multi-head self-attention mechanism and a feedforward network to perform feature fusion and enhancement on the stitching result to obtain three-dimensional image data.

[0071] Step 230: Determine the Gaussian parameter set for each Gaussian point based on the three-dimensional image data.

[0072] The Gaussian parameter set includes: the center position coordinates of the Gaussian point, the rotation angle, the rotation scale, and the color attribute.

[0073] Step 240: Determine the local curvature of each Gaussian point based on the Gaussian parameter set, and adjust the point cloud density of the neighborhood of the Gaussian point based on the local curvature to obtain a three-dimensional Gaussian point cloud.

[0074] Optionally, the point cloud density of the Gaussian point neighborhood is adjusted according to the local curvature to obtain a three-dimensional Gaussian point cloud, including: determining a first curvature threshold and a second curvature threshold for point cloud density adjustment based on each local curvature; wherein the first curvature threshold is greater than the second curvature threshold; when the target local curvature is greater than the first curvature threshold, inserting a new Gaussian point in the neighborhood of the target Gaussian point corresponding to the target local curvature; when the target local curvature is less than the second curvature threshold, merging adjacent Gaussian points in the neighborhood of the target Gaussian point corresponding to the target local curvature.

[0075] Step 250: Establish the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy in the 3D human body modeling tool; and perform image rendering based on the binding relationship to obtain the digital human model.

[0076] Optionally, a binding relationship is established between the 3D Gaussian point cloud and the skeletal hierarchy in the 3D human body modeling tool, including: determining the vertices corresponding to each Gaussian point in the 3D Gaussian point cloud in the 3D human body modeling tool, and determining the skin weights corresponding to the vertices; determining the transformation matrices of each joint step by step from the root node according to the skeletal hierarchy in the 3D human body modeling tool; and performing weighted calculations on the transformation matrices of each Gaussian point, the corresponding vertex, the skin weights, and the corresponding joints to obtain the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy.

[0077] Step 260: Generate corresponding digital human voice data based on the response data during the digital human interaction process.

[0078] Optionally, based on the response data during the digital human interaction process, corresponding digital human voice data is generated, including: collecting voice samples of the target user and extracting acoustic features from the voice samples; identifying the target user's emotional state based on the acoustic features and generating a style control vector for the target user based on the emotional state; performing speech synthesis based on the style control vector and voice samples to obtain the target user's voice waveform; and generating corresponding digital human voice data based on the voice waveform and response data.

[0079] Based on the digital human speech data generation method provided by this invention, style control vectors are incorporated into the speech data generation process, reducing the required number of speech samples from the target user. For example, only 3-6 speech samples are needed to generate highly similar digital human speech data. First, acoustic features can be extracted from the speech samples. Acoustic features can include timbre, pitch, fundamental frequency, speech rate, volume intensity, energy, and emotional characteristics. For example, acoustic features can be extracted from the speech samples using a neural network model. An emotional state set can be pre-constructed, and by combining acoustic features with the corresponding text data of the speech samples and an attention mechanism, the emotional state of the target user can be identified. . , This is a pre-constructed set of emotional states. Therefore, style control vectors are generated based on these emotional states. , , For data dimensions, This refers to the mapping relationship of style control vectors. For example, style control vectors can be obtained based on the emotional state of the target user through neural networks or function mappings.

[0080] When generating speech waveforms, style control vectors can be combined with speech samples, such as through formulas. Generate the speech waveform. Where T is the input text corresponding to the speech sample. This indicates a speech generation model that supports zero-sample speech cloning, maintaining speech similarity and naturalness, and quickly adapting to the target user's speaking characteristics under conditions of minimal sample size. Therefore, in practical applications, by combining response data from the digital human interaction process with speech waveforms, corresponding digital human speech data can be generated, achieving a more natural and realistic interaction effect.

[0081] Step 270: Combine and display the digital human model and digital human voice data to obtain a digital human business card.

[0082] For example, during digital human interaction, the voice command of the user asking the question can be obtained, the voice command can be converted into text through speech recognition, the corresponding response data can be obtained by processing the text, and a vivid and realistic digital human model and digital human voice data can be displayed through the target user's digital human business card. The answer can be given based on the target user's image and voice characteristics, thereby improving the dynamic interaction performance and emotional expression ability during information interaction, improving information transmission efficiency, establishing an emotional connection with the user, and improving customer conversion rate.

[0083] To maintain the synergy between voice and visuals in the digital human business card and achieve a more natural and realistic effect, optionally, the digital human model and digital human voice data can be combined for display to obtain the digital human business card. This includes: extracting key acoustic features from the digital human voice data and mapping the key acoustic features to corresponding mouth movement parameters in the digital human model; aligning the digital human voice data and mouth movement parameters in time; and playing the digital human voice data while displaying the digital human model to obtain the digital human business card.

[0084] Key acoustic features can include phonemes, syllable boundaries, and energy variations. Lip movement modeling can be used to map these key acoustic features to corresponding mouth movement parameters in the digital human model. Speech alignment algorithms, such as those based on dynamic time warping or attention mechanisms, can achieve temporal alignment between the digital human's speech data and mouth movement parameters, resulting in precise matching of speech and mouth shape. Simultaneously displaying the digital human model and playing its speech data enables high-quality speech cloning and a highly natural interactive experience, making the digital human more vivid and realistic during voice interaction, enhancing user immersion and satisfaction.

[0085] When using digital personal business cards, you can configure the interaction methods, such as voice interaction, action demonstration, product introduction, etc., and set the playback video duration, resolution and format of the digital personal business card, which is convenient for use in social media and conference scenarios.

[0086] The technical solution of this invention involves acquiring the input image of the target user in digital human business card generation and preprocessing the input image to obtain image feature data; aligning and mapping the image feature data in three-dimensional space coordinates to obtain three-dimensional image data; determining the Gaussian parameter set for each Gaussian point based on the three-dimensional image data; determining the local curvature of each Gaussian point based on the Gaussian parameter set and adjusting the point cloud density of the neighborhood of the Gaussian point based on the local curvature to obtain a three-dimensional Gaussian point cloud; establishing a binding relationship between the three-dimensional Gaussian point cloud and the skeletal hierarchy structure in a three-dimensional human modeling tool; rendering the image based on the binding relationship to obtain a digital human model; generating corresponding digital human voice data based on the response data during the digital human interaction process; and finally, using the digital human model... By combining digital human voice data with digital human business cards, the system solves the problems of high cost and insufficient accuracy in digital human model generation. It aligns the input image with 3D spatial coordinates and reconstructs it into a 3D Gaussian point cloud for digital human model creation, enabling single-image realistic digital human technology. This reduces the amount of data required for digital human production, lowers production costs, and enhances the realism of the digital human image, deeply binding it to the target user's image. Through low-sample voice cloning and temporal alignment of voice and lip movements, it provides realistic and vivid voice data. Based on the target user's real image and voice, it can generate highly realistic digital human business cards, improving user experience and information delivery efficiency through intelligent interaction and personalized services. Simultaneously, its ability to clone low-sample voice and generate digital human models from a single image significantly reduces production costs and increases production efficiency, making it widely applicable in multiple industries such as financial services, telecommunications, retail, and education. This effectively improves user reach efficiency and conversion rates, providing a more competitive digital marketing tool. Replacing traditional business cards with digital human business cards enables real-time digital interaction and can meet personalized consultation needs.

[0087] Example 3

[0088] Figure 3 This is a schematic diagram of a digital personal business card generation device according to Embodiment 3 of the present invention. Figure 3 As shown, the device includes: an image feature data generation module 310, a Gaussian point cloud generation module 320, a digital human model generation module 330, and a digital human business card generation module 340. Wherein:

[0089] The image feature data generation module 310 is used to acquire the input image of the target user in the digital human business card generation process and to preprocess the input image to obtain image feature data.

[0090] The Gaussian point cloud generation module 320 is used to perform three-dimensional spatial coordinate alignment and mapping on image feature data to obtain three-dimensional image data; and to reconstruct the three-dimensional image data into a three-dimensional Gaussian point cloud.

[0091] The digital human model generation module 330 is used to establish the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy structure in the 3D human body modeling tool; and to perform image rendering based on the binding relationship to obtain the digital human model.

[0092] The digital human business card generation module 340 is used to generate corresponding digital human voice data based on the response data during the digital human interaction process; the digital human model and digital human voice data are combined and displayed to obtain the digital human business card.

[0093] Optional, the Gaussian point cloud generation module 320 includes:

[0094] The Gaussian parameter set determination unit is used to determine the Gaussian parameter set of each Gaussian point based on the three-dimensional data of the image. The Gaussian parameter set includes: the center position coordinates of the Gaussian point, the rotation angle, the rotation scale, and the color attribute.

[0095] The Gaussian point cloud generation unit is used to determine the local curvature of each Gaussian point based on the Gaussian parameter set, and adjust the point cloud density of the neighborhood of the Gaussian point according to the local curvature to obtain a three-dimensional Gaussian point cloud.

[0096] Optional, Gaussian point cloud generation unit, including:

[0097] The curvature threshold determination subunit is used to determine the first curvature threshold and the second curvature threshold when adjusting the point cloud density based on each local curvature; wherein, the first curvature threshold is greater than the second curvature threshold;

[0098] The Gaussian point insertion sub-unit is used to insert a new Gaussian point in the neighborhood of the target Gaussian point corresponding to the target local curvature when the target local curvature is greater than the first curvature threshold.

[0099] The Gaussian point merging sub-unit is used to merge adjacent Gaussian points in the neighborhood of the target Gaussian point corresponding to the target local curvature when the target local curvature is less than the second curvature threshold.

[0100] Optional, the digital human business card generation module 340 includes:

[0101] The acoustic feature extraction unit is used to collect speech samples from the target user and extract acoustic features from the speech samples.

[0102] The style control vector generation unit is used to identify the target user's emotional state based on acoustic features and generate the target user's style control vector based on the emotional state.

[0103] The speech waveform determination unit is used to perform speech synthesis based on style control vectors and speech samples to obtain the speech waveform of the target user.

[0104] The digital human voice data generation unit is used to generate corresponding digital human voice data based on the voice waveform and response data.

[0105] Optional, the digital human business card generation module 340 includes:

[0106] The mouth movement parameter determination unit is used to extract key acoustic features from digital human speech data and map the key acoustic features to the corresponding mouth movement parameters in the digital human model.

[0107] The timing alignment unit is used to align the digital human's voice data with the mouth movement parameters in a timely manner, and play the digital human's voice data while displaying the digital human model to obtain the digital human's business card.

[0108] Optional, the digital human model generation module 330 includes:

[0109] The vertex determination unit is used to determine the vertices corresponding to each Gaussian point in the 3D Gaussian point cloud in the 3D human body modeling tool, and to determine the skin weights corresponding to the vertices.

[0110] The transformation matrix determination unit is used to determine the transformation matrix of each joint step by step, starting from the root node, based on the skeletal hierarchy in the 3D human body modeling tool.

[0111] The binding relationship determination unit is used to perform weighted calculations on each Gaussian point, the corresponding vertex, the skin weight, and the transformation matrix of the corresponding joint to obtain the binding relationship between the 3D Gaussian point cloud and the bone hierarchy structure.

[0112] Optionally, the image feature data generation module 310 includes:

[0113] The feature vector determination unit is used to divide the input image into multiple image blocks, extract the corresponding feature vectors for each image block, and add the position code to the feature vectors.

[0114] The image feature data determination unit is used to perform layer normalization on the feature sub-vectors to obtain a feature token set, which is then used as the image feature data.

[0115] Optional, the Gaussian point cloud generation module 320 includes:

[0116] The stitching result determination unit is used to stitch together the image feature data and the UV tokens in the pre-learned texture coordinate system to obtain the stitching result;

[0117] The image 3D data generation unit is used to perform feature fusion and enhancement on the stitching results using a multi-head self-attention mechanism and a feedforward network to obtain image 3D data.

[0118] The digital human business card generation device provided in this embodiment of the invention can execute the digital human business card generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0119] In the technical solutions of this invention, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information (such as the target user's input images, voice samples, acoustic features, emotional states, and style control vectors) all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0120] The information collected is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.

[0121] When providing digital human business card generation services, users are given a corresponding entry point to choose whether to agree to or reject the automated decision-making result; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0122] Example 4

[0123] Figure 4 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0124] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0125] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0126] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the method for generating a digital human business card.

[0127] In some embodiments, the method for generating a digital human business card may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the digital human business card generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the digital human business card generation method by any other suitable means (e.g., by means of firmware).

[0128] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0129] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0130] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0131] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0132] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0133] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0134] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for generating a digital personal business card, characterized in that, include: The input image of the target user in the digital human business card generation is obtained, and the input image is preprocessed to obtain image feature data; The image feature data is aligned and mapped in three-dimensional space to obtain three-dimensional image data; the three-dimensional image data is then reconstructed into a three-dimensional Gaussian point cloud. Establish the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy in the 3D human body modeling tool; and perform image rendering based on the binding relationship to obtain a digital human model. Based on the response data during the digital human interaction process, corresponding digital human voice data is generated; the digital human model and the digital human voice data are combined and displayed to obtain a digital human business card.

2. The method according to claim 1, characterized in that, Reconstructing the three-dimensional image data into a three-dimensional Gaussian point cloud includes: Based on the three-dimensional data of the image, the Gaussian parameter set of each Gaussian point is determined. The Gaussian parameter set includes: the center position coordinates of the Gaussian point, the rotation angle, the rotation scale, and the color attribute. The local curvature of each Gaussian point is determined based on the Gaussian parameter set, and the point cloud density of the neighborhood of the Gaussian point is adjusted according to the local curvature to obtain a three-dimensional Gaussian point cloud.

3. The method according to claim 2, characterized in that, The point cloud density of the Gaussian point neighborhood is adjusted according to the local curvature to obtain a three-dimensional Gaussian point cloud, including: A first curvature threshold and a second curvature threshold are determined based on the local curvatures described above when adjusting the point cloud density; wherein the first curvature threshold is greater than the second curvature threshold. When the target local curvature is greater than the first curvature threshold, a new Gaussian point is inserted in the neighborhood of the target Gaussian point corresponding to the target local curvature. When the target local curvature is less than the second curvature threshold, adjacent Gaussian points are merged in the neighborhood of the target Gaussian point corresponding to the target local curvature.

4. The method according to claim 1, characterized in that, Based on the response data during the digital human interaction process, corresponding digital human voice data is generated, including: Collect voice samples from the target user and extract the acoustic features from the voice samples; The target user's emotional state is identified based on the acoustic features, and a style control vector for the target user is generated based on the emotional state. Speech synthesis is performed based on the style control vector and the speech sample to obtain the speech waveform of the target user; Based on the speech waveform and the response data, corresponding digital human speech data is generated.

5. The method according to claim 1, characterized in that, By combining and displaying the digital human model and the digital human voice data, a digital human business card is obtained, including: Extract key acoustic features from the digital human's speech data and map the key acoustic features to corresponding mouth movement parameters in the digital human model; The digital human voice data is time-aligned with the lip-sync parameters, and the digital human voice data is played while the digital human model is displayed to obtain a digital human business card.

6. The method according to claim 1, characterized in that, Establishing the binding relationship between the 3D Gaussian point cloud and the skeletal hierarchy in the 3D human body modeling tool includes: In the 3D human body modeling tool, the vertices corresponding to each Gaussian point in the 3D Gaussian point cloud are determined, and the skinning weights corresponding to the vertices are determined. Based on the skeletal hierarchy in the 3D human body modeling tool, the transformation matrix of each joint is determined step by step, starting from the root node. The binding relationship between the three-dimensional Gaussian point cloud and the bone hierarchy is obtained by weighting and calculating the transformation matrix of each Gaussian point, the corresponding vertex, the skin weight, and the corresponding joint.

7. The method according to claim 1, characterized in that, The input image is preprocessed to obtain image feature data, including: The input image is divided into multiple image blocks, and the corresponding feature sub-vectors are extracted from each image block. The position code is then added to the feature sub-vectors. The feature subvectors are subjected to layer normalization to obtain a feature token set, which is then used as image feature data.

8. The method according to claim 1, characterized in that, The image feature data is aligned and mapped in three-dimensional space to obtain three-dimensional image data, including: The image feature data and the UV tokens in the pre-learned texture coordinate system are concatenated to obtain the concatenation result; A multi-head self-attention mechanism and a feedforward network are used to perform feature fusion and enhancement on the stitching result to obtain three-dimensional image data.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for generating a digital human business card according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method for generating a digital human business card according to any one of claims 1-8.

11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method for generating a digital human business card according to any one of claims 1-8.