VPA image generation method and system based on multi-modal AI large model

Through multimodal AI big model identification and generation technology, the flexibility and expressiveness problems of non-standard geometric virtual image generation are solved, and personalized intelligent virtual image generation is realized, suitable for the fields of education and entertainment.

CN120451354AInactive Publication Date: 2025-08-08FORYOU GENERAL ELECTRONICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510963960.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing intelligent virtual image generation technology is difficult to adapt to non-standard geometric forms, with low flexibility in bone binding and poor performance, which cannot meet the personalized needs of different users.

Method used

The VPA image generation method based on multimodal AI big model is adopted to identify body features through deep learning networks, combine diffusion models and conditional constraint networks to generate stylized images, and use graph neural networks to automatically generate bone structures to achieve automated bone binding and personalized adjustments.

Benefits of technology

It realizes flexible adaptive generation of any shape, improves the intelligence and interactivity of virtual images, and meets the personalized needs of the education and entertainment fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451354A_ABST
    Figure CN120451354A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent interaction, and provides a VPA image generation method and system based on a multi-modal AI large model, any body can be identified and a VPA image can be adaptively generated based on the AI large model, the head-body ratio is forced to be constrained according to a conditional constraint network, and the scaling and displacement of parts of the input body are dynamically adjusted, so that automatic modeling is realized; the intelligent skeleton binding and self-learning optimization technology is combined, self-learning adapts to the rare body, the hierarchical skeleton is automatically generated based on the body structure, a preset template is not needed, the flexibility is high, and the expressive force is good; through a closed loop of geometric understanding, stylized generation, intelligent binding and continuous personalized optimization, the intelligence and interactivity of VPA image generation are improved, the image is continuously optimized based on user feedback, and then the intelligent virtual image demonstration requirement in the education and entertainment field is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent interaction technology, and in particular to a method and system for generating a VPA image based on a multimodal AI large model. Background Art

[0002] With the widespread application of large AI models in intelligent interaction and personalized services, VPAs (Virtual Personal Assistants) have become a key player in human-computer interaction, finding widespread application in areas such as smart cockpits, smart homes, social platforms, and gaming and entertainment. Current VPA image generation relies primarily on preset templates or manual design, making it difficult to adapt to the personalized needs of different users and lacking the ability to adapt to different geometric shapes.

[0003] Existing technologies for intelligent avatar generation typically focus on realistic or cartoon-like characters with fixed proportions. However, adaptive and cute design for non-standard geometric shapes (such as ellipses, triangles, and cones) remains challenging. Rigging is a crucial step in VPA modeling, connecting a 3D model to a series of virtual skeletons that can be manipulated and animated to achieve dynamic effects for the character or object. Current VPA rigging often requires manual adjustment, resulting in high development costs and a lack of flexibility. Furthermore, custom skeleton hierarchies are required for non-standard body shapes, potentially disrupting common rigging templates and affecting the model's animation expressiveness and naturalness. Summary of the Invention

[0004] The present invention provides a VPA image generation method and system based on a multimodal AI large model, which solves the technical problems in existing intelligent virtual image generation technology that rely on preset templates or manual design, cannot adapt to non-standard modeling designs, have low flexibility and poor expressiveness when performing bone binding, and cannot meet the personalized needs of different users.

[0005] To solve the above technical problems, the present invention provides a VPA image generation method based on a multimodal AI large model, comprising: Obtain input shape and perform classification and recognition to obtain shape feature parameters; An initial VPA image is constructed based on the diffusion model with the body feature parameters as conditions, and the head-to-body ratio is constrained by a conditional constraint network to achieve style transfer to generate a stylized VPA image; Binding the character geometry of the VPA image to the skeleton; The VPA image is personalized based on the user interaction data.

[0006] This basic solution is based on a large AI model, which can recognize any shape and adaptively generate a VPA image. It dynamically adjusts the component scaling and displacement of the input shape according to the head-to-body ratio enforced by the conditional constraint network to achieve automatic modeling. Combining intelligent bone binding and self-learning optimization technology, it self-learns to adapt to rare shapes and automatically generates hierarchical skeletons based on the shape structure without the need for preset templates, with high flexibility and good expressiveness. Through the closed loop of geometric understanding → stylized generation → intelligent binding → continuous personalized optimization, the intelligence and interactivity of VPA image generation are improved, and the image is continuously optimized based on user feedback, thereby meeting the needs of intelligent virtual image demonstration in the fields of education and entertainment.

[0007] In a further embodiment, obtaining an input body and performing classification and recognition to obtain body feature parameters includes: Acquire an input shape, identify and classify the input shape based on a deep learning network, and determine the shape category of the input shape; Perform feature extraction on the corresponding input shape according to the shape category to obtain a corresponding geometric feature vector; Using a deformable convolutional layer to identify key points of the geometric feature vector to generate a feature tensor; The corresponding shape feature parameters are read from the feature tensor and normalized.

[0008] The hybrid architecture of the deep learning network (CNN+Transformer) in this solution can simultaneously capture local details and global context, improve classification accuracy, effectively identify irregular or dynamic shapes, and automatically classify user-input shapes to achieve flexible and personalized image generation, significantly improving the naturalness and creativity of human-computer interaction; through feature extraction, the geometric feature vector describing the input shape is obtained, and it provides basic data for subsequent key point identification as global information; the deformable convolutional layer is used to identify key points of the geometric feature vector, and the resulting feature tensor provides local information that can be used to describe the body structure in detail, ensuring the accuracy of subsequent generation and skeletal binding.

[0009] In a further embodiment, a VPA image is constructed based on a diffusion model with the body feature parameters as a condition, specifically: a virtual image generator G is constructed based on the diffusion model, and an initial VPA image is generated with the body feature parameters as a condition; wherein the diffusion model is as follows;

[0010] Where, 、 Represents the time step and the state of the image or avatar when using the device; Represents the time step The noise attenuation strength, is the cumulative decay value of all time steps; Indicates the intensity of noise added in each time step; Indicates that at the current time step and Noise prediction under state; represents the standard deviation of the noise at each time step; represents noise sampled from a standard normal distribution.

[0011] This scheme builds a virtual image generator G based on the diffusion model, and generates the initial VPA image based on the shape feature parameters; the diffusion model generates the initial VPA image from the noise image through multiple time steps of iteration. At the beginning, the denoising is gradually carried out through the reverse process to generate a clear target virtual image ,The diffusion model generates high-fidelity images through progressive denoising, ,which can delicately depict the texture, light and shadow, and geometric details of the ,virtual image.

[0012] In a further embodiment, the head-to-body ratio is constrained according to a conditional constraint network to implement style transfer to generate a stylized VPA image, including: According to the body feature parameters, adjusting the head-to-body ratio of the VPA image through a conditional constraint network; Based on a generative adversarial network or a diffusion model, the VPA image is mapped to a pre-defined design space to generate a VPA image that meets the preset target parameters; The deformation algorithm is used to optimize and adjust the geometric proportions of the VPA image.

[0013] This solution combines forced head-to-body ratio constraints with style transfer to ensure that the output VPA image strictly conforms to preset target parameters (and preset style specifications, such as the standard head-to-body ratio). Based on a generative adversarial network or diffusion model, the VPA image is mapped to a pre-set design space (such as a Q-version design space), separating body parameters (head-to-body ratio, limb length) from style parameters (color, texture), supporting seamless style switching within a fixed body structure. This solution can not only meet the precision requirements of professional fields, but also serve the personalized creative needs of general users.

[0014] In a further embodiment, binding the character geometry of the VPA image to a skeleton comprises: The motion propagation matrix of each vertex in the VPA image is calculated based on a graph neural network, and the skeleton structure is automatically generated using an automatic skeleton matching algorithm. Predicting the distribution of bone nodes in the skeletal structure through a hybrid density network, and adaptively adjusting the positions of the bone points according to the input geometric features and shape data; The VPA image is skinned and the influence weight of each bone node on the surface vertex is calculated to ensure natural deformation of non-rigid binding.

[0015] This solution uses a graph neural network to calculate the motion propagation matrix for each vertex in the aforementioned VPA avatar and automatically generates the skeletal structure using an automatic bone matching algorithm. The motion of each skeletal node is propagated based on the relationship between adjacent nodes, optimizing the skeletal binding effect and achieving natural dynamic propagation of the skeleton. Based on the training of the motion propagation matrix, the dynamic behavior of the skeleton can be adjusted according to the structural and motion relationships between nodes, achieving accurate movement transmission, harmonious body proportions, and ensuring the natural and smooth movement of the avatar. By using traditional skinning binding methods and dynamic weight distribution technology, the influence weight of each skeletal node on the surface vertex is calculated to ensure the natural and smooth deformation of the vertex during the avatar animation process.

[0016] In a further embodiment, the prediction formula for the skeletal node distribution is as follows:

[0017] Where, It is the input data, including geometric features and shape data; Represents a given input Under the condition of , the distribution probability of each bone node of the predicted VPA image; Represents the number of mixed Gaussian distributions; Represents the weight of the Gaussian distribution at each position of each bone node of the VPA image, the weight of all Gaussian distributions The sum is 1; The probability density function representing the distribution of each skeletal node of the VPA image; is the mean of the current Gaussian distribution; is the covariance matrix, which represents the strength of relationships or interdependence between nodes.

[0018] This solution uses mixed Gaussian distribution (GMM) modeling to simultaneously predict multiple possible positions of skeletal nodes (such as the number of Gaussian distributions of the positions of each skeletal node in the VPA image), rather than forcing the output of a single fixed value. By modeling the distribution of skeletal nodes through MDN, the virtual image system can more realistically reflect the inherent uncertainty of biological movement.

[0019] The present invention also provides a VPA image generation system based on a multimodal AI large model, which is used to implement the VPA image generation method based on a multimodal AI large model as described above, comprising: The geometric feature analysis module is used to obtain the input shape and classify and identify it to obtain shape feature parameters; An image generation module connected to the geometric feature analysis module is used to construct an initial VPA image based on a diffusion model with the body feature parameters as conditions, and to implement style transfer by forcibly constraining the head-to-body ratio according to a conditional constraint network to generate a stylized VPA image; A skeleton dynamic binding module connected to the image generation module, used to bind the character geometry of the VPA image to the skeleton; The personalized adjustment module connected to the skeletal dynamics binding module is used to perform personalized adjustment on the VPA image according to user interaction data.

[0020] Among them, the personalized adjustment module can continuously adjust the characteristics of the VPA image according to user needs, making the VPA more personalized and intelligent.

[0021] In a further embodiment, the geometric feature analysis module includes a shape classification unit, a feature vector recognition unit, a key point recognition unit and a parameter output unit connected in sequence; The shape classification unit is used to obtain an input shape, identify and classify the input shape based on a deep learning network, and determine the shape category of the input shape; The feature vector recognition unit is used to extract features of the corresponding input shape according to the shape category through the AI large model to obtain the corresponding geometric feature vector; The key point recognition unit is used to perform key point recognition on the geometric feature vector using a deformable convolution layer to generate a feature tensor; The parameter output unit is used to read the corresponding shape feature parameters from the feature tensor and perform normalization processing.

[0022] In a further embodiment, the image generation module includes a generator construction unit, a head-to-body ratio constraint unit, and a model output unit connected in sequence; The generator construction unit is used to construct a virtual image generator G based on the diffusion model, and generate an initial VPA image based on the body feature parameters; The head-to-body ratio constraint unit is used to adjust the head-to-body ratio of the VPA image based on the body feature parameters through a conditional constraint network; and based on a generative adversarial network or a diffusion model, map the VPA image to a pre-set design space to generate a VPA image that meets the preset target parameters; and is also used to optimize and adjust the body geometric proportions of the VPA image using a deformation algorithm.

[0023] The model output unit is used to output the generated stylized VPA image.

[0024] In a further embodiment, the skeleton dynamics binding module comprises a GNN unit, a skeleton node distribution unit and a weight distribution unit connected in sequence; The GNN unit is used to calculate the motion propagation matrix of each vertex in the VPA image based on a graph neural network, and automatically generate a skeleton structure using an automatic skeleton matching algorithm; The skeleton node distribution unit is used to predict the skeleton node distribution in the skeleton structure through a hybrid density network, and adaptively adjust the positions of the skeleton points according to the input geometric features and body data; The weight distribution unit is used to perform skinning on the VPA image and calculate the influence weight of each bone node on the surface vertex to ensure natural deformation of non-rigid binding.

[0025] A technology that integrates multiple modal data such as text, images, voice, and actions, and uses deep learning models to automatically generate or drive virtual characters.

[0026] The GNN unit of this solution uses a graph neural network (GNN) to learn the motion propagation matrix between skeletal nodes and generate a skeletal structure that conforms to natural movement based on this matrix. The connection relationship and motion pattern of each skeletal node are automatically calculated to ensure that the generated skeletal structure can adapt to different body characteristics. The skeletal node distribution unit adaptively adjusts the position of the skeletal points based on the input geometric features and body data to ensure the naturalness and stability of the skeletal structure in different postures. The weight distribution unit uses skinning to enable each skeletal node to naturally influence the surrounding mesh vertices during movement, achieving a natural flow of the skeleton during deformation, avoiding unreasonable distortion or stiffness, and allowing the VPA image to be animated smoothly. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a workflow diagram of a VPA image generation method based on a multimodal AI large model provided by an embodiment of the present invention; Figure 2 This is a system framework diagram of a VPA image generation system based on a multimodal AI large model provided by an embodiment of the present invention; Among them: geometric feature analysis module 1, shape classification unit 11, feature vector recognition unit 12, key point recognition unit 13, parameter output unit 14; image generation module 2, generator construction unit 21, head-to-body ratio constraint unit 22, model output unit 23; bone dynamic binding module 3, GNN unit 31, bone node distribution unit 32, weight allocation unit 33; personalized adjustment module 4. DETAILED DESCRIPTION

[0028] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of protection of the present invention. Many changes may be made to the present invention without departing from the spirit and scope of the present invention.

[0029] Example 1 The embodiment of the present invention provides a VPA image generation method based on a multimodal AI large model, such as Figure 1 As shown, in this embodiment, steps S1 to S4 are included: S1. Obtain input shape and perform classification and recognition to obtain shape feature parameters, including: S11, obtaining an input shape, identifying and classifying the input shape based on a deep learning network, and determining a shape category of the input shape; S12. Based on a large AI model (e.g., PointNet++), perform feature extraction on the corresponding input shape according to the shape category to obtain corresponding geometric feature vectors, including at least one or more geometric feature vectors of vertex curvature, axis of symmetry, and center of mass; Among them, the geometric feature vector provides global information and basic data for subsequent key point recognition.

[0030] S13. Use a deformable convolution layer to identify key points of the geometric feature vector to generate a feature tensor, which includes at least one or more feature tensors of vertex curvature, symmetry axis, and center of mass.

[0031] Among them, the feature tensor provides local information for finely describing the body structure, ensuring the accuracy of subsequent generation and bone binding.

[0032] S14. Read the corresponding shape feature parameters from the feature tensor and perform normalization processing.

[0033] Among them, the deep learning network includes a hybrid architecture consisting of "CNN+Transformer"; shape categories include but are not limited to ellipses, triangles, and cones.

[0034] The vertex curvature of a shape is a parameter that describes the degree of curvature of the surface. Through the vertex curvature, the model can identify whether the shape is smooth, round, or has a more complex geometry.

[0035] The axis of symmetry is used to describe the symmetry of a shape and helps the model generate more aesthetically pleasing symmetrical images.

[0036] The centroid is the geometric center of a shape and is used to determine the shape's overall balance and posture.

[0037] These normalized body feature parameters are then used in the subsequent character generation module 2, playing a particularly important role in adjusting the "2-head-to-body" ratio and cuteness. The normalized parameters ensure that the generated character meets specific geometric and aesthetic requirements.

[0038] The hybrid architecture of the deep learning network (CNN+Transformer) in this solution can simultaneously capture local details and global context, improve classification accuracy, effectively identify irregular or dynamic shapes, and automatically classify user-input shapes to achieve flexible and personalized image generation, significantly improving the naturalness and creativity of human-computer interaction; through feature extraction, the geometric feature vector describing the input shape is obtained, and it provides basic data for subsequent key point identification as global information; the deformable convolutional layer is used to identify key points of the geometric feature vector, and the resulting feature tensor provides local information that can be used to describe the body structure in detail, ensuring the accuracy of subsequent generation and skeletal binding.

[0039] S2. constructing an initial VPA image based on the diffusion model and the body feature parameters as conditions, and forcibly constraining the head-to-body ratio according to the conditional constraint network to achieve style transfer to generate a stylized VPA image, including: S21. Constructing a virtual image generator G based on a diffusion model, and generating an initial VPA image based on the normalized body feature parameters; wherein the diffusion model is as follows;

[0040] Where: 、 Represents the time step and the state of the image or avatar when using the device; and Both are coefficients for controlling noise attenuation; Represents the time step The noise attenuation strength, It is the cumulative attenuation value of all time steps, which affects the removal of noise during the generation process.

[0041] It controls the intensity of noise added in each time step, affecting the randomness of the generation process.

[0042] Indicates that at the current time step and The noise prediction in the state is the noise prediction term of the model. The model learns how to predict noise through training and uses it to remove the noise in generation.

[0043] Indicates the standard deviation of the noise at each time step, controlling the intensity of the noise.

[0044] Represents noise sampled from a standard normal distribution, which helps generate different avatar details.

[0045] This scheme builds a virtual image generator G based on the diffusion model, and generates the initial VPA image based on the shape feature parameters; the diffusion model generates the initial VPA image from the noise image through multiple time steps of iteration. At the beginning, the denoising is gradually carried out through the reverse process to generate a clear target virtual image ,The diffusion model generates high-fidelity images through progressive denoising, ,which can delicately depict the texture, light and shadow, and geometric details of the ,virtual image.

[0046] S22, adjusting the head-to-body ratio of the VPA image through a conditional constraint network based on the body feature parameters; For example, in children's education, the shape of a VPA is adjusted to a "2-head-to-body" ratio through a conditional constraint network.

[0047] S23. Mapping the VPA image to a pre-set design space (e.g., a Q-version design space) based on a generative adversarial network (GAN) or a diffusion model to generate a VPA image that meets pre-set target parameters (e.g., a large eye coefficient > 0.7, a curvature smoothness < 0.3); In this step, it is mainly used to generate facial and body details of the VPA image.

[0048] S24. Use a deformation algorithm (such as Morphing) to optimize and adjust the geometric proportions of the VPA image.

[0049] During the image generation process, the deformation algorithm adjusts the geometric proportions according to different body features (such as the shape of the head, torso, and limbs), making the transition of the VPA image smoother and more natural, thereby giving the stylized VPA image a natural anthropomorphic appearance.

[0050] This solution combines forced head-to-body ratio constraints with style transfer to ensure that the output VPA image strictly conforms to preset target parameters (and preset style specifications, such as the standard head-to-body ratio). Based on a generative adversarial network or diffusion model, the VPA image is mapped to a pre-set design space (such as a Q-version design space), separating body parameters (head-to-body ratio, limb length) from style parameters (color, texture), supporting seamless style switching within a fixed body structure. This solution can not only meet the precision requirements of professional fields, but also serve the personalized creative needs of general users.

[0051] S3. Binding the character geometry of the VPA image to the skeleton, including: S31. Calculate the motion propagation matrix of each vertex in the VPA image based on a graph neural network (GNN, used to learn the dependencies between vertices), and automatically generate the skeleton structure using an automatic skeleton matching algorithm; Among them, the motion propagation matrix of the vertex is expressed as , where N is the number of nodes in the skeleton network. The element Wg[i][j] in the matrix represents the connection strength between node i and node j.

[0052] Through the graph neural network (GNN), the system can learn the motion propagation matrix between skeletal nodes , and generates a bone structure that conforms to natural movement based on this matrix.

[0053] For example, in the rigging process of a VPA character, let's assume that node i represents the elbow, node j represents the shoulder, and node k represents the wrist. The GNN will learn how to appropriately adjust the movements of shoulder node j and wrist node k when the movement of elbow node i changes, thereby achieving natural arm movements.

[0054] S32, predicting the distribution of bone nodes in the skeletal structure through a hybrid density network, and adaptively adjusting the positions of the bone points according to the input geometric features and body data; In this embodiment, the prediction formula for the skeleton node distribution is as follows:

[0055] Where, is input data, usually representing certain features or parameters of a shape, and in this embodiment includes geometric features and shape data; Represents a given input Under the condition of , the distribution probability of each bone node of the predicted VPA image; Indicates the number of mixed Gaussian distributions, that is, the number of Gaussian distributions used to predict the positions of each bone node in the VPA image. In this case, You can determine the complexity of the distribution of each bone node.

[0056] The weight of the Gaussian distribution of each position of each bone node of the VPA image indicates the importance of the Gaussian distribution in the overall prediction. Among them, the weight of all Gaussian distributions is The sum is 1; It is a Gaussian distribution, which represents the probability density function of the distribution of each bone node of the VPA image; is the mean of the current Gaussian distribution (i.e. the position of the bone node); is the covariance matrix, which represents the strength of relationships or interdependence between nodes.

[0057] Through these parameters, the probability of bone nodes at different positions can be predicted.

[0058] For example, if the upper body skeleton model of a VPA character includes three key nodes: head, shoulders, and arms, during the rigging process, the system needs to predict the optimal position of the shoulder node. This can be done through MDN. The specific process is as follows: A) Input data : The input data x includes information such as the relative position of the shoulder and the head, the angle between the shoulder and the arm, etc.

[0059] B) Determine the Gaussian distribution of the target part based on the input data x : The MDN model learns multiple Gaussian distribution models to predict the possible location of the target part (such as the shoulder).

[0060] For example, suppose the system discovers through training that the optimal position of a target part (such as the shoulder) can vary within a certain range, then the probabilities of different shoulder positions are calculated based on multiple Gaussian distribution models.

[0061] C) Determine the weight of the Gaussian distribution .

[0062] The weights of different Gaussian distributions determine the reliability and importance of each model. For example, if the distribution of shoulders may be different in different postures, in some cases the shoulders may be closer to the neck, and in other cases they may be closer to the chest, the MDN will learn the weights of each distribution.

[0063] This solution uses mixed Gaussian distribution (GMM) modeling to simultaneously predict multiple possible positions of skeletal nodes (such as the number of Gaussian distributions of the positions of each skeletal node in the VPA image), rather than forcing the output of a single fixed value. By modeling the distribution of skeletal nodes through MDN, the virtual image system can more realistically reflect the inherent uncertainty of biological movement.

[0064] S33. Skin binding is performed on the VPA image, and the influence weight of each bone node on the surface vertex is calculated to ensure natural deformation of non-rigid binding.

[0065] Weight distribution is typically calculated based on factors such as the distance, angle, and velocity between the bones and vertices, and the weights are dynamically adjusted based on the motion of the skeletal nodes. To avoid unnatural deformations, the algorithm also includes optimization steps such as smoothing and weight normalization. This technique is a common method in the existing field and will not be described in detail in this embodiment.

[0066] This solution uses a graph neural network to calculate the motion propagation matrix for each vertex in the aforementioned VPA avatar and automatically generates the skeletal structure using an automatic bone matching algorithm. The motion of each skeletal node is propagated based on the relationship between adjacent nodes, optimizing the skeletal binding effect and achieving natural dynamic propagation of the skeleton. Based on the training of the motion propagation matrix, the dynamic behavior of the skeleton can be adjusted according to the structural and motion relationships between nodes, achieving accurate movement transmission, harmonious body proportions, and ensuring the natural and smooth movement of the avatar. By using traditional skinning binding methods and dynamic weight distribution technology, the influence weight of each skeletal node on the surface vertex is calculated to ensure the natural and smooth deformation of the vertex during the avatar animation process.

[0067] S4. Personalize the VPA image based on the user interaction data, for example: A) Obtain user interaction data and use reinforcement learning algorithms to adjust and optimize the VPA's appearance. This includes, but is not limited to, the VPA's geometric features, movement performance, and animation effects.

[0068] Alternatively, B) obtaining user interaction data, determining the user's behavior pattern and preference, and adjusting and optimizing the distribution and motion pattern of the skeleton nodes based on the behavior pattern and preference.

[0069] The embodiment of the present invention is based on an AI large model, which can recognize any shape and adaptively generate a VPA image. It dynamically adjusts the component scaling and displacement of the input shape according to the conditional constraint network's mandatory head-to-body ratio constraints to achieve automatic modeling. It combines intelligent bone binding and self-learning optimization technology to self-learn and adapt to rare shapes and automatically generate hierarchical skeletons based on the body structure without the need for preset templates, with high flexibility and good expressiveness. Through the closed loop of geometric understanding → stylized generation → intelligent binding → continuous personalized optimization, the intelligence and interactivity of VPA image generation are improved, and the image is continuously optimized based on user feedback, thereby meeting the needs of intelligent virtual image demonstration in the fields of education and entertainment.

[0070] Example 2 The embodiment of the present invention further provides a VPA image generation system based on a multimodal AI large model, which is used to implement the VPA image generation method based on a multimodal AI large model as described above. Figure 2 ,include: Geometric feature analysis module 1, used to obtain input shapes and classify and identify them to obtain shape feature parameters; An image generation module 2 connected to the geometric feature analysis module 1 is used to construct an initial VPA image based on a diffusion model with the body feature parameters as a condition, and to implement style transfer by forcibly constraining the head-to-body ratio according to a conditional constraint network to generate a stylized VPA image; A skeleton dynamic binding module 3 connected to the image generation module 2, for binding the character geometry of the VPA image to the skeleton; The personalized adjustment module 4 connected to the skeleton dynamics binding module 3 is used to perform personalized adjustment on the VPA image according to user interaction data.

[0071] Among them, the personalized adjustment module 4 can continuously adjust the characteristics of the VPA image in combination with user needs, making the VPA more personalized and intelligent.

[0072] In this embodiment, the geometric feature analysis module 1 includes a shape classification unit 11, a feature vector recognition unit 12, a key point recognition unit 13 and a parameter output unit 14 connected in sequence; The shape classification unit 11 is used to obtain an input shape, identify and classify the input shape based on a deep learning network, and determine the shape category of the input shape; The feature vector recognition unit 12 is used to extract features of the corresponding input shape according to the shape category through the AI large model to obtain the corresponding geometric feature vector; The key point recognition unit 13 is used to perform key point recognition on the geometric feature vector using a deformable convolution layer to generate a feature tensor; The parameter output unit 14 is used to read the corresponding shape feature parameters from the feature tensor and perform normalization processing.

[0073] In this embodiment, the image generation module 2 includes a generator construction unit 21, a head-to-body ratio constraint unit 22, and a model output unit 23 connected in sequence; The generator construction unit 21 is used to construct a virtual image generator G based on the diffusion model, and generate an initial VPA image based on the body feature parameters; The head-to-body ratio constraint unit 22 is used to adjust the head-to-body ratio of the VPA image based on the body feature parameters through a conditional constraint network; and based on a generative adversarial network or a diffusion model, map the VPA image to a pre-set design space to generate a VPA image that meets the preset target parameters; and is also used to optimize and adjust the body geometric proportions of the VPA image using a deformation algorithm.

[0074] The model output unit 23 is used to output the generated stylized VPA image.

[0075] In this embodiment, the skeleton dynamics binding module 3 includes a GNN unit 31, a skeleton node distribution unit 32 and a weight distribution unit 33 connected in sequence; The GNN unit 31 is used to calculate the motion propagation matrix of each vertex in the VPA image based on a graph neural network, and automatically generate a skeleton structure using an automatic skeleton matching algorithm; The skeleton node distribution unit 32 is used to predict the skeleton node distribution in the skeleton structure through a hybrid density network, and adaptively adjust the positions of the skeleton points according to the input geometric features and body data; The weight distribution unit 33 is used to perform skinning on the VPA image and calculate the influence weight of each bone node on the surface vertex to ensure natural deformation of non-rigid binding.

[0076] A technology that integrates multiple modal data such as text, images, voice, and actions, and uses deep learning models to automatically generate or drive virtual characters.

[0077] This solution's GNN unit 31 uses a graph neural network (GNN) to learn the motion propagation matrix between skeletal nodes and, based on this matrix, generates a skeletal structure that conforms to natural movement. The connection relationships and motion patterns of each skeletal node are automatically calculated to ensure that the generated skeletal structure can adapt to different body characteristics. The skeletal node distribution unit 32 adaptively adjusts the positions of skeletal points based on the input geometric features and body data, ensuring the naturalness and stability of the skeletal structure in different postures. The weight distribution unit 33 uses skinning to ensure that each skeletal node naturally affects the surrounding mesh vertices during movement, achieving a natural flow of the skeleton during deformation, avoiding unreasonable distortion or rigidity, and enabling smooth animation of the VPA image.

[0078] The embodiment of the present invention is based on an AI big model to create a VPA image generation system based on a multimodal AI big model. Steps S1 to S4 in the above-mentioned embodiment 1 are implemented through the combination of various module functions. By recognizing arbitrary shapes and adaptively generating cute VPA smart driving partners (for example, in early childhood education, multiple geometric structures can be automatically added and combined into a Q-version anthropomorphic image after selection in the interactive interface), combined with intelligent bone binding and self-learning optimization technology, it can recognize arbitrary shapes (whether elliptical, triangular, cone, etc.) without manual modeling, and adaptively convert them into a standard "two-head body" cute image. At the same time, bone binding is automatically performed, thereby improving the intelligence and interactivity of VPA image generation, and continuously optimizing the image based on user feedback to enhance immersion.

[0079] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A VPA image generation method based on a multimodal AI large model, characterized in that: include: Obtain input shape and perform classification and recognition to obtain shape feature parameters; An initial VPA image is constructed based on the diffusion model with the body feature parameters as conditions, and the head-to-body ratio is constrained by a conditional constraint network to achieve style transfer to generate a stylized VPA image; Binding the character geometry of the VPA image to the skeleton; The VPA image is personalized based on the user interaction data.

2. The VPA image generation method based on a multimodal AI large model according to claim 1, characterized in that: Obtain input shapes and perform classification and recognition to obtain shape feature parameters, including: Obtaining an input shape, identifying and classifying the input shape based on a deep learning network, and determining a shape category of the input shape; Perform feature extraction on the corresponding input shape according to the shape category to obtain a corresponding geometric feature vector; Using a deformable convolutional layer to identify key points of the geometric feature vector to generate a feature tensor; The corresponding shape feature parameters are read from the feature tensor and normalized.

3. The VPA image generation method based on a multimodal AI large model according to claim 1, characterized in that: The VPA image is constructed based on the diffusion model and the body feature parameters are used as a condition, specifically: a virtual image generator G is constructed based on the diffusion model, and an initial VPA image is generated based on the body feature parameters; wherein the diffusion model is as follows; Where, 、 Represents the time step and the state of the image or avatar when using the device; Represents the time step The noise attenuation strength, is the cumulative decay value of all time steps; Indicates the intensity of noise added in each time step; Indicates that at the current time step and Noise prediction under state; represents the standard deviation of the noise at each time step; represents noise sampled from a standard normal distribution.

4. The VPA image generation method based on a multimodal AI large model according to claim 3, characterized in that: The head-to-body ratio is constrained by the conditional constraint network, and style transfer is achieved to generate a stylized VPA image, including: According to the body feature parameters, adjusting the head-to-body ratio of the VPA image through a conditional constraint network; Based on a generative adversarial network or a diffusion model, the VPA image is mapped to a pre-defined design space to generate a VPA image that meets the preset target parameters; The deformation algorithm is used to optimize and adjust the geometric proportions of the VPA image.

5. The VPA image generation method based on a multimodal AI large model according to claim 1, characterized in that: Bind the character geometry of the VPA image to the skeleton, including: The motion propagation matrix of each vertex in the VPA image is calculated based on a graph neural network, and the skeleton structure is automatically generated using an automatic skeleton matching algorithm. Predicting the distribution of bone nodes in the skeletal structure through a hybrid density network, and adaptively adjusting the positions of the bone points according to the input geometric features and shape data; The VPA image is skinned and the influence weight of each bone node on the surface vertex is calculated to ensure natural deformation of non-rigid binding.

6. The VPA image generation method based on a multimodal AI large model according to claim 5, characterized in that: The prediction formula for the skeleton node distribution is as follows: Where, It is the input data, including geometric features and shape data; Represents a given input Under the condition of , the distribution probability of each bone node of the predicted VPA image; Represents the number of mixed Gaussian distributions; Represents the weight of the Gaussian distribution at each position of each bone node of the VPA image, the weight of all Gaussian distributions The sum is 1; The probability density function representing the distribution of each skeletal node of the VPA image; is the mean of the current Gaussian distribution; is the covariance matrix, which represents the strength of relationships or interdependence between nodes.

7. A VPA image generation system based on a multimodal AI large model, used to implement a VPA image generation method based on a multimodal AI large model as described in any one of claims 1 to 6, characterized in that: include: The geometric feature analysis module is used to obtain the input shape and classify and identify it to obtain shape feature parameters; An image generation module connected to the geometric feature analysis module is used to construct an initial VPA image based on a diffusion model with the body feature parameters as conditions, and to implement style transfer by forcibly constraining the head-to-body ratio according to a conditional constraint network to generate a stylized VPA image; A skeleton dynamic binding module connected to the image generation module, used to bind the character geometry of the VPA image to the skeleton; The personalized adjustment module connected to the skeletal dynamics binding module is used to perform personalized adjustment on the VPA image according to user interaction data.

8. The VPA image generation system based on a multimodal AI large model according to claim 7, characterized in that: The geometric feature analysis module includes a shape classification unit, a feature vector recognition unit, a key point recognition unit and a parameter output unit connected in sequence; The shape classification unit is used to obtain an input shape, identify and classify the input shape based on a deep learning network, and determine the shape category of the input shape; The feature vector recognition unit is used to extract features of the corresponding input shape according to the shape category through the AI large model to obtain the corresponding geometric feature vector; The key point recognition unit is used to perform key point recognition on the geometric feature vector using a deformable convolution layer to generate a feature tensor; The parameter output unit is used to read the corresponding shape feature parameters from the feature tensor and perform normalization processing.

9. A VPA image generation system based on a multimodal AI large model as claimed in claim 8, characterized in that: The image generation module includes a generator construction unit, a head-to-body ratio constraint unit and a model output unit connected in sequence; The generator construction unit is used to construct a virtual image generator G based on the diffusion model, and generate an initial VPA image based on the body feature parameters; The head-to-body ratio constraint unit is configured to adjust the head-to-body ratio of the VPA image based on the body feature parameters through a conditional constraint network; and to map the VPA image to a pre-set design space based on a generative adversarial network or a diffusion model to generate a VPA image that meets preset target parameters; and is further configured to optimize and adjust the body geometric proportions of the VPA image using a deformation algorithm; The model output unit is used to output the generated stylized VPA image.

10. The VPA image generation system based on a multimodal AI large model according to claim 9, characterized in that: The skeleton dynamics binding module includes a GNN unit, a skeleton node distribution unit and a weight distribution unit connected in sequence; The GNN unit is used to calculate the motion propagation matrix of each vertex in the VPA image based on a graph neural network, and automatically generate a skeleton structure using an automatic skeleton matching algorithm; The skeleton node distribution unit is used to predict the skeleton node distribution in the skeleton structure through a hybrid density network, and adaptively adjust the positions of the skeleton points according to the input geometric features and body data; The weight distribution unit is used to perform skinning on the VPA image and calculate the influence weight of each bone node on the surface vertex to ensure natural deformation of non-rigid binding.