Skeletal binding method, computer program product, electronic device
Patent Information
- Application Number
- CN202610958852.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本公开提供一种骨骼绑定方法,以至少在一定程度上解决相关技术中无法对非人形虚拟角色进行骨骼绑定的问题
[0009]本公开实施例提供的一种骨骼绑定方法,获取虚拟对象模型的多视角图像以及虚拟对象模型的预设骨架规范;基于预设骨架规范,生成虚拟对象模型的目标骨骼的提示信息;将提示信息以及与目标骨骼对应的多视角图像输入至多模态大语言模型,得到目标骨骼的骨骼节点位置估计信息,基于骨骼节点位置估计信息得到目标骨骼的骨骼节点的三维位置;基于预设骨架规范,将骨骼节点的三维位置映射为骨架实例,根据骨架实例得到骨骼绑定资产。一方面,在获取到虚拟对象模型的预设骨架规范后,基于该预设骨架规范生成虚拟对象模型的目标骨骼的提示信息,将该提示信息以及虚拟对象模型的多视角图像输入至多模态大语言模型中,得到目标骨骼的骨骼节点位置估计信息,实现了无需针对龙、四足兽、机甲等物种收集标注数据,多模态大语言模型通过自然语言描述即可理解关节语义,显著降低数据成本与模型维护负担;再一方面,基于骨骼节点位置估计信息得到目标骨骼的骨骼节点的三维位置,基于预设骨架规范,将三维位置映射为骨架实例,并根据该骨骼实例得到骨骼绑定资产,实现了对虚拟对象模型骨骼的自动绑定,降低了人工骨骼标注与绑定依赖,提高了骨骼绑定效率。
Smart Images

Figure CN122820931A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a skeletal binding method, computer program products, and electronic devices. Background Technology
[0002] In game development, non-humanoid virtual characters must undergo skeletal rigging before they can play animations. While the cost of AI-generated 3D characters has decreased and the production capacity of non-humanoid virtual characters has increased, rigging remains a bottleneck.
[0003] In related technologies, 2D pose estimators rely on supervised training of human / humanoid joints, can only recognize human joints, and cannot generalize to non-humanoid topologies. They require the collection of labeled data for each species, which is costly and has limited coverage. While 3D autoregressive rigging methods support non-humanoids, their output is a free skeleton, which cannot meet the skeleton specifications of projects, making it difficult to reuse animation libraries and requiring extensive post-processing and manual mapping. Multimodal large language models can only output object-level bounding boxes, lack the ability to output fine-grained joint-level coordinates, and have limited localization accuracy, making it difficult to directly support high-quality rigging. Summary of the Invention
[0004] This disclosure provides a skeletal rigging method to at least partially solve the problem in related technologies that it is impossible to perform skeletal rigging on non-humanoid virtual characters.
[0005] According to a first aspect of this disclosure, a skeletal rigging method is provided, the method comprising: Obtain multi-view images of the virtual object model and the preset skeleton specifications of the virtual object model; Based on the preset skeleton specifications, generate prompt information for the target skeleton of the virtual object model; The prompt information and the multi-view image corresponding to the target skeleton are input into the multimodal big language model to obtain the bone node position estimation information of the target skeleton, and the three-dimensional position of the bone node of the target skeleton is obtained based on the bone node position estimation information. Based on the preset skeleton specification, the three-dimensional position of the bone node is mapped to the skeleton instance, and the skeleton-bound asset is obtained according to the skeleton instance.
[0006] According to a second aspect of this disclosure, a skeletal binding device is provided, the device comprising: The multi-view image acquisition module is used to acquire multi-view images of the virtual object model and the preset skeleton specifications of the virtual object model; The prompt information generation module is used to generate prompt information for the target skeleton of the virtual object model based on the preset skeleton specifications; The skeletal node position determination module is used to input the prompt information and the multi-view image corresponding to the target skeleton into the multimodal big language model to obtain the skeletal node position estimation information of the target skeleton, and obtain the three-dimensional position of the skeletal node of the target skeleton based on the skeletal node position estimation information. The asset binding output module is used to map the 3D position of bone nodes to skeleton instances based on preset skeleton specifications, and obtain the skeleton binding assets based on the skeleton instances.
[0007] According to a third aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method of the first aspect described above and possible implementations thereof.
[0008] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the method of the first aspect and possible implementations thereof by executing the executable instructions.
[0009] This disclosure provides a skeleton binding method that acquires multi-view images of a virtual object model and a preset skeleton specification of the virtual object model; generates prompt information for the target skeleton of the virtual object model based on the preset skeleton specification; inputs the prompt information and the multi-view images corresponding to the target skeleton into a multimodal large language model to obtain bone node position estimation information of the target skeleton; obtains the three-dimensional position of the bone nodes of the target skeleton based on the bone node position estimation information; maps the three-dimensional position of the bone nodes to skeleton instances based on the preset skeleton specification; and obtains the skeleton binding asset based on the skeleton instances. On the one hand, after obtaining the preset skeleton specification of the virtual object model, the system generates prompt information for the target skeleton of the virtual object model based on the preset skeleton specification. The prompt information and the multi-view images of the virtual object model are then input into the multimodal large language model to obtain the estimated position information of the bone nodes of the target skeleton. This achieves the goal of understanding joint semantics through natural language description without collecting annotation data for species such as dragons, quadrupeds, and mechs, significantly reducing data costs and model maintenance burden. On the other hand, the system obtains the three-dimensional position of the bone nodes of the target skeleton based on the bone node position estimation information. Based on the preset skeleton specification, the three-dimensional position is mapped to a skeleton instance, and the skeleton binding asset is obtained based on the skeleton instance. This achieves automatic binding of the skeleton of the virtual object model, reduces the dependence on manual skeleton annotation and binding, and improves the efficiency of skeleton binding. Attached Figure Description
[0010] Figure 1 A flowchart illustrating a skeletal binding method in this exemplary embodiment is shown; Figure 2This exemplary embodiment shows a method for generating prompt information about the target skeleton of a virtual object model based on a preset skeleton specification. Figure 3 This exemplary embodiment shows a method for inputting prompt information and multi-view images into a multimodal large language model to obtain bone node position estimation information of the target bone when the target bone is a visible bone. Figure 4 This exemplary embodiment shows a flowchart of a method for obtaining the three-dimensional position of the bone nodes of a target bone based on bone node position estimation information. Figure 5 This exemplary embodiment shows a method for evaluating the credibility of skeletal node position estimation information to obtain the target two-dimensional observation position in the skeletal node position estimation information. Figure 6 This example implementation shows a method flowchart for inputting prompt information and multi-view images into a multimodal large language model to obtain bone node position estimation information of the target bone when the bone is invisible. Figure 7 This example embodiment shows a flowchart of a method for obtaining the three-dimensional position of the bone nodes of the target bone based on bone node position estimation information. Figure 8 This example implementation shows a flowchart of a method for mapping the three-dimensional position of bone nodes to bone instances based on a preset skeleton specification, and obtaining bone-bound assets based on the skeleton instances. Figure 9 A block diagram of a skeletal binding device is shown in this exemplary embodiment; Figure 10 A schematic diagram of the structure of an electronic device in this exemplary embodiment is shown. Detailed Implementation
[0011] Exemplary embodiments of this disclosure will be described more fully below with reference to the accompanying drawings.
[0012] The accompanying drawings are schematic illustrations of this disclosure and are not necessarily drawn to scale. Some block diagrams shown in the drawings may be functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in hardware modules or integrated circuits, or in networks, processors, or microcontrollers. Implementations can be carried out in various forms and should not be construed as limited to the examples set forth herein. The features, structures, or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough description of embodiments of this disclosure. However, those skilled in the art will recognize that one or more specific details may be omitted when implementing the technical solutions of this disclosure, or other methods, components, apparatuses, steps, etc., may be used to replace one or more specific details.
[0013] In game development, non-humanoid virtual characters such as dragons, mythical beasts, and mechs must undergo skeletal rigging before they can play animations. While AI modeling can generate character meshes in batches, binding these non-humanoid meshes to project skeletal specifications still relies heavily on manual labor.
[0014] In related technologies, 2D pose estimators rely on supervised training of human / humanoid joints, can only recognize human joints, and cannot generalize to non-humanoid topologies. They require the collection of labeled data for each species, which is costly and has limited coverage. Although 3D autoregressive rigging methods support non-humanoids, their output is a free skeleton, which cannot meet the skeleton specifications of projects, making it difficult to reuse animation libraries and requiring a lot of post-processing and manual mapping. The localization granularity of multimodal large language models is object-level bounding boxes, which cannot achieve joint-level localization accuracy, and the localization accuracy is limited, making it difficult to directly support high-quality rigging. Furthermore, the two-dimensional coordinates output by single-view multimodal large language models have output biases such as point offset and false observations. Existing methods do not have cross-view reprojection error constraints and cannot use multi-view geometric consistency to filter unreliable observations. Furthermore, in skeleton specifications, attachment bones, dynamic bones, and subdivided spinal bones cannot be directly located visually, and existing methods do not have a completion mechanism based on semantic reasoning.
[0015] In related technologies, if multi-view multimodal large language models, joint localization, and triangulation are directly combined, the following unsolvable problems still exist: multi-view multimodal large language models have no joint-level output; existing joint localization reliability can only characterize the reliability of single-view point locations and lacks a cross-view geometric consistency verification mechanism; conventional triangulation can only recover spatial points from two-dimensional observations, which cannot solve the VLM illusion problem of two-dimensional observations themselves, nor can it guarantee that the output skeleton meets the project skeleton specifications.
[0016] In view of the above problems, an exemplary embodiment of this disclosure provides a skeletal rigging method. (See reference...) Figure 1As shown, the skeleton rigging method may include the following steps: Step S110: Obtain multi-view images of the virtual object model and the preset skeleton specifications of the virtual object model; Step S120: Based on the preset skeleton specification, generate prompt information for the target skeleton of the virtual object model; Step S130: Input the prompt information and the multi-view image corresponding to the target skeleton into the multimodal big language model to obtain the bone node position estimation information of the target skeleton, and obtain the three-dimensional position of the bone node of the target skeleton based on the bone node position estimation information. Step S140: Based on the preset skeleton specification, map the three-dimensional position of the bone node to a skeleton instance, and obtain the skeleton-bound asset based on the skeleton instance.
[0017] In the aforementioned skeleton binding method, on the one hand, after obtaining the preset skeleton specification of the virtual object model, the target skeleton of the virtual object model is generated based on the preset skeleton specification. The prompt information and the multi-view image of the virtual object model are input into the multimodal large language model to obtain the estimated position information of the bone nodes of the target skeleton. This achieves the goal of understanding the joint semantics through natural language description without collecting annotation data for species such as dragons, quadrupeds, and mechs, significantly reducing data costs and model maintenance burden. On the other hand, the three-dimensional position of the bone nodes of the target skeleton is obtained based on the estimated position information of the bone nodes. Based on the preset skeleton specification, the three-dimensional position is mapped to a skeleton instance, and the skeleton binding asset is obtained based on the skeleton instance. This achieves automatic binding of the skeleton of the virtual object model, reduces the dependence on manual skeleton annotation and binding, and improves the efficiency of skeleton binding.
[0018] The following will provide further explanation and description of steps S110-S140.
[0019] In step S110, multi-view images of the virtual object model and the preset skeleton specifications of the virtual object model are obtained.
[0020] The virtual object model can be a non-humanoid virtual object model, such as a dragon model, mythical beast model, mecha model, etc. This disclosure does not specifically limit the non-humanoid virtual object model. The virtual object model is rendered from multiple perspectives to obtain multi-view images, which may include multiple images of the virtual object model. The preset skeleton specification may include bone naming, parent-child hierarchy, functional semantics (primary motion bone / attachment bone / dynamic bone), symmetry relationships, etc., this disclosure does not specifically limit the preset skeleton specification.
[0021] In one implementation, acquiring multi-view images of a virtual object model includes: Get the set of preset rendering perspectives; The virtual object model is rendered according to each rendering perspective in the preset rendering perspective set to obtain multi-view images.
[0022] Specifically, a preset set of rendering perspectives is obtained, which includes multiple rendering perspectives, such as front, back, left, right, top, and bottom, etc., which are not specifically limited in this disclosure. After obtaining the rendering perspectives, the virtual object model is rendered in multiple perspectives using a renderer in the offline stage to obtain an image corresponding to each rendering perspective. When outputting the rendered image, auxiliary channels such as depth map, visibility buffer, and semantic segmentation can be optionally output, which are not specifically limited in this disclosure.
[0023] In one embodiment, the method further includes: When obtaining the rendering virtual object model, the camera projection matrix corresponding to each rendering viewpoint.
[0024] Specifically, when rendering virtual object models from multiple perspectives, the camera projection matrix of the virtual camera under each rendering perspective can also be output. .
[0025] In step S120, based on the preset skeleton specification, prompt information for the target skeleton of the virtual object model is generated.
[0026] After obtaining the preset skeleton specification, the preset skeleton specification can be parsed to generate natural language description information for each bone in the virtual object model.
[0027] In one implementation, reference Figure 2 As shown, the target skeleton includes visible and invisible bones. Based on the preset skeleton specification, the tooltip information for generating the target skeleton of the virtual object model includes: Step S210: Parse the preset skeleton specification to generate visual language prompts for visible bones; wherein, the visual language prompts include at least one of the following: bone name, hierarchical relationship, and visual description; Step S220: Parse the preset skeleton specification to obtain the hierarchical relationship and functional description of the invisible bones. Based on the hierarchical relationship and functional description, obtain the semantic reasoning prompt information of the invisible bones.
[0028] The following will further explain and illustrate steps S210 and S220. Specifically, the virtual object model includes two types of skeletons: visible skeletons and invisible skeletons. Visible skeletons are those whose position can be directly identified based on visual appearance features in the multi-view image corresponding to the virtual object model. Invisible skeletons are those whose position cannot be directly determined based on visual appearance in the multi-view image corresponding to the virtual object model. Invisible skeletons include at least one of hidden skeletons, control skeletons, attachment skeletons, auxiliary skeletons, and dynamic skeletons.
[0029] After obtaining the preset skeleton specification, for visible bones, the preset skeleton specification can be parsed to obtain the bone name, bone hierarchy, and visual description information of the corresponding bone nodes, and visual language prompts can be generated based on the parsing results. Among them, the visual description information refers to the appearance, shape, or positional features of the target bone's corresponding region in the virtual object model, etc., which are not specifically limited in this disclosure.
[0030] For invisible skeletons, the preset skeleton specification can be parsed to obtain the hierarchical relationship and functional description of the invisible skeletons. Semantic reasoning prompts can be generated based on the hierarchical relationship and functional description. The functional description is used to characterize the functional role of the invisible skeleton in the virtual object model, such as for weapon mounting, for wing dynamic control, for tail swing constraint, and for character animation inverse kinematics control.
[0031] For example, for the visible bone of the left wing root, parsing the preset skeleton specification yields the bone name LeftWingRoot; the hierarchical relationship is Spine → LeftWingRoot → LeftWingMid (LeftWingRoot is a child node of Spine and serves as the starting node of the left wing bone chain). The visual description is: located in the area where the left wing connects to the torso, typically at the turning point where the wing root connects to the back when spread. Based on this, the generated visual language prompt could be: Please locate the bone node LeftWingRoot. This bone is located between Spine and LeftWingMid, belongs to the starting bone of the left wing, and is typically located in the area where the left wing connects to the torso, representing the turning point where the wing root connects to the back. Please output the two-dimensional position of this bone node.
[0032] For invisible skeletons attached to other skeletons, parsing the preset skeleton specification yields the hierarchy: RightHand → WeaponSocket. The function is described as being used to attach weapons, typically located on the front right side. The generated semantic reasoning prompt could be: "Please infer the position of the skeleton node WeaponSocket. This skeleton is a child node of RightHand, used for weapon attachment, typically located in the front region of the right-hand skeleton, offset along the right-hand direction. Please infer the target position based on the right-hand position in the image."
[0033] In step S130, the prompt information and the multi-view image corresponding to the target skeleton are input into the multimodal big language model to obtain the bone node position estimation information of the target skeleton, and the three-dimensional position of the bone node of the target skeleton is obtained based on the bone node position estimation information.
[0034] After obtaining the cue information for visible and invisible bones, this cue information, along with the multi-view image corresponding to the target bone, can be input into a multimodal large language model. This model then yields the estimated bone node positions for the target bone. When the target bone is visible, the estimated bone node positions include two-dimensional estimates and corresponding confidence levels for each viewpoint. When the target bone is invisible, the estimated bone node positions include the semantic relative positional relationship or positional constraints between the invisible bone and associated bone nodes. After obtaining the estimated bone node positions, the three-dimensional positions of the target bone's bone nodes can be determined based on this information.
[0035] In one implementation, reference Figure 3 As shown, when the target skeleton is a visible skeleton, the prompt information and the multi-view image corresponding to the target skeleton are input into the multimodal large language model to obtain the estimated bone node position information of the target skeleton, including: Step S310: Input the visual language prompts and the multi-view images corresponding to the visible skeletons into the multimodal large language model; Step S320: Obtain the output of the multimodal large language model, perform structured parsing on the output to obtain the two-dimensional coordinate information of the skeletal nodes; Step S330: Convert the two-dimensional coordinate information into image pixel coordinates, verify the image pixel coordinates, and obtain the estimated position information of the skeletal nodes.
[0036] The following will further explain and illustrate steps S310-S330. Specifically, for visible skeletons, since the structural regions corresponding to the target skeletons can form obvious visual features in multi-view images, the two-dimensional positions of the skeleton nodes can be directly determined based on the visual content using a multimodal large language model. That is, the visual language prompts generated for the target skeletons and the multi-view images corresponding to the target skeletons can be input into the multimodal large language model. After obtaining the output results of the multimodal large language model, the output results are structured and parsed to obtain the two-dimensional coordinate information of the skeleton nodes included in the output results. After obtaining the two-dimensional coordinate information, the two-dimensional coordinate information is converted into image pixel coordinates. To avoid abnormal output results of the multimodal large language model affecting subsequent three-dimensional position recovery, the image pixel coordinates can also be verified. For example, it can be detected whether the coordinates are within the image boundary range and whether they meet the reasonable position region constraints corresponding to the target skeletons. Based on the verified image pixel coordinates, the estimated position information of the visible skeleton nodes of the virtual object model is obtained.
[0037] The estimated bone node positions of the visible bones include multiple two-dimensional positions of the visible bones and their corresponding confidence scores. The multimodal large language model understands and localizes each input viewpoint image separately, and outputs a specific two-dimensional coordinate for the visible bones in that viewpoint. Since the set of rendering viewpoints includes multiple rendering viewpoints, for a single visible bone, the multimodal large language model will output N two-dimensional coordinates, that is, obtain N bone node position estimates, where N is the number of rendering viewpoints.
[0038] In one implementation, reference Figure 4 As shown, the three-dimensional positions of the bone nodes of the target bone are obtained based on the bone node position estimation information, including: Step S410: Evaluate the credibility of the skeletal node position estimation information to obtain the target two-dimensional observation position in the skeletal node position estimation information; Step S420: Obtain the three-dimensional position based on the two-dimensional observation position of the target.
[0039] The following will further explain and illustrate steps S410 and S420. Specifically, since the skeletal node position estimation information output by the multimodal large language model may have false detection, offset, or hallucination problems, after obtaining the skeletal node position estimation information for each target skeleton, reliable observation results can be screened through a credibility assessment to obtain the target's two-dimensional observation position. After obtaining the target's two-dimensional observation position, the three-dimensional position of the skeletal node is obtained based on the target's two-dimensional observation position.
[0040] In one implementation, reference Figure 5As shown, the credibility of the skeletal node position estimation information is evaluated to obtain the target's two-dimensional observation position in the skeletal node position estimation information, including: Step S510: Obtain the confidence level of the multimodal large language model output corresponding to the skeletal node position estimation information; Step S520: Determine the geometric reliability of the bone node position estimation information based on the reprojection error of the bone node position estimation information; Step S530: Determine the joint confidence of the skeletal node position estimation information based on the confidence level and geometric confidence level, and filter the skeletal node position estimation information based on the joint confidence level to obtain the target two-dimensional observation position.
[0041] The following will further explain and illustrate steps S510-S530. Specifically, firstly, the conditional probability of the multimodal large language model when generating the current output token is obtained. This can be based on the geometric mean or minimum value of the probabilities of multiple tokens that make up the two-dimensional coordinates to determine the confidence level of the skeletal node position estimation information. When the multimodal large language model does not directly provide the probability of the coordinate tokens, the output text perplexity can be used, and the reciprocal of the output text perplexity can be used as the confidence level of the skeletal node position estimation information. Secondly, the reprojection error of the skeletal node position estimation information is calculated, and the geometric confidence level of the skeletal node position estimation information is determined based on this reprojection error. This reprojection error is used to measure whether the skeletal node position estimation information from multiple perspectives satisfies spatial geometric consistency. This includes: recovering temporary 3D points based on the skeletal node position estimation information from multiple perspectives, reprojecting the recovered temporary 3D points onto the images of each perspective to obtain the theoretical 2D projection position, comparing the theoretical 2D projection position with the skeletal node position estimation information, and obtaining the difference between the two. This difference is the reprojection error. The larger the reprojection error, the smaller the geometric confidence level. Finally, based on the confidence level and geometric confidence level, the joint confidence level of the skeletal node position estimation information is determined, which can be expressed as: ,in, From the perspective Lower skeletal nodes Confidence level, For reprojection error, Here, represents the temperature coefficient, and represents an empirical parameter. After obtaining the joint confidence level, the estimated skeletal node positions corresponding to joint confidence levels below a preset threshold can be filtered to obtain the target's two-dimensional observation position.
[0042] In one implementation, obtaining the three-dimensional position based on the target's two-dimensional observation position includes: Based on the target's two-dimensional observation position and the camera projection matrix corresponding to the multi-view images, determine the spatial geometric constraints corresponding to the bone nodes of the target skeleton; Based on the joint confidence level corresponding to the two-dimensional observation position of the target, a weighted geometric solution is performed on the spatial geometric constraints to obtain the three-dimensional position of the bone nodes of the target skeleton.
[0043] Specifically, the spatial geometric constraints of the target skeleton nodes can be constructed based on the target's two-dimensional observation position and the camera projection matrix. For the i-th rendering viewpoint, these constraints can be based on the two-dimensional observation position. With camera projection matrix Constructing constraints ,in, Let be the three-dimensional position to be solved. After obtaining the spatial geometric constraints, the weights corresponding to each viewpoint can be determined based on the joint confidence level. We construct a weighted least squares problem and solve it using singular value decomposition (SVD) to obtain the three-dimensional positions of the bone nodes of the target skeleton.
[0044] In one implementation, the number of valid rendering viewpoints after screening is counted. If the number of valid rendering viewpoints is lower than a preset threshold, the multimodal large language model is called again for joint localization by expanding the set of rendering viewpoints and / or adjusting the prompt information until the triangulation solution requirements are met.
[0045] Specifically, after completing the joint credibility assessment of the skeletal node position estimation information, the number of valid rendering viewpoints retained after screening is counted. Because virtual object models commonly exhibit issues such as limb overlap, wing occlusion, mecha armor coverage, and complex multi-segment skeletal chain occlusion, the multimodal large language model cannot achieve accurate joint localization in some viewpoints, and the corresponding 2D observations will be discarded due to low credibility. If the number of remaining valid viewpoints after screening is lower than a preset threshold, sufficient and geometrically reasonable 2D constraint equations cannot be provided for weighted analytical triangulation. Therefore, when the number of valid rendering viewpoints does not reach the preset threshold, the set of rendering viewpoints can be expanded and / or the prompt information adjusted. After expanding the viewpoints or adjusting the prompt information, the multimodal large language model is called again to perform 2D joint localization, and the credibility assessment and hallucination anomaly observation screening are performed again, and the number of valid viewpoints is counted again. This process is repeated until the number of valid viewpoints meets the preset threshold requirement. The preset threshold can be 2 or 3, and this disclosure does not specifically limit it.
[0046] In one implementation, for invisible bones, since they are usually located inside the virtual object model or occluded by external structures, it is difficult to form clear and identifiable visual features in multi-view images. Therefore, it is impossible to obtain highly reliable two-dimensional localization results directly through a multimodal large language model. (Reference) Figure 6As shown, when the skeleton is invisible, the prompt information and the multi-view image corresponding to the target skeleton are input into the multimodal large language model to obtain the estimated bone node position information of the target skeleton, including: Step S610: Input the semantic reasoning prompts and the multi-view images corresponding to the invisible skeletons into the multimodal big language model, and obtain the semantic relative positional relationship of the invisible skeletons relative to the associated skeleton nodes based on the multimodal big language model. Step S620: Obtain the estimated position information of the skeletal nodes based on the semantic relative positional relationship.
[0047] The following will further explain and illustrate steps S610 and S620. Specifically, after obtaining the semantic reasoning prompts for the invisible skeleton, these prompts and the multi-view images corresponding to the invisible skeleton can be input into a multimodal large language model to enable the model to perform semantic reasoning and obtain the semantic relative positional relationship between the invisible skeleton and its associated skeleton nodes. The associated skeleton node is a known spatial reference point that, in the preset skeleton specification, has a direct parent-child connection or explicit topological constraint relationship with the invisible skeleton, and whose three-dimensional coordinates have been recovered by visual positioning. After obtaining the semantic relative positional relationship between the invisible skeleton and its associated skeleton nodes, this semantic relative positional relationship is determined as the skeleton node position estimation information.
[0048] In one implementation, reference Figure 7 As shown, the three-dimensional positions of the bone nodes of the target bone are obtained based on the bone node position estimation information, including: Step S710: Obtain the 3D position of at least one associated bone node that has a topological association with the invisible bone, and generate position constraint information based on the semantic relative position relationship; Step S720: Determine the spatial positional relationship between the invisible bone and the associated bone node based on the positional constraint information corresponding to the invisible bone and the skeleton completion constraints included in the preset skeleton specification. Step S730: Based on the spatial positional relationship and the three-dimensional position of the associated bone nodes, determine the three-dimensional position of the bone nodes of the invisible bones.
[0049] The following will further explain and illustrate steps S710-S730. Specifically, when determining the three-dimensional coordinates of the skeletal nodes of the invisible skeleton, the three-dimensional position of at least one associated skeletal node with a topological association with the invisible skeleton can be obtained. This topological association refers to the connection relationships within the skeleton structure, including parent-child relationships, child-child relationships, and adjacent-node relationships. The associated skeletal node is a known spatial reference point in the preset skeleton specification that has a direct parent-child connection or explicit topological constraint relationship with the invisible skeleton, and whose three-dimensional coordinates have been recovered by visual positioning. Furthermore, positional constraint information can be generated based on the semantic relative positional relationships in the position estimation information of the skeletal nodes of the invisible skeleton. This positional constraint information essentially transforms the fuzzy natural language reasoning results into a structured constraint instruction with a clear reference point. For example, the structured result "outside the X direction" can be recorded as: {Reference point: associated skeletal node A, constraint direction: positive X-axis}. Secondly, based on the positional constraint information corresponding to the invisible bone and the skeleton completion constraints included in the preset skeleton specification, the spatial positional relationship between the invisible bone and its associated bone nodes is determined. This spatial positional relationship can be at least one of relative direction relationship, relative distance relationship, offset ratio, and hierarchical direction constraint, which is not specifically limited in this disclosure. Skeleton completion constraints are used to describe the structural rules of the invisible bone in the standard skeleton, and can include at least one of bone hierarchy constraint, bone direction constraint, bone length ratio constraint, parent-child connection constraint, topological connection constraint, and symmetry constraint, which is not specifically limited in this disclosure. After obtaining the spatial positional relationship, the three-dimensional position of the bone nodes of the invisible bone can be determined based on this spatial positional relationship and the three-dimensional position of the associated bone nodes.
[0050] Based on the above implementation method, even when invisible bones lack direct visual observation information, their reasonable three-dimensional positions that satisfy the skeletal topology can still be restored, thereby avoiding skeletal breakage, binding distortion, or animation abnormalities caused by missing internal bones, and improving the integrity of non-humanoid virtual object model skeleton restoration and the stability of binding results.
[0051] In step S140, based on the preset skeleton specification, the three-dimensional position of the bone node is mapped to a skeleton instance, and the skeleton binding asset is obtained according to the skeleton instance.
[0052] After obtaining the 3D position of the bone nodes of the target bone, the 3D position can be mapped to a skeleton instance based on the preset skeleton specification, and the bone binding asset can be obtained based on the skeleton instance.
[0053] In one implementation, reference Figure 8 As shown, based on a preset skeleton specification, the 3D positions of bone nodes are mapped to skeleton instances, and the skeleton-bound assets are obtained from the skeleton instances, including: Step S810: Based on the bone node topology, bone naming rules and bone hierarchy in the preset skeleton specification, perform a standard mapping of the three-dimensional position of the bone node to obtain a skeleton instance corresponding to the preset skeleton specification. Step S820: Determine the bone transformation information corresponding to the bone node based on the three-dimensional position of the bone node and the bone hierarchy relationship and bone direction constraints in the preset skeleton specification. Step S830: Based on the skeleton instance, establish the skinning association between the skeleton nodes and the mesh vertices of the virtual object model; Step S840: Based on the skeleton instance, skinning relationship and skeleton transformation information, obtain the bound asset.
[0054] The following will further explain and illustrate steps S810-S840. Specifically, the three-dimensional position of the bone nodes can describe the distribution structure of the skeleton in space, but relying solely on discrete spatial coordinates cannot form a standard skeleton structure that meets the requirements of the animation system. Therefore, after obtaining the three-dimensional coordinates, the three-dimensional position of the bone nodes can be mapped according to the bone node topology, bone naming rules, and bone hierarchy in the preset skeleton specification to obtain a skeleton instance that meets the preset skeleton specification. After obtaining the skeleton instance, the bone transformation information corresponding to the bone nodes can be determined according to the three-dimensional position of the bone nodes and the bone hierarchy and bone direction constraints in the preset skeleton specification. The bone transformation information is used to describe the posture state of the bones in three-dimensional space, including: bone translation information, bone rotation information, bone local coordinate system, bone direction information, etc. After obtaining the skeleton instance, the skinning relationship between the bone nodes and the mesh vertices of the virtual object model can be established based on the skeleton instance. The skinning relationship is used to describe the degree of influence of the bone nodes on the surface mesh vertices of the virtual object model. When determining skinning relationships, for each mesh vertex in the non-humanoid virtual object model, the distance from that mesh vertex to multiple bone segments can be calculated, and the influence weight of the bone node on the corresponding mesh vertex can be determined based on the distance. The bone segment can be a bone connection segment composed of a parent bone node and a child bone node. The distance can be geodesic distance or approximate distance. Geodesic distance describes the distance on the mesh surface path; approximate distance can use Euclidean distance or an approximate estimate based on spatial neighborhood, which is not specifically limited in this disclosure. After obtaining the distance, candidate weights corresponding to the bone nodes can be determined based on a distance decay model. To avoid increasing computational complexity or animation instability due to excessive bone influence on a single mesh vertex, the top-K bone nodes with the highest candidate weights can be retained as target influence weights. The distance decay model can be Gaussian decay, thermal diffusion, bone segment influence field, etc., which is not specifically limited in this disclosure. After obtaining the skeleton instance, skinning relationships, and bone transformation information, the skeleton instance, skinning relationships, and bone transformation information can be output as a binding asset. The format of the binding asset can be FBX or glTF, which is not specifically limited in this disclosure.
[0055] Exemplary embodiments of this disclosure also provide a skeletal binding device, with reference to Figure 9 As shown, it includes: The multi-view image acquisition module 910 is used to acquire multi-view images of the virtual object model and the preset skeleton specifications of the virtual object model; The prompt information generation module 920 is used to generate prompt information for the target skeleton of the virtual object model based on the preset skeleton specification; The skeletal node position determination module 930 is used to input the prompt information and the multi-view image corresponding to the target skeleton into the multimodal big language model to obtain the skeletal node position estimation information of the target skeleton, and obtain the three-dimensional position of the skeletal node of the target skeleton based on the skeletal node position estimation information. The asset binding output module 940 is used to map the three-dimensional position of the bone node to a skeleton instance based on a preset skeleton specification, and obtain the skeleton binding asset based on the skeleton instance.
[0056] In one embodiment, the multi-view image acquisition module includes: The rendering view set determination module is used to obtain the preset rendering view set; The multi-view image generation module is used to render the virtual object model according to each rendering view in the preset rendering view set to obtain multi-view images.
[0057] In one embodiment, the multi-view image acquisition module further includes: The camera projection matrix acquisition module is used to obtain the camera projection matrix corresponding to each rendering viewpoint when rendering the virtual object model.
[0058] In one embodiment, the target skeleton includes visible skeletons and invisible skeletons, and the prompt information generation module includes: The visual language prompt information generation module is used to parse the preset skeleton specification and generate visual language prompt information for visible bones; wherein, the visual language prompt information includes at least one of the following: bone name, hierarchical relationship, and visual description; The semantic reasoning prompt information generation module is used to parse the preset skeleton specification to obtain the hierarchical relationship and functional description of the invisible skeleton, and to obtain the semantic reasoning prompt information of the invisible skeleton based on the hierarchical relationship and functional description.
[0059] In one implementation, when the target skeleton is a visible skeleton, the skeleton node position determination module includes: The first input module is used to input visual language prompts and the multi-view images corresponding to visible skeletons into the multimodal large language model; The output result parsing module is used to obtain the output results of the multimodal large language model, perform structured parsing on the output results, and obtain the two-dimensional coordinate information of the skeletal nodes; The skeletal node position estimation information acquisition module is used to convert two-dimensional coordinate information into image pixel coordinates, verify the image pixel coordinates, and obtain skeletal node position estimation information.
[0060] In one embodiment, the skeletal node position determination module includes: The credibility assessment module is used to assess the credibility of the skeleton node position estimation information and obtain the target two-dimensional observation position in the skeleton node position estimation information. The three-dimensional position determination module is used to obtain the three-dimensional position of the target based on the two-dimensional observation position.
[0061] In one implementation, the credibility assessment module includes: The confidence acquisition module is used to obtain the confidence scores of the multimodal large language model output corresponding to the skeletal node position estimation information. The geometric reliability determination module is used to determine the geometric reliability of the bone node position estimation information based on the reprojection error of the bone node position estimation information. The target two-dimensional observation location acquisition module is used to determine the joint confidence of the skeletal node location estimation information based on confidence level and geometric confidence level, and to filter the skeletal node location estimation information based on the joint confidence level to obtain the target two-dimensional observation location.
[0062] In one embodiment, the three-dimensional position determination module includes: The spatial geometric constraint construction module is used to determine the spatial geometric constraints corresponding to the bone nodes of the target skeleton based on the target's two-dimensional observation position and the camera projection matrix corresponding to the multi-view images. The geometry solution module is used to perform weighted geometric solution of spatial geometric constraints based on the joint confidence level corresponding to the two-dimensional observation position of the target, so as to obtain the three-dimensional position of the bone nodes of the target skeleton.
[0063] In one implementation, when the skeleton is invisible, the skeleton node position determination module includes: The second input module is used to input semantic reasoning prompts and the multi-view images corresponding to the invisible skeletons into the multimodal big language model, and obtain the semantic relative positional relationship of the invisible skeletons relative to the associated skeleton nodes based on the multimodal big language model. The skeletal node position estimation information acquisition module is used to obtain skeletal node position estimation information based on semantic relative positional relationships.
[0064] In one embodiment, the skeletal node position determination module includes: The associated bone node position acquisition module is used to acquire the three-dimensional position of at least one associated bone node that has a topological association with an invisible bone, and generate position constraint information based on semantic relative positional relationship; The relative position relationship determination module is used to determine the spatial position relationship of the invisible bone relative to the associated bone node based on the position constraint information corresponding to the invisible bone and the skeleton completion constraints included in the preset skeleton specification. The 3D position determination module is used to determine the 3D position of the bone nodes of invisible bones based on spatial positional relationships and the 3D position of associated bone nodes.
[0065] In one implementation, the asset binding output module includes: The skeleton instance mapping module is used to perform standardized mapping of the three-dimensional position of the bone nodes according to the bone node topology, bone naming rules and bone hierarchy in the preset skeleton specification, so as to obtain the skeleton instance corresponding to the preset skeleton specification. The skeleton transformation information determination module is used to determine the skeleton transformation information corresponding to the skeleton node based on the three-dimensional position of the skeleton node and the skeleton hierarchy relationship and bone direction constraints in the preset skeleton specification. The skinning association determination module is used to establish skinning associations between skeleton nodes and mesh vertices of the virtual object model based on the skeleton instance; The bound asset output module is used to obtain bound assets based on skeleton instances, skinning relationships, and skeleton transformation information.
[0066] The specific details of each part of the above-mentioned device have been described in detail in the method section of the implementation plan. For any undisclosed details, please refer to the implementation plan of the method section, and therefore will not be repeated here.
[0067] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to exemplary embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0068] Furthermore, although the steps of the method in this invention are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.
[0069] Exemplary embodiments of this disclosure also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the above-described skeletal rigging method.
[0070] In one implementation, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium can be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid-state drive (SSD), etc. For example, the computer program product can be implemented as a non-volatile storage medium storing a computer program, such as read-only memory, NAND flash memory, etc.
[0071] In one implementation, the computer program product can be an intangible product containing a computer program. For example, the computer program product can be implemented as a virtual digital product, such as an executable file, installation package, or other digital file storing the computer program.
[0072] Computer program code can be written in one or more programming languages. Examples of programming languages include C, Java, and C++. Program code can execute entirely on the user's computing device, partially on the user's computing device, or as a standalone software package. It can also execute partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, such as a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via an internet connection provided by a mobile network operator).
[0073] Computer programs can be carried or transmitted via signals such as electricity, magnetism, light, electromagnetic fields, and infrared radiation. Electronic devices can convert signals carrying computer programs into digital signals, thereby running the computer programs. When a computer program runs on an electronic device, its code is used to cause the electronic device to execute (more specifically, the processor of the electronic device to execute) the method steps of various exemplary embodiments of this disclosure, such as the skeleton binding method described above, which includes the following steps: Step S110: Acquire multi-view images of a virtual object model and a preset skeleton specification of the virtual object model; Step S120: Generate prompt information for the target skeleton of the virtual object model based on the preset skeleton specification; Step S130: Input the prompt information and the multi-view images corresponding to the target skeleton into a multimodal large language model to obtain the bone node position estimation information of the target skeleton, and obtain the three-dimensional position of the bone nodes of the target skeleton based on the bone node position estimation information; Step S140: Map the three-dimensional position of the bone nodes to a skeleton instance based on the preset skeleton specification, and obtain the skeleton binding asset based on the skeleton instance.
[0074] The above method steps are implemented by a computer program. On the one hand, after obtaining the preset skeleton specification of the virtual object model, the program generates prompt information for the target skeleton of the virtual object model based on the preset skeleton specification. The prompt information and the multi-view images of the virtual object model are input into a multimodal large language model to obtain the estimated position information of the bone nodes of the target skeleton. This achieves the goal of understanding joint semantics through natural language description without collecting annotation data for species such as dragons, quadrupeds, and mechs, significantly reducing data costs and model maintenance burden. On the other hand, the program obtains the three-dimensional position of the bone nodes of the target skeleton based on the estimated position information of the bone nodes. Based on the preset skeleton specification, the three-dimensional position is mapped to a skeleton instance, and the skeleton binding asset is obtained based on the skeleton instance. This achieves automatic binding of bones of non-humanoid virtual object models, reduces the dependence on manual bone annotation and binding, and improves the efficiency of bone binding.
[0075] Exemplary embodiments of this disclosure also provide an electronic device. The electronic device may include a processor and a memory. The memory stores executable instructions for the processor, such as a computer program. The processor executes the executable instructions to perform the method steps of various exemplary embodiments of this disclosure. Furthermore, the electronic device may also include a display for displaying a graphical user interface.
[0076] The following is for reference. Figure 10 The electronic device is illustrated by way of a general-purpose computing device. It should be understood that... Figure 10 The electronic device 1000 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0077] like Figure 10 As shown, the electronic device 1000 may include: a processor 1010, a memory 1020, a bus 1030, an I / O (input / output) interface 1040, a network adapter 1050, and a display 1060.
[0078] The memory 1020 may include volatile memory, such as RAM 1021 and cache unit 1022, and may also include non-volatile memory, such as ROM 1023. The memory 1020 may also include one or more program modules 1024, such program modules 1024 including, but not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. For example, program module 1024 may include the modules described above.
[0079] The processor 1010 may include one or more processing units, such as an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit).
[0080] The processor 1010 can be used to execute executable instructions stored in the memory 1020, such as the above-described skeleton binding method, which includes the following steps: Step S110: Acquire multi-view images of the virtual object model and the preset skeleton specification of the virtual object model; Step S120: Generate prompt information for the target skeleton of the virtual object model based on the preset skeleton specification; Step S130: Input the prompt information and the multi-view images corresponding to the target skeleton into a multimodal large language model to obtain the bone node position estimation information of the target skeleton, and obtain the three-dimensional position of the bone node of the target skeleton based on the bone node position estimation information; Step S140: Map the three-dimensional position of the bone node to a skeleton instance based on the preset skeleton specification, and obtain the skeleton binding asset based on the skeleton instance.
[0081] The processor 1010 executes the above method steps. On the one hand, after obtaining the preset skeleton specification of the virtual object model, it generates the target skeleton prompt information of the virtual object model based on the preset skeleton specification. The prompt information and the multi-view image of the virtual object model are input into the multimodal large language model to obtain the bone node position estimation information of the target skeleton. This achieves the goal of understanding joint semantics through natural language description without collecting annotation data for species such as dragons, quadrupeds, and mechs, significantly reducing data costs and model maintenance burden. On the other hand, the three-dimensional position of the bone node of the target skeleton is obtained based on the bone node position estimation information. Based on the preset skeleton specification, the three-dimensional position is mapped to a skeleton instance, and the skeleton binding asset is obtained based on the skeleton instance. This achieves automatic binding of the skeleton of the virtual object model, reduces the dependence on manual skeleton annotation and binding, and improves the efficiency of skeleton binding.
[0082] Bus 1030 is used to connect different components of electronic device 1000, and may include data bus, address bus and control bus.
[0083] Electronic device 1000 can communicate with one or more external devices 1100 (such as keyboard, mouse, external controller, etc.) through I / O interface 1040.
[0084] Electronic device 1000 can communicate with one or more networks via network adapter 1050. For example, network adapter 1050 can provide mobile communication solutions such as 3G / 4G / 5G, or wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication. Network adapter 1050 can communicate with other modules of electronic device 1000 via bus 1030.
[0085] Electronic device 1000 can display a graphical user interface via display 1060.
[0086] although Figure 10 As not shown in the diagram, other hardware and / or software modules may also be configured in the electronic device 1000, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0087] As can be seen from the above, the technical solutions disclosed herein can be implemented as methods, apparatus, systems, computer program products, storage media, electronic devices, etc. Those skilled in the art will understand that various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which may be referred to as "circuit," "module," or "system," respectively.
[0088] It should be understood that this disclosure is not limited to the specific methods, steps, or structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. Those skilled in the art will readily conceive of other embodiments based on the specific implementations provided in this disclosure. Therefore, the specific implementations provided in this disclosure are merely exemplary, and the scope and spirit of this disclosure are indicated by the claims, and should cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary technical means in the art not disclosed in this disclosure.
Claims
1. A method for skeletal binding, characterized in that, include: Obtain multi-view images of the virtual object model and the preset skeleton specifications of the virtual object model; Based on the preset skeleton specification, generate prompt information for the target skeleton of the virtual object model; The prompt information and the multi-view image corresponding to the target skeleton are input into a multimodal large language model to obtain the bone node position estimation information of the target skeleton, and the three-dimensional position of the bone node of the target skeleton is obtained based on the bone node position estimation information. Based on the preset skeleton specification, the three-dimensional position of the bone node is mapped to a skeleton instance, and the skeleton-bound asset is obtained according to the skeleton instance.
2. The method according to claim 1, characterized in that, Acquire multi-view images of the virtual object model, including: Get the set of preset rendering perspectives; The virtual object model is rendered according to each rendering perspective in the preset rendering perspective set to obtain the multi-view image.
3. The method according to claim 1, characterized in that, The method further includes: When rendering the virtual object model, obtain the camera projection matrix corresponding to each rendering viewpoint.
4. The method according to claim 3, characterized in that, The target skeleton includes visible and invisible skeletons. The prompt information for generating the target skeleton of the virtual object model based on the preset skeleton specification includes: The preset skeleton specification is parsed to generate visual language prompts for the visible bones; wherein, the visual language prompts include at least one of the following: bone name, hierarchical relationship, and visual description; The preset skeleton specification is parsed to obtain the hierarchical relationship and functional description of the invisible bones. Based on the hierarchical relationship and the functional description, semantic reasoning prompts for the invisible bones are obtained.
5. The method according to claim 4, characterized in that, When the target skeleton is a visible skeleton, the step of inputting the prompt information and the multi-view image corresponding to the target skeleton into a multimodal large language model to obtain the skeleton node position estimation information of the target skeleton includes: The visual language prompts and the multi-view images corresponding to the visible skeletons are input into the multimodal large language model; The output of the multimodal large language model is obtained, and the output is subjected to structured parsing to obtain the two-dimensional coordinate information of the skeletal nodes; The two-dimensional coordinate information is converted into image pixel coordinates, and the image pixel coordinates are verified to obtain the estimated position information of the skeletal node.
6. The method according to claim 5, characterized in that, The step of obtaining the three-dimensional positions of the bone nodes of the target bone based on the bone node position estimation information includes: The credibility of the estimated bone node position information is evaluated to obtain the two-dimensional observation position of the target in the estimated bone node position information. The three-dimensional position is obtained based on the two-dimensional observation position of the target.
7. The method according to claim 6, characterized in that, The process of evaluating the credibility of the skeletal node position estimation information to obtain the target two-dimensional observation position in the skeletal node position estimation information includes: Obtain the confidence level of the multimodal large language model output corresponding to the skeletal node position estimation information; The geometric reliability of the bone node position estimation information is determined based on the reprojection error of the bone node position estimation information. Based on the confidence level and the geometric confidence level, the joint confidence level of the skeletal node position estimation information is determined. The skeletal node position estimation information is then filtered based on the joint confidence level to obtain the target two-dimensional observation position.
8. The method according to claim 7, characterized in that, The process of obtaining the three-dimensional position based on the two-dimensional observation position of the target includes: Based on the target's two-dimensional observation position and the camera projection matrix corresponding to the multi-view image, determine the spatial geometric constraints corresponding to the bone nodes of the target skeleton; Based on the joint confidence level corresponding to the two-dimensional observation position of the target, a weighted geometric solution is performed on the spatial geometric constraints to obtain the three-dimensional position of the bone nodes of the target skeleton.
9. The method according to claim 4, characterized in that, When the bone is invisible, the prompt information and the multi-view image corresponding to the target bone are input into a multimodal large language model to obtain the bone node position estimation information of the target bone, including: The semantic reasoning prompts and the multi-view images corresponding to the invisible skeletons are input into the multimodal large language model, and the semantic relative positional relationship of the invisible skeletons relative to the associated skeleton nodes is obtained based on the multimodal large language model. The position estimation information of the skeletal nodes is obtained based on the semantic relative positional relationship.
10. The method according to claim 9, characterized in that, The step of obtaining the three-dimensional positions of the bone nodes of the target bone based on the bone node position estimation information includes: Obtain the 3D position of at least one associated bone node that has a topological association with the invisible bone, and generate position constraint information based on the semantic relative positional relationship; Based on the positional constraint information corresponding to the invisible bone and the skeleton completion constraints included in the preset skeleton specification, the spatial positional relationship of the invisible bone relative to the associated bone node is determined. Based on the spatial relationship and the three-dimensional position of the associated skeletal node, the three-dimensional position of the skeletal node of the invisible bone is determined.
11. The method according to claim 1, characterized in that, The process of mapping the 3D positions of the bone nodes to skeleton instances based on the preset skeleton specification, and obtaining the skeleton-bound assets based on the skeleton instances, includes: Based on the bone node topology, bone naming rules, and bone hierarchy in the preset skeleton specification, the three-dimensional positions of the bone nodes are mapped in a standardized manner to obtain a skeleton instance corresponding to the preset skeleton specification. Based on the three-dimensional position of the bone node and the bone hierarchy and bone direction constraints in the preset skeleton specification, determine the bone transformation information corresponding to the bone node; Based on the skeleton instance, establish the skinning association between the skeleton nodes and the mesh vertices of the virtual object model; The bound asset is obtained based on the skeleton instance, the skinning relationship, and the skeleton transformation information.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 11.
13. An electronic device, characterized in that, include: processor; Memory for storing the executable instructions of the processor; The processor is configured to execute the method of any one of claims 1 to 11 by executing the executable instructions.