Derivation model construction method, derivation model construction device, recording medium, construction device and construction method

By building an inference model and machine learning, the process of generating 3D representations from 2D images is simplified, solving the problem of heavy workload for designers when manually adjusting deformations, and achieving efficient 3D representation generation.

CN115461784BActive Publication Date: 2025-09-23LIVE2D
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180031690.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-18
Publication Date
2025-09-23
Estimated Expiration
2041-02-18

AI Technical Summary

Technical Problem

In the prior art, when generating a 3D model, the designer needs to manually define and adjust the deformation of each component to achieve the desired 3D representation, which results in a large workload and may be burdensome for unfamiliar designers.

Method used

By building an inference model and using machine learning to deform the components of the 2D image to generate a 3D representation, the distribution of control points and machine learning are used to build the inference model to achieve the 3D representation expected by the designer.

Benefits of technology

It simplifies the process of designers generating desired 3D expressions, reduces the workload of manual adjustment, and improves efficiency and convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115461784B_ABST
    Figure CN115461784B_ABST
Patent Text Reader

Abstract

The inference model construction method enables a computer to construct an inference model for the following expression model, wherein the expression model realizes a depiction expression corresponding to a state different from the reference state by deforming each component of a two-dimensional image of an object object corresponding to a reference state, the inference model infers the deformation of each component of the two-dimensional image under a defined state different from the reference state, the expression model is constructed so as to become defined by defining the deformation of each component of the two-dimensional image under at least one defined state, and is at least capable of realizing a depiction expression corresponding to a state between the reference state and the defined state, the deformation of each component in the expression model is controlled by the distribution form of control points set in the component, the inference model construction method includes: an acquisition step, for the defined expression model, obtaining the distribution of control points in the reference state and the defined state; an extraction step, based on the distribution of control points in the reference state obtained in the acquisition step, extracting a first feature value; and a learning step, using the first feature value extracted in the extraction step as a label, performing machine learning on the distribution of control points in the defined state obtained in the acquisition step, and constructing the inference model based on the learning results obtained by learning multiple defined expression models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an inference model construction method, an inference model construction device, a program, a recording medium, a construction device, and a construction method, and particularly to a technology for performing three-dimensional rendering and expression using a two-dimensional image. Background Art

[0002] In recent years, in the technical field of computer graphics represented by electronic games, depiction using 3D models has become mainstream. This is because, for example, when depicting frames at multiple consecutive time points while moving the same character and background as in animation, the time required for repeated depictions and light source calculations can be reduced. In particular, for interactive content such as the actions of characters in electronic games that change in real time relative to operational input, if 3D models, animation data, etc. are prepared in advance, depiction from various viewpoints and directions can be handled. For such 3D graphics, a 3D model is usually constructed based on multiple 2D images (cutouts) prepared by designers, and depiction is performed by applying textures to the model.

[0003] On the other hand, 3D graphics created through the application of such textures can sometimes create an impression different from the original cutouts prepared by designers. 3D graphics are essentially textured 3D models that are "accurately" rendered for a specific viewpoint. Therefore, it's difficult to reproduce an effective "appearance" from a specific viewing direction, as with cutouts drawn in 2D images. Therefore, the unique expressive appeal of 2D graphics is prioritized, and 2D graphics are primarily used in game content, for example, to maintain a certain degree of support.

[0004] Patent Document 1 discloses a rendering technique capable of expressing (3D rendering) a three-dimensional animation while maintaining the atmosphere and charm of a two-dimensional image created by a designer. Specifically, Patent Document 1 decomposes a two-dimensional image into its components, such as hair, eyebrows, eyes, and outline (face). A curved surface is then simply assigned to each outline component, consistent with the appearance of the two-dimensional image. The remaining two-dimensional image components are then geometrically deformed and moved to correspond to the spherical surface of the outline, rotated according to the orientation of the face to be rendered. Various adjustments are applied to achieve the desired rendering from different directions without disrupting the original impression of the two-dimensional image. In other words, unlike rendering methods that employ texture alone, the method described in Patent Document 1 employs a method for deforming the two-dimensional image to achieve the desired rendering.

[0005] Prior art literature

[0006] Patent Literature

[0007] Patent Document 1: Japanese Patent Application Laid-Open No. 2009-104570 Summary of the Invention

[0008] Problems to be solved by the invention

[0009] In the rendering technology described in Patent Document 1, for example, for a 2D image rendered in a reference orientation, such as the front view, deformations of each component are defined to produce a rendering representation in directions (angles) different from the reference orientation. This allows for the generation of a 3D representation for an angular range encompassing the reference orientation. Specifically, to create the desired 3D representation, the designer must define the deformations of each component at the angle (direction) corresponding to the end of the desired angular range, and then make fine adjustments while confirming that the desired 3D representation is achieved at other angles (directions) within the range.

[0010] However, such definition and adjustment may cause a lot of work, and designers who are not particularly familiar with these processes may worry about the load on them.

[0011] The present invention has been completed in view of the above-mentioned problems, and its purpose is to provide an inference model construction method, an inference model construction device, a program, a recording medium, a composition device and a composition method, which can easily generate the expression desired by the designer in the method of deforming a two-dimensional image to obtain a three-dimensional expression.

[0012] Means for solving problems

[0013] To achieve the above-mentioned object, the inference model construction method of the present invention causes a computer to construct an inference model for the following representation model, wherein the representation model realizes a depiction representation corresponding to a state different from the reference state by deforming each component of a two-dimensional image of an object corresponding to a reference state of the object, the inference model deducing the deformation of each component of the two-dimensional image under a defined state different from the reference state, characterized in that the representation model is configured to be defined by defining the deformation of each component of the two-dimensional image under at least one defined state, and is capable of realizing a depiction representation corresponding to at least a state between the reference state and the defined state, and the deformation of each component in the representation model is controlled by the distribution form of control points set in the component. The inference model construction method includes: an acquisition step of acquiring the distribution of control points in the reference state and the defined state for the defined representation model; an extraction step of extracting a first feature value based on the distribution of control points in the reference state obtained in the acquisition step; and a learning step of performing machine learning on the distribution of control points in the defined state obtained in the acquisition step using the first feature value extracted in the extraction step as a label, and constructing the inference model based on the learning results obtained by learning a plurality of defined representation models.

[0014] Effects of the Invention

[0015] With such a configuration, according to the present invention, in the method of deforming a two-dimensional image to obtain a three-dimensional representation, it is possible to easily generate a representation desired by a designer.

[0016] Other features and advantages of the present invention will become clear from the following description with reference to the accompanying drawings. In the accompanying drawings, the same or equivalent structures are denoted by the same reference numerals. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are incorporated in and constitute a part of the specification, illustrate embodiments of the present invention, and together with the description, serve to explain the principles of the present invention.

[0018] Figure 1 1 is a block diagram showing the functional structure of the construction device 100 according to the embodiment of the present invention.

[0019] Figure 2 It is a block diagram showing the functional configuration of the component device 200 according to the embodiment of the present invention.

[0020] Figure 3A This is a diagram illustrating a depiction expression (reference state) of a representation model according to an embodiment of the present invention.

[0021] Figure 3B This is a diagram illustrating a depiction expression (definition state) of a representation model according to an embodiment of the present invention.

[0022] Figure 4A It is a diagram for explaining a modification of the component parts of the embodiment of the present invention (reference state).

[0023] Figure 4B It is a diagram for explaining a modification of the components of the embodiment of the present invention (definition state).

[0024] Figure 5A This is a diagram (reference state) of a control point group for explaining deformation of a component part according to an embodiment of the present invention.

[0025] Figure 5B This is another diagram (defined state) of a control point group for explaining a modification of a component according to an embodiment of the present invention.

[0026] Figure 5C It is a diagram of a control point group for explaining deformation of a component part according to an embodiment of the present invention.

[0027] Figure 6A These are diagrams for explaining differences in the movable ranges of expression models according to the embodiments of the present invention.

[0028] Figure 6B This is another diagram for explaining the difference in the movable range of the expression model according to the embodiment of the present invention.

[0029] Figure 6C This is another diagram for explaining the difference in the movable range of the expression model according to the embodiment of the present invention.

[0030] Figure 7A This is a diagram for explaining the standardization of the expression model according to the embodiment of the present invention.

[0031] Figure 7B This is another diagram for explaining the standardization of the representation model according to the embodiment of the present invention.

[0032] Figure 8A This is a diagram for explaining the second characteristic amount according to the embodiment of the present invention.

[0033] Figure 8B This is another diagram for explaining the second feature value according to the embodiment of the present invention.

[0034] Figure 9 This is a flowchart illustrating the construction process executed in the construction device 100 according to the embodiment of the present invention.

[0035] Figure 10 This is a diagram illustrating a GUI for adjusting component deformation after derivation according to an embodiment of the present invention.

[0036] Figure 11AThis is a diagram for explaining a process of matching arrangement relationships when adjusting deformation of component parts according to an embodiment of the present invention.

[0037] Figure 11B This is another diagram for explaining the process of matching the arrangement relationship when adjusting the deformation of the component parts according to the embodiment of the present invention.

[0038] Figure 11C This is another diagram for explaining a process of matching arrangement relationships when adjusting deformation of component parts according to an embodiment of the present invention.

[0039] Figure 12 This is a flowchart illustrating a configuration process executed in the configuration device 200 according to the embodiment of the present invention.

[0040] Figure 13A This is a diagram illustrating the data structure of a presentation model according to an embodiment of the present invention.

[0041] Figure 13B This is a diagram illustrating the data structure of the expression model (texture information) according to the embodiment of the present invention.

[0042] Figure 13C This is a diagram illustrating the data structure of the presentation model (state information) according to the embodiment of the present invention. DETAILED DESCRIPTION

[0043] [Implementation Method]

[0044] The following embodiments are described in detail with reference to the accompanying drawings. The following embodiments do not limit the inventions described in the claims, and not all combinations of features described in the embodiments are essential to the invention. Any combination of two or more of the multiple features described in the embodiments may be used. Identical or equivalent structures are denoted by the same reference numerals, and duplicate descriptions are omitted.

[0045] In one embodiment described below, the present invention is applied to an example of a construction device that constructs an inference model by performing machine learning on multiple defined representation models, and a composition device that constructs a representation model using the inference model constructed by the construction device. In this embodiment, the inference model construction device and the composition device are described as separate devices, but the present invention is not limited to this embodiment; the present invention may also be implemented in a single device having these functions.

[0046] In this specification, a "representation model" is described as data that achieves a representation corresponding to a desired state by deforming a two-dimensional image of a reference state of an object (target object) for each component, thereby generating a two-dimensional image corresponding to that state, and achieves a three-dimensional representation by continuously representing the process of state transitions. More specifically, the representation model is constructed to include information on the deformation of the two-dimensional image of each component for states (defined states) different from the reference state in order to achieve the representation of the desired state. For example, by using the technology described in Patent Document 1, the representation model is capable of achieving a representation corresponding to at least states that can transition (interpolate) between the reference state and the defined state. However, the representation model is not limited to achieving a representation by interpolating states between two defined states and can also be constructed to achieve a representation of states not included in the defined state by extrapolation.

[0047] In this specification, "derive" refers to applying a predetermined input to an inference model and deriving a predetermined output based on, for example, a neural network constructed using the inference model. In contrast, "estimate" refers to deriving the state assumed by the input by applying a predetermined operation to the input without using an inference model.

[0048] The structure of the construction device

[0049] Figure 1 This is a block diagram illustrating the functional structure of the construction device 100 according to this embodiment. Construction device 100 may be, for example, a server managed at a point of sale for an editing application used to construct a presentation model, described later. Construction device 100 is configured to connect to construction device 200 via a network (not shown) and to provide the constructed derivation model.

[0050] The control unit 101 is a control device such as a CPU, and controls the operation of each module included in the construction device 100. Specifically, the control unit 101 reads the operation program of each module stored in the recording medium 102, loads it into the memory 103, and executes it, thereby controlling the operation of each module.

[0051] The recording medium 102 is, for example, a non-volatile memory such as a rewritable ROM, or a recording device such as an HDD that can be detachably connected to the construction device 100. In addition to recording the operation programs of each module possessed by the construction device 100, the recording medium 102 also records information such as parameters required for the operation of each module. In addition, in the present embodiment, the recording medium 102 records a variety of representation models used in machine learning (representation models as learning objects). The memory 103 can be, for example, a volatile memory such as a RAM. The memory 103 is not only used as a loading area for loading programs read from the recording medium 102, but also as a storage area for temporarily storing intermediate data output during the operation of each module.

[0052] The normalization unit 104 performs a normalization process on the representation model to be learned in order to converge the machine learning performed by the learning unit 108. The details of the normalization process performed by the normalization unit 104 will be described later. The normalized representation model can be stored in the recording medium 102.

[0053] During machine learning of the expression model, the acquisition unit 105 reads and acquires the expression model normalized by the normalization unit 104 from the recording medium 102. The normalized expression model acquired by the acquisition unit 105 is sent to the extraction unit 106, the estimation unit 107, and the learning unit 108.

[0054] The extraction unit 106 extracts three feature quantities as the first feature quantity of the present invention based on the information of the reference state of the normalized representation model. The details of these three feature quantities will be described later. The first feature quantity is information obtained by quantifying the features of the appearance of the target object.

[0055] Based on the information about the reference state and the defined state of the standardized representation model, the estimation unit 107 estimates two feature quantities as the second feature quantity of the present invention. Details of these two feature quantities will be described later. However, unlike the first feature quantity, the second feature quantity is information obtained by estimating and quantifying the main causes of deformation defined for the object.

[0056] The learning unit 108 performs machine learning on the standardized representation model transmitted from the acquisition unit 105 based on the feature values ​​obtained for the representation model by the extraction unit 106 and the estimation unit 107. The learning unit 108 constructs an inference model based on the learning results obtained by learning multiple standardized representation models. The inference model can be obtained using a neural network, but it can also be obtained using other methods.

[0057] The communication unit 109 is a communication interface provided by the construction device 100 for communicating with external devices. The communication unit 109 is connected to the external device via a network (which can be wired or wireless) to transmit and receive data. The network can be a communication network such as the Internet or a local area network that connects devices.

[0058] The structure of the device

[0059] Next, refer to Figure 2 The functional structure of the composition device 200 of this embodiment will be described. Here, the composition device 200 may be a user terminal that executes an editing application for constructing a representation model, and is configured to obtain a derivation model from the construction device 100 via a network (not shown). Furthermore, in the structure of the composition device 200 of this embodiment, components that implement the same functions as those of the construction device 100 are distinguished by being prefixed with "composition."

[0060] The configuration control unit 201 is a control device such as a CPU, and controls the operation of each module included in the configuration device 200. Specifically, the configuration control unit 201 reads the operation program of each module stored in the configuration recording medium 202 and the program of the editing application for configuring the performance model (described later), loads the program into the configuration memory 203, and executes the program, thereby controlling the operation of each module.

[0061] The configuration recording medium 202 is, for example, a nonvolatile memory such as a rewritable ROM, or a recording device such as an HDD detachably connected to the configuration device 200. In addition to recording the operating programs and editing applications of each module included in the configuration device 200, the configuration recording medium 202 also records information such as parameters required for the operation of each module. Furthermore, in this embodiment, the configuration recording medium 202 records the inference model constructed by the construction device 100. The configuration memory 203 can be, for example, a volatile memory such as RAM. The configuration memory 203 is used not only as a loading area for loading programs read from the configuration recording medium 202, but also as a storage area for temporarily storing intermediate data output during the operation of each module.

[0062] The rendering unit 204 may be a rendering device such as a GPU, for example, and generates an image (screen) displayed in the display area of ​​the display unit 220. In the configuration device 200 of this embodiment, when the editing application is executed, the rendering unit 204 renders, for at least the representation model being edited, a two-dimensional image that realizes the appearance of the object represented by the representation model, that is, realizes the representation of the object in a specified state.

[0063] The display control unit 205 performs display control related to the display of the screen generated by the rendering unit 204 on the display unit 220. The display unit 220 may be, for example, a display device such as an LCD, and may be integrally formed with the component device 200 or may be an external device that is attachable to and detachable from the component device 200.

[0064] The setting unit 206 sets control points that serve as a basis for deformation of each component of the two-dimensional image of the object (the object constituting the object) that constitutes the representation model. Regarding the deformation of each component involved in depicting the change in representation, the two-dimensional image of the component is applied (mapped) to a curved surface, and the positions of the control points set for the curved surface are changed to change the shape of the curved surface, thereby achieving a change in the appearance of the component. The details will be described later. Therefore, in the editing application, the setting unit 206 sets the distribution of control points in the reference state based on user input based on the definition of the reference state of the object constituting the object. In addition, in this embodiment, the user is requested to set control points when defining the reference state in order to achieve the user's desired representation. However, the implementation of the present invention is not limited to this. The setting unit 206 can also set control points without user operation based on image recognition, layer structure analysis, etc.

[0065] When constructing a representation model of an object to be constructed, the determination unit 207 determines a feature quantity related to a definition state to be derived by the derivation unit 208. The feature quantity determined by the determination unit 207 includes a first feature quantity and a second feature quantity.

[0066] The derivation unit 208 derivates the distribution of control points related to the definition state of the object to be constructed using the feature quantity determined by the determination unit 207. The derivation unit 208 uses the derivation model constructed by the construction device 100.

[0067] The composition unit 209 constructs a representation model of the object to be composed based on the derivation result of the derivation unit 208. Specifically, the composition unit 209 constructs the representation model so as to include the distribution of control points in the reference state set by the setting unit 206 and the distribution of control points in the defined state deduced by the derivation unit 208.

[0068] The operation input unit 210 is, for example, a user interface such as a mouse, keyboard, or tablet included in the component device 200. Upon detecting an operation input via any of the interfaces, the operation input unit 210 outputs a control signal corresponding to the operation input to the component control unit 201. Alternatively, the operation input unit 210 notifies the component control unit 201 of the occurrence of an event corresponding to the operation input.

[0069] The communication unit 211 is a communication interface provided by the device 200 for communicating with external devices. The communication unit 211 is connected to the external device via a network (which can be wired or wireless) to transmit and receive data. The network can be a communication network such as the Internet or a local area network that connects devices.

[0070] The Structure of the Representation Model

[0071] Next, the detailed structure of the representation model used in this embodiment will be described. Furthermore, in this embodiment, it is assumed that the object provided for depiction by the representation model is a character including at least a head, and a two-dimensional image representing the appearance of the character is displayed by presenting the representation model for illustration.

[0072] like Figure 3A As shown, the appearance of the character represented by the representation model is constructed using 2D images of the various components of the character's head (face, left eye, right eye, nose, mouth, front hair, etc.) configured for the character's frontal orientation. The 2D images of each component are configured to deform independently or in conjunction with the deformation of other components, and this deformation is controlled according to the state of the displayed character. In other words, the representation model is configured so that the character can be depicted in states different from the base state by deforming the 2D image of the character in the frontal orientation state, which serves as the base state, rather than switching to a fixed 2D image for each state of the displayed character.

[0073] In order to achieve a depiction expression in a state different from the reference state, the user needs to define what kind of depiction expression the expression model will become in at least one state other than the reference state. The user defines the state (defined state) in which the desired depiction expression is presented in a state other than the reference state, and deforms the various components of the character in the reference state, thereby defining the depiction expression that should be presented in the defined state. As a result, the expression model can at least derive the deformation form of each component for the state within the range specified by the reference state and the defined state (hereinafter referred to as the movable range), and can achieve the corresponding depiction expression. Hereinafter, the expression model in which the deformation of at least one defined state is defined in this way and the state in which the depiction expression in a state different from the reference state is achieved will be referred to as a "defined expression model", and the expression model in which the deformation of any defined state is not defined and the state in which the depiction expression in the reference state can only be achieved will be referred to as an "undefined expression model".

[0074] For example, in Figure 3A In the case where a character in a front-facing state is defined as a reference state, the character in the reference state can be deformed by deforming the two-dimensional image of each component of the character. Figure 3B The character representation in the obliquely oriented state as shown is defined, and corresponding representation can be achieved for each state within the movable range between the two states. In addition, in this embodiment, in order to facilitate the understanding of the invention, the representation model is defined as follows: Figure 3A and Figure 3B state, and is configured to enable the depiction of a character's head rotating in the Yaw direction for illustration. Here, the so-called rotational movement in the Yaw direction is not limited to movements such as shaking the head, in which the position of the character does not move, but only the head (or the entire character) rotates in the Yaw direction, but also includes the following situation: a change in the way of viewing the character caused by the sideways movement of the character (a way of viewing the character from an oblique direction by deviating from the optical axis (sight direction) of the camera performing the depiction). However, the embodiments of the present invention are not limited to this. The expression model can also be configured to define multiple defined states such as posture changes in other directions, translational movements, or specific actions and expressions, and can achieve a transition in the depiction expression from a reference state to a defined state, or can achieve a transition in the depiction expression from a reference state to a state obtained by combining these defined states.

[0075] In this embodiment, the deformation of the two-dimensional image representing each component of the model is achieved by applying the two-dimensional image as a texture to a curved surface (including a plane) and changing the shape of the curved surface. More specifically, in order to define the shape of the curved surface, a plurality of control points for determining the shape are set on each curved surface, and the shape of the curved surface and the deformation of the applied two-dimensional image are achieved by making the distribution of the control points different. In this embodiment, a method of applying the two-dimensional image as a texture to a curved surface and achieving the deformation of the two-dimensional image by changing the shape of the curved surface is described, but the embodiments of the present invention are not limited to this. That is, instead of using an indirect concept such as a curved surface, a method of directly controlling the deformation of the two-dimensional image of the component can be adopted only by the distribution of the control points set for the two-dimensional image.

[0076] For example, Figure 4A As shown in FIG, a curved surface and control points are set for the two-dimensional image of each component in the reference state. Here, the number and distribution of control points set for each component can be determined according to the required level of detail for the desired deformation expression of the component. For example, regarding the reference state and definition state of the facial component, as shown in FIG. Figure 5A and 5BAs shown in FIG, in addition to setting control point groups for the four corners (vertices) of the applied curved surface (rectangle), control point groups are also set for the middle points that divide the edge of the curved surface into a predetermined number (in the example of the figure, it is divided into two parts in the horizontal direction and three parts in the vertical direction) and the intersection points determined by connecting them. Figure 5C As shown, in Figure 5A and 5B The control point group set for each vertex, midpoint, and intersection point can include control point 501 and control point 502. Control point 501 is equivalent to a positioning point that specifies the position of the point corresponding to the surface, and control point 502 is used to define the direction line of the curve (broken line, Bezier curve, etc.) connecting adjacent control points (or other control points on its extension line).

[0077] The definition of the character's depiction in the definition state is as follows Figure 4B As shown in , this is done by changing the distribution of these control points for each component. Figure 4A and 4B In the example shown, the distribution of control points (the shape of the curved surface) is shown only for the face, eyes, and mouth of the character in order to facilitate visual confirmation of the distribution of control points. However, deformation can of course be defined in the same way for other components.

[0078] When the deformation of the 2D image of each component is defined in this way, the depiction expression in any state included in the movable range specified by the reference state and the defined state can be derived by interpolating the configuration coordinates of the control points in these states. More specifically, the state of the object being depicted (object state) can be specified by the internal ratio between the reference state and the defined state, so by weighted addition of the configuration coordinates of the same control point in the reference state and the defined state based on the internal ratio, the configuration coordinates of the control point related to the object state can be derived. That is, when depicting a character using a defined representation model, the object state can be determined based on the internal ratio in the movable range, and the depiction expression of the object state can be generated using this internal ratio as input.

[0079] So, for example Figure 13AAs shown, the representation model can be configured to include texture information 1302, reference state information 1303, and defined state information 1304, all associated with a model ID 1301. Model ID 1301 uniquely identifies the representation model; texture information 1302 represents various information about the two-dimensional images of various components of the character involved in the representation model; reference state information 1303 represents the distribution of control points defined in various components for the reference state; and defined state information 1304 represents the distribution of control points defined in various components for the defined state. As described above, in this embodiment, to facilitate understanding of the invention, only one defined state other than the reference state is defined for the representation model. Therefore, the representation model data is described as including only one type of defined state information 1304. However, the embodiments of the present invention are not limited to this. Alternatively, the representation model data may include as many defined state information 1304 as desired. In this case, each defined state information 1304 includes identification information uniquely identifying the defined state.

[0080] Here, the texture information 1302 defining the state may be configured to, for example, Figure 13B As shown, the component ID 1311 for identifying the component includes: a role ID 1312 indicating the role of the component in the character's appearance (right eye, nose, face, etc.), a two-dimensional image 1313 of the component, detailed information 1314 storing various information such as the size of the two-dimensional image of the component, and applied surface information 1315 storing the size of the surface to which the two-dimensional image of the component is applied, a list of set control points, etc. In addition, the reference state information 1303 and the definition state information 1304 are configured so that, for each component, Figure 13C As shown, a control point ID 1322 and the arrangement coordinates 1323 of the control point in the object state are managed in association with a component ID 1321 that identifies the component. The control point ID 1322 uniquely identifies each control point set in the curved surface related to the component.

[0081] In addition, detailed description of the configuration coordinates of the control point is omitted. The configuration coordinates of the control point can be absolutely specified with the specified origin set for the entire character as the center, for example, the configuration coordinates of the control point can also be relatively specified with the specified coordinates of the component in which the control point is set and the specified coordinates of other components associated with the component as the center.

[0082] Construction of Derivative Model

[0083] Next, we will describe the construction of a derivation model performed by the construction device 100 of this embodiment. This derivation model is based on machine learning using multiple representation models that define the distribution of control points for the reference state and the defined state. By using the derivation model constructed by the construction device 100, the distribution of control points associated with a new defined state can be derived from a representation model (undefined representation model) that only defines the distribution of control points associated with the reference state. Details will be described later.

[0084] Standardization of the presentation model

[0085] The derivation model is constructed by using multiple defined performance models as teacher data and having the learning unit 108 perform machine learning. However, since the movable ranges specified by the performance models are not necessarily consistent, it is difficult to directly perform machine learning on the distribution of control points represented by various performance models. Figures 6A to 6C As shown, even if the baseline state shows a character facing forward, the form of the depiction of the defined state is not common to all representation models. Therefore, even if only the distribution of control points related to the defined state is learned, a good inference model cannot be constructed.

[0086] This is because the depiction expression in the definition state varies depending on the expression method adopted by the designer and the purpose of the expression model.

[0087] In the case where the expression model is configured to prompt the rotation movement of the character's head in the yaw direction as in the present embodiment, the angle range of the rotation movement that can be achieved by the expression model will also differ depending on the environment in which the expression model is used. For example, when used in a chat system that prompts a character in a quarrel scene obtained by controlling the movement of the user's real-world movements, it is preferable to make the character's movement more exaggerated than the user's real-world movement, or to increase the amplitude of the character's movement to make it more eye-catching, so the angle range required for achieving the depiction in the expression model becomes larger. On the other hand, when used in an adventure game that prompts a roughly full-body (medium shot, full view) of a character such as a so-called standing picture, the amplitude of the character's movement is not so required, so the angle range required for achieving the depiction in the expression model becomes narrower.

[0088] Furthermore, especially when depicting characters viewed from an oblique angle, strictly speaking, the design is limited to specifying a specific numerical value for the rotation angle from the front (the viewing angle relative to the front direction). In most cases, the resulting depiction is based on the designer's perception. Therefore, when attempting to learn an arbitrary representation model, it is difficult to derive the value of the rotation angle indicated by the representation in the defined state of that representation model. On the other hand, even when designing with a specified rotation angle value, different designers may use different representation methods. Therefore, even for representation models with the same purpose, the representation in the defined state (the arrangement and deformation of the components) may differ.

[0089] Furthermore, in a representation model that specifies changes in viewing angles associated with sideways character movement, while including the same rendering as yaw rotation, the rendering in the defined state is defined by shifting all components in the direction of movement. Therefore, even in a representation model that implements yaw rotation, differences in the rendering in the defined state may occur depending on whether or not translation is used.

[0090] Therefore, even if they are all defined expression models, there are differences in the depiction expressions they can achieve. Therefore, in order to obtain appropriate learning results, in the construction device 100 of this embodiment, the defined expression model is standardized and used as teacher data.

[0091] Here, the standardization of the defined expression model can be performed mainly through two steps: "standardization of scale" and "standardization of deformation amount".

[0092] First, the size (dimensions) of the two-dimensional image of the character used in the representation model is preferably optimized according to the intended use of the representation model. Specifically, since the size of the range within which the control points are distributed varies for each representation model, the standardization unit 104 standardizes the scale based on the configuration of the character's eye components, which constitute the first category of components in the present invention, in a reference state.

[0093] First, if Figure 7A As shown, the normalization unit 104 refers to the reference state information 1303 in the representation model data and derives the interocular distance 701 of the representation model based on the arrangement coordinates of the left-eye component and the right-eye component (e.g., the center coordinates of the curved surface and the arrangement coordinates of the control points assigned to the pupil). The normalization unit 104 then divides the values ​​of the arrangement coordinates 1323 in the reference state information 1303 and the definition state information 1304 by the interocular distance, thereby constructing the scaled and normalized representation model data.

[0094] In addition, in this embodiment, the scale normalization is described based on the distance between the left-eye component and the right-eye component, but the implementation of the present invention is not limited to this. Of course, the scale normalization can also be performed based on the configuration of other components.

[0095] Next, as described above, the depiction of the character represented by the definition state is different for each expression model, that is, even if the component is the same, the deformation amount in the definition state (from the reference state) is different, so the standardization unit 104 standardizes the deformation amount based on the movement amount of the control point set on the nose tip of the nose component of the character as the second category of the component of the present invention for the standardized model after the scale is standardized. This is because, in the manner in which the expression model is configured to indicate the rotation movement of the character's head in the yaw direction as in the present embodiment, the component of the character's head farthest from the rotation axis is the nose tip, as Figure 7B As shown, movement based on a rotational motion can be characteristically exhibited (702).

[0096] The normalization unit 104 adjusts the definition state information 1304 of the scaled and normalized representation model data so that the amount of movement of the configuration coordinates 1323 of the control point ID 1322 set on the nose from the reference state to the definition state becomes a predetermined value (fixed value), thereby forming the defined representation model data as the training data. Here, if the amount of movement of the control point of the nose from the reference state to the definition state in the representation model data before deformation normalization is M, the configuration coordinates of any control point in the reference state and the definition state are pf and p, respectively, and the movement amount (fixed value) of the control point of the nose after normalization is D, then the configuration coordinates p' of any control point in the definition state in the representation model data after deformation normalization can be derived from the equation: p' = (p-pf) × D / M. In this way, the deformation amount based on the rotational motion can be standardized for the depicted representations in the definition states of different representation models, and a representation model that partially accommodates the differences in the range of motion between the representation models can be obtained as the training data. Hereinafter, a defined representation model in a state where the scale and the amount of change are standardized will be simply referred to as a “standardized representation model”.

[0097] Furthermore, the control point of the tip of the nose can be configured to be identifiable by associating the control point of the tip of the nose with information indicating the tip of the nose, similarly to the action ID, when defining the control point. For example, the control point of the tip of the nose can be determined as the control point with the largest horizontal coordinate movement among the control points set for the components of the nose.

[0098] In this embodiment, the deformation amount is normalized based on the amount of movement of the control points set at the nose tip. However, as with scale normalization, the present invention is not limited to this. Normalization to equalize the range of motion of each representation model may also be performed based on specific control points included in a specific component.

[0099] 〈Feature Quantities of Standardized Expression Model〉

[0100] When performing machine learning on a standardized representation model (the distribution of control points associated with a defined state), the characteristic quantities represented by the standardized representation model are assigned as labels. In this embodiment, the labels are composed of a first characteristic quantity extracted from the reference state for the standardized representation model and a second characteristic quantity estimated based on the reference state and the defined state. Each characteristic quantity is described below with reference to the accompanying drawings.

[0101] The first feature quantity is information obtained by quantifying the facial features of the character involved in the standardized representation model (e.g., long face, large eyes, etc.). The extraction unit 106 extracts the dimensions of the two-dimensional image of each component in the reference state, the dimensions of the curved surface to which the two-dimensional image is applied (e.g., based on the coordinates of the control points set at the four corners of the curved surface), and the center coordinates (e.g., the coordinates of the control points set at the center position) as the first feature quantity of the standardized representation model.

[0102] Here, by including the size of the curved surface to which the 2D image is applied, in addition to the size of the 2D image of each component, the feature quantity allows the content of the deformation to be learned as a feature. For example, if the size of the 2D image applied is small relative to the size of the curved surface, even if the curved surface is deformed, the deformation effect on the 2D image is small. On the other hand, for example, if the size of the 2D image applied is large relative to the size of the curved surface, the deformation effect of the curved surface on the 2D image becomes greater. Therefore, the degree of deformation displayed varies depending on the ratio of the size of the 2D image of each component to the size of the curved surface to which the 2D image is applied. Therefore, in the construction device 100 of this embodiment, by including these in the first feature quantity, the characteristics of the deformation displayed can be more finely classified and learned.

[0103] On the other hand, the second feature quantity is not a feature quantity explicitly shown in the normalized expression model, but is information obtained by estimating what kind of action causes the deformation of each component and quantifying it.

[0104] Here, if we imagine using the derivation results of the derivation model constructed by the construction device 100 to generate the distribution of control points in the defined state of an undefined performance model, it can be imagined that it is preferable to be able to change the obtained derivation results according to what kind of action the designer expects. That is, when generating the distribution of control points of the undefined performance model in the defined state, the deformation amount of the component generated by the rotation of the character and the deformation amount of the component generated by the translation of the character are derived separately according to the main cause of the rotation action required by the defined state (rotation / translation), and the deformation amount obtained by summing these derivation results can be determined. This method can be said to be a measure that improves the convenience of the designer. That is, the deformation of the component of the performance model as the teacher data is preferably learned in a manner that can be separated into a component generated by the rotation of the character (rotation component) and a component generated by the translation of the character (translation component).

[0105] However, as mentioned above, the deformation of components defined in the representation model varies depending on the purpose of the representation model. Even the deformation form of components in the defined state can differ depending on the representation method used by the designer. Therefore, it is difficult to determine, based solely on the representation model, information that can be used to determine which action causes the deformation of components in the defined state during the design phase. For example, it is difficult to determine, within the deformation of a component in the defined state, to what extent the deformation is caused by the character's rotation and to what extent by the character's translation. Furthermore, it is difficult to determine whether the deformation is caused not only by the action but also by the representation method used by the designer.

[0106] Therefore, in the construction device 100 of this embodiment, for the standardized expression model serving as teacher data, the distribution of control points related to the defined state is not separated according to the main cause of deformation (translation / rotation). Instead, for the distribution of control points in the state combining these main causes, the second feature value representing the translation amount and the rotation amount estimated by the estimation unit 107 is used as a label for learning.

[0107] The information indicating the translation amount is estimated based on the control points set for the facial components as the third category of components of the present invention. Figure 8A As shown, the estimation unit 107 derives the movement amount 801 of the control point set at the center of the facial component of the standardized expression model from the reference state to the defined state as information indicating the translation amount.

[0108] Furthermore, in this embodiment, the translation amount of an action related to a defined state is estimated based on the amount of movement of a control point set at the center of a facial component. However, the present invention is not limited to this. For example, estimation may be performed based on the amount of movement of at least one control point set for a facial component, such as control points set at the four corners of a curved surface associated with the facial component, or based on control points set for a component other than the face that represents the position of the character's head.

[0109] In contrast, the information indicating the amount of rotation is estimated based on the control points set for the left eyebrow / right eyebrow / nose component, which is the fourth category of components of the present invention. Figure 8B As shown, the estimation unit 107 derives the average of the movement amounts of the control points set at the brows of the left and right eyebrow components from the reference state to the defined state as the movement amount 802 related to the glabella, and derives the difference between this movement amount and the movement amount 803 of the control points set at the nose tip from the reference state to the defined state as information indicating the amount of rotation. This is achieved by utilizing the fact that the length of the trajectory traced by a point arranged in three-dimensional space when rotating about an arbitrary rotation axis varies depending on the distance from the rotation axis (rotation radius). That is, in this embodiment, the positions of the glabella and nose tip, which are arranged on the character's midline and represent the concavities and convexities of the character's face, exhibit a difference in movement amount due to the difference in rotation radius, which is obtained as information indicating the amount of rotation.

[0110] In addition, similar to the information representing the translation amount, the information representing the rotation amount is not limited to being estimated based on the control points set for the left eyebrow / right eyebrow / nose components, but can also be estimated based on the control points set for any component that represents the concavity and convexity of the character's head.

[0111] Using the first and second feature quantities obtained in this manner as labels, the learning unit 108 performs machine learning on the distribution of control points associated with the standardized representation model and constructs a derived model based on the learning results. The derived model constructed by the construction device 100 of this embodiment takes as input a representation model that defines a reference state (more specifically, a two-dimensional image of a character constituting the representation model corresponding to the reference state and the distribution of control points in that reference state) and derives and outputs the distribution of control points in the defined state. Details will be described later.

[0112] Build Process

[0113] Hereinafter, regarding the construction process executed in association with the construction of the derivation model based on the standardized representation model in the construction device 100 of this embodiment, Figure 9 The specific processing is described with reference to the flowchart. The processing corresponding to the flowchart can be implemented by the control unit 101 reading the corresponding processing program recorded in the recording medium 102, loading it into the memory 103, and executing it. The description is based on the premise that this construction process is initiated when, for example, a command to construct a derived model is received after specifying multiple standardized representation models recorded in the recording medium 102. Furthermore, when executing this construction process, the representation model to be learned is previously standardized by the standardization unit 104, and the resulting standardized representation model is stored in the recording medium 102.

[0114] In S901 , under the control of the control unit 101 , the acquisition unit 105 reads out one standardized expression model (target model) for which feature values ​​have not yet been acquired from among the plurality of standardized expression models to be learned from the recording medium 102 .

[0115] In S902, the extraction unit 106 extracts a first feature value from information related to the reference state of the object model under the control of the control unit 101. More specifically, based on the texture information 1302 and the reference state information 1303 of the object model data, the extraction unit 106 extracts the size of the two-dimensional image of each component, the size of the curved surface to which the two-dimensional image is applied, and information on the center coordinates as the first feature value.

[0116] In S903, under the control of the control unit 101, the estimation unit 107 estimates the second feature quantity based on the information related to the reference state and the definition state of the object model. More specifically, the estimation unit 107 estimates the translation amount and rotation amount related to the deformation of the component part based on the reference state information 1303 and the definition state information 1304 of the object model data, and obtains information representing the translation amount and rotation amount as the second feature quantity.

[0117] In S904, the control unit 101 determines whether the first feature value and the second feature value have been acquired for all of the plurality of standardized representation models to be learned. If the control unit 101 determines that the first feature value and the second feature value have been acquired for all of the plurality of standardized representation models to be learned, the process shifts to S905. If the control unit 101 determines that there is a standardized representation model for which the first feature value and the second feature value have not been acquired, the process returns to S901.

[0118] In S905, the learning unit 108, under the control of the control unit 101, performs machine learning using the distribution of control points in the definition state of each of the multiple standardized expression models to be learned as the following teacher data, where the first feature quantity extracted in S902 and the second feature quantity obtained in S903 are used as labels, thereby constructing an inference model. The machine learning performed in this step is repeated until the difference (loss function) between the distribution of control points output by the inference model for the labels of the teacher data and the distribution of control points of the teacher data converges.

[0119] In S906 , under the control of the control unit 101 , the learning unit 108 outputs the derivation model constructed based on the learning results obtained by learning the plurality of standardized representation models as learning targets, thereby completing the construction process.

[0120] Thus, according to the construction device of this embodiment, a plurality of expression models defining depiction expressions in various definition states can be used as teacher data, thereby constructing an inference model capable of inferring the deformation form of each component.

[0121] Construction of a performance model using a deductive model

[0122] Next, the structure of the representation model that has been defined using the derivation results of the derivation model in the constituent device 200 of this embodiment will be described.

[0123] By acquiring the derivation model constructed by the construction device 100 of this embodiment as described above, the construction device 200 can infer the deformation of each component for an undefined defined state. More specifically, given a two-dimensional image of each component and the distribution of control points associated with the surface to which the two-dimensional image is applied, which is set for a reference state of the object being constructed, the first feature value extracted based on this reference state is provided as input to the derivation model, thereby obtaining the distribution of control points associated with the defined state.

[0124] Here, the derivation model is a model obtained by machine learning using the first feature quantity and the second feature quantity as labels, so it is necessary to assign an appropriate second feature quantity during derivation. However, since the second feature quantity is information representing the translation and rotation related to the distribution of the following control points, and the distribution of the control points is a distribution related to the definition state that is predetermined to be defined by derivation, these parameters do not exist for the undefined definition state. In addition, the deformation form of the component involved in the undefined definition state depends on the expectations of the user (designer) who instructs the derivation to be performed, and there is no absolute scale for the information representing the translation and rotation, so it is unrealistic to ask the designer to specify a specific numerical value before derivation. Therefore, in the construction device 100 of this embodiment, the average value (translation and rotation) of the second feature quantity involved in all the standardized expression models learned when constructing the derivation model is output, and when the derivation is performed in the construction device 200, this average value can be used as an initial value for derivation.

[0125] In addition, in this embodiment, the average value of the second feature values ​​of all the standardized representation models used in learning is used as the initial value for the second feature value when deriving the undefined definition state. However, the embodiments of the present invention are not limited to this. For example, the second feature value can be derived based on a predetermined feature exhibited by the reference state, or the second feature value can be set according to the application of the representation model.

[0126] Then, the derivation model is used to perform derivation based on the first feature quantity obtained for the reference state of the object being constructed and the initial value of the second feature quantity obtained from the construction device 100. The distribution of control points in a predetermined definition state (dependent on the second feature quantity) of the object being constructed is output as a derivation result. The construction unit 209 can be configured to store the distribution of control points, which is the derivation result, in the definition state information 1304 of the data of the representation model related to the object being constructed, thereby setting the representation model to a defined state.

[0127] On the other hand, after setting the distribution of control points in the defined state, the two-dimensional image of each component can be deformed to display the representation model. In other words, the defined state of the object constituting the object can be visually indicated, allowing the designer to understand whether the desired depiction is achieved. Therefore, in the editing application of this embodiment, after deduction based on the derivation model, a graphical user interface (GUI) is provided for adjusting the deformation of the component in the defined state based on the distribution of control points obtained by the deduction, thereby adjusting the depiction in the defined state from the initial form based on the second feature value to a different form.

[0128] As described above, because it is difficult to identify the primary cause of deformation of a specified component within a defined representation model, the derivation result of the derivation model constructed by the construction device 100 of this embodiment can include both translational and rotational components of control point movement. In this editing application, to easily adjust the derivation result to the designer's desired rendering, the construction device 200 separates the distribution of control points in the derivation result into translational and rotational components. After adjusting the translation and rotation levels, these control point distributions are combined to obtain the adjusted distribution of control points in the defined state. Separation of the translational and rotational components can be performed, for example, by deriving a translation amount based on the amount of movement (from a reference state) of the center position of a specific component, such as a facial component, and subtracting this translation amount from the coordinates of all control points in the derivation result to obtain the rotational component. The difference between the derivation result and the rotational component, i.e., the distribution obtained by adding the coordinates of all control points to the translation amount, is used as the translational component.

[0129] For example, Figure 10 As shown, the GUI can be configured to accept adjustments to at least one of the translation amount (translation level) and rotation amount (rotation level) for deformation in the definition state of each component. When adjustments are made (slider values ​​are changed) via the GUI, the definition state information 1304 of the target representation model data is changed, and the depiction of the changed definition state is displayed on the display unit 220. When at least one of the translation amount and the rotation amount is adjusted, the distribution of the control points of the corresponding component is changed to a value corresponding to the adjustment value for the distribution of the control points of the translation component and the distribution of the control points of the rotation component separated from the derivation result. The changed control point distributions are then added together to derive the distribution of the control points in the adjusted definition state. The definition state information 1304 of the representation model data is then updated based on the distribution of the control points associated with the adjusted definition state.

[0130] Here, in use Figure 10 If the translation and rotation of each component are adjusted after derivation in the GUI shown in the example, there is a possibility that the components will not match after the adjustment. For example, if the distribution of control points of the character's face component is adjusted by the translation, the distribution of control points of the ear component will not be changed. Figure 11A As shown, there is a mismatch in the arrangement relationship between the ear component and the face component. Therefore, constraints are defined for the arrangement relationship between the ear component and the face component, and the changes to the definition state information 1304 are controlled so that the arrangement relationship is maintained before and after the adjustment.

[0131] More specifically, first, for a surface of a 2D image to which corresponding components (hereinafter, facial components will be described), for example, Figure 11B As shown, a grid 1101 is defined for the surface shape in the reference state. Here, the grid 1101 can be used to define the local deformation of the two-dimensional image applied in a specific manner, but in this embodiment, the concept of the grid 1101 is set to be used to determine the partial area in the surface represented by the distribution of control points (a rectangular area after being subdivided by the grid 1101). Then, the following information is pre-stored, which is used to determine to which partial area the connection position of the facial component and the ear component in the distribution of the control points in the defined state obtained by derivation (for example, the position in the surface of the facial component that overlaps with the control point in the center of the ear component) belongs, and at which position in the partial area it is configured. The information for determining which position in the partial area is configured can be the following information: the corresponding partial area in the surface of the facial component before adjustment is in Figure 11C When expanded in a two-dimensional plane as shown, in the circumscribed rectangle of the partial area (a parallelogram whose opposite sides represent the diagonal directions of the adjusted partial area), information is composed of the internal ratios (a, b) in the two side directions corresponding to the connection position 1102.

[0132] Furthermore, when the distribution of control points related to the derived facial component is adjusted by at least one of the translation amount and the rotation amount, it is sufficient to determine the position determined by the internal ratio (a, b) in two side directions in the circumscribed rectangle of the corresponding partial area in the curved surface related to the adjusted facial component, and change the distribution of the control points of the ear component in such a manner as to connect the ear component to the position.

[0133] This ensures that the arrangement relationship between some components is maintained, thereby reducing the burden on the designer involved in adjustments.

[0134] Composition Processing

[0135] Next, regarding the composition process of the composition device 200 of this embodiment using the derivation model to compose the expression model, Figure 12 The specific processing will be described with reference to the flowchart. The processing corresponding to the flowchart can be realized by the configuration control unit 201 reading a program related to the editing application recorded in the configuration recording medium 202, for example, and loading it into the configuration memory 203 for execution.

[0136] Furthermore, this structural processing is described with the assumption that it begins when editing operations related to the composition of a representation model are initiated for illustration data obtained by separating the object, which is the composition target, into two-dimensional images for each component, and defining a layer structure that represents the anteroposterior relationship between the components when depicted. Furthermore, in an editing application corresponding to this structural processing, in order to present the designer with the information required when constructing the representation model of the composition target object, the display control unit 205, for example, executes the following processing each time, according to the display update frequency of the display unit 220: displaying a two-dimensional image based on the illustration data or displaying a depiction representation generated by the depiction unit 204 for a specified state of the constructed representation model. In the following description, description of non-characteristic display control processing related to the description of the present invention is omitted.

[0137] In S1201, under the control of the composition control unit 201, the setting unit 206 sets the curved surfaces and control points in the reference state for the two-dimensional images of each component of the object being composed. Specifically, based on the operation input received via the operation input unit 210, the setting unit 206 sets the curved surfaces to which the two-dimensional images of each component are applied, as well as the control points used to control the deformation of the curved surfaces, for the two-dimensional images of the components included in the illustration data of the object being composed. Once the setting of the curved surfaces and control points in the reference state is complete, the setting unit 206 transmits the defined information to the composition unit 209, thereby composing the data for the representation model of the object being composed, which includes this information as the reference state information 1303.

[0138] In step S1202, the determination unit 207 determines a first feature quantity and a second feature quantity for the representation model of the object to be composed, which is composed in step S1201, under the control of the composition control unit 201. Specifically, the determination unit 207 determines the first feature quantity based on the texture information 1302 and the reference state information 1303 of the representation model, and determines the second feature quantity (initial value) based on average information related to the teacher data associated with the derivation model.

[0139] In S1203 , under the control of the composition control unit 201 , the derivation unit 208 takes the first feature value and the second feature value as input and uses the derivation model to derive the distribution of control points in the definition state of the representation model of the object to be composed.

[0140] In step S1204, under the control of the composition control unit 201, the composition unit 209 stores the information on the distribution of the control points in the definition state derived in step S1203 as the definition state information 1304 of the data of the expression model of the object constituting the object, thereby completing the composition of the data of the expression model.

[0141] In S1205, the rendering unit 204, under the control of the composition control unit 201, generates a rendering representation in the defined state based on the data of the constructed representation model. Then, under the control of the composition control unit 201, the display control unit 205 causes the display unit 220 to display the generated rendering representation in the defined state together with a GUI as an editing application related to the adjustment of each derived component.

[0142] In S1206, the composition control unit 201 determines whether at least one of the translation amount and the rotation amount related to the distribution of the control points in the defined state has been adjusted for any component of the object being composed. If the composition control unit 201 determines that the adjustment has been made, the process proceeds to S1207; if not, the process proceeds to S1208.

[0143] In S1207, under the control of the composition control unit 201, the composition unit 209 derives the distribution of the control points of the corresponding component after the change based on the adjusted translation and rotation amounts, and updates the definition status information 1304 of the data representing the model. Furthermore, if a component for which a configuration constraint has been set has been adjusted, the composition unit 209 also changes the distribution of the control points of the corresponding associated component accordingly and updates the definition status information 1304.

[0144] In S1208, the composition control unit 201 determines whether the editing operation related to the representation model of the object to be composed has been completed. If the composition control unit 201 determines that the editing operation has been completed, it stores the data of the representation model in the composition recording medium 202, completing the composition process. If it determines that the editing operation has not been completed, the process returns to S1206.

[0145] Thus, according to the composition processing of this embodiment, it is possible to construct an expression model of an object including composition objects in a definition state representing a desired rendering expression with a small number of working hours.

[0146] [Variation 1]

[0147] In the above embodiment, in order to reduce the amount of computation in the component device 200, derivation using the derivation model is performed only once when deriving the distribution of control points in an undefined state. However, the present invention is not limited to this. For example, it is also possible to derive the distribution of control points in the modified definition state using a second feature value based on information about the adjusted translation and rotation amounts for the corresponding component. In this case, it is believed that a more natural adjustment result can be obtained compared to a method that separates the control point distribution obtained through a single derivation into translational and non-translational components, adjusts and synthesizes these components separately.

[0148] [Variation 2]

[0149] Furthermore, in the above-described embodiment and modified examples, the second feature value is assigned when constructing the inference model (machine learning), but implementation of the present invention is not limited to this.

[0150] By assigning a second characteristic value, it is possible to more precisely classify and learn the deformations exhibited in the performance model being learned. This allows the constructed inference model to be used to obtain a distribution of control points in a defined state that more closely matches (and more accurately matches) the undefined performance model as a derivation result. This assumes that there are a sufficient number of samples of the performance model being learned. In other words, increasing the number of labels makes it easier to learn the performance model being learned based on different performances. However, this also reduces the number of performance models that can be learned for a single label combination. Therefore, to avoid the problem of overlearning, a larger absolute number of samples is preferable.

[0151] On the other hand, when the absolute number of samples is small, the more labels are added, the higher the probability of overlearning. As a result, it is possible that appropriate inference results cannot be obtained from the inference model. Therefore, when implementing the present invention, it is also possible to assign only the first feature value as a label to perform machine learning and thus construct the inference model. In the construction process, only the first feature value is also assigned to obtain the inference result.

[0152] [Variation 3]

[0153] In the above-described embodiment and modified examples, the size of the two-dimensional image is described as a component included in the first feature quantity, but implementation of the present invention is not limited thereto.

[0154] As described above, the deformation form of a component changes depending on the ratio of the size of the component's two-dimensional image to the size of the curved surface to which the two-dimensional image is applied. However, when constructing a representation model, if a curved surface (distribution of control points) in a reference state corresponding to the shape and size of each component's two-dimensional image is routinely defined, rather than defining an extremely large curved surface, the present invention can be implemented even without including the size of the component's two-dimensional image or the size of the curved surface to which the two-dimensional image is applied as the first feature value.

[0155] [Variation 4]

[0156] Furthermore, while the aforementioned embodiments and variations illustrate a method for constructing an inference model by performing machine learning on a representation model configured to enable rendering of a character's head rotating in the yaw direction, the embodiments of the present invention are not limited to this. Specifically, the representation model targeted for machine learning may also include a representation model configured to enable rendering of rotations in directions other than the yaw direction. For example, to achieve three-dimensional rendering of a character's head, a rendering may be defined for a rotation in the pitch direction in addition to the yaw direction. In this case, the inference model may be constructed separately for the deformation of the components in the yaw and pitch directions (one inference model for each dimension), or a single inference model may be constructed for the deformation of the components in a composite state (a inference model capable of inferring the distribution of control points associated with the two-dimensional deformation). Furthermore, as described above, the renderings achievable by the representation model are not limited to these rotational movements, so the specific representation or combination of representations for which the inference model is constructed can be modified as appropriate.

[0157] Furthermore, in the above-described embodiment and variations, the distribution of control points resulting from derivation is described as consisting of translational and rotational components in order to construct a rendering model that realizes the rendering of a character's head rotating in the yaw direction. However, implementation of the present invention is not limited to this. Specifically, the primary cause of component deformation varies depending on the rendering of the object, and thus deformations other than translation are not limited to rotations. Specifically, in a method that allows for the distribution of control points resulting from derivation to be separated and adjusted, the components to be adjusted can also be divided into translational and non-translational components.

[0158] In addition, in the above-mentioned embodiments and variations, a method is described in which a performance model including the same depiction performance (rotation performance in the Yaw direction) is used as teacher data when constructing a derivation model. However, in addition to this, for example, by using a performance model for the same purpose as teacher data, a performance model designed by the same designer as teacher data, or a performance model designed using the same performance method as teacher data, a more appropriate derivation result can be obtained.

[0159] In addition, while the aforementioned embodiments and variations describe the definition of the roles of the components separated in the representation model, the present invention is not limited to this. Defining the roles of the components enables effective machine learning and appropriate derivation, but even without defining the roles of the components, the invention can be implemented similarly by estimating the roles based on the configuration relationships of the components. In other words, in the present invention, the representation model that is the subject of machine learning does not necessarily require that the component structure be identical, or that the number and configuration of control points assigned to the surface be in a predetermined state. Similarly, the reference state of the object that constitutes the object for derivation of the defined state does not necessarily require that the component structure be specific, or that the number and configuration of control points assigned to the surface be in a predetermined state. In other words, as long as the representation model that specifies a similar depiction representation is constructed by learning the deformation of components (or component groups) that are considered to be the same, the distribution of control points representing the deformation can be transformed into a distribution with a specific fineness (resolution of representation) for learning. In contrast, when using the derivation results based on the derivation model, the derivation results can also be adjusted to match the fineness of the deformation that can be achieved by the control points set for the reference state of the object constituting the object, thereby defining the distribution of the control points related to the defined state.

[0160] [Other embodiments]

[0161] The present invention is not limited to the above-described embodiments, and various modifications and variations are possible within the scope of the invention. Furthermore, the construction device and the construction device of the present invention can also be implemented by a program that causes one or more computers to function as the device. The program can be recorded on a computer-readable recording medium or provided / distributed via a telecommunications line.

Claims

1. A method for constructing an inference model, comprising constructing an inference model for a representation model that achieves a depiction representation of an object corresponding to a state different from the reference state by deforming each component of a two-dimensional image of the object corresponding to a reference state, the inference model inferring the deformation of each component of the two-dimensional image in a defined state different from the reference state. It is characterized by: The representation model is configured to be defined by defining the deformation of at least one of the components of the two-dimensional image in the defined state, and to be capable of achieving a depiction representation corresponding to a state between the reference state and the defined state. The deformation of each component in the expression model is controlled by the distribution of control points set for the component. The derivation model construction method comprises: an acquisition step of acquiring, for the defined performance model, distributions of the control points in the reference state and the defined state; an extraction step of extracting a first feature value based on the distribution of control points in the reference state acquired in the acquisition step; and A learning process uses the first feature quantity extracted in the extraction process as a label, performs machine learning on the distribution of control points in the defined state obtained in the acquisition process, and constructs an inference model based on the learning results obtained by learning multiple defined performance models.

2. The method for constructing an inference model according to claim 1, wherein: Each component in the expression model is composed of a two-dimensional image of the component, a curved surface to which the two-dimensional image is applied, and the control points that determine the shape of the curved surface. The first feature value includes information on the center position and size of the curved surface associated with each component in the distribution of control points in the reference state.

3. The method for constructing an inference model according to claim 2, wherein: The first feature value further includes information on the size of a two-dimensional image applied to the curved surface of each component in the reference state.

4. The method for constructing an inference model according to claim 1, wherein: The derivation model construction method further includes a standardization process, in which the defined performance model is standardized. In the acquisition step, the expression model normalized in the normalization step is acquired.

5. The method for constructing an inference model according to claim 4, wherein: The object includes the character's head. The standardization process includes the following steps: normalizing the scales of distributions of the control points in the reference state and the defined state based on a distance between two components included in a first category of components constituting the character's head; as well as The deformation amount from the reference state to the defined state is standardized based on the amount of movement between the distribution of the control points set for the second category of components constituting the character's head in the reference state after the scale is standardized and the distribution of the control points in the defined state after the scale is standardized.

6. The method for constructing an inference model according to claim 1, wherein: The derivation model construction method further includes an estimation step in which a second feature value is estimated based on the distribution of the control points in the reference state and the distribution of the control points in the definition state obtained in the acquisition step. In the learning step, the second feature value estimated in the estimating step is further added to the label to perform machine learning.

7. The method for constructing an inference model according to claim 6, wherein: The second feature amount includes information indicating the amount of deformation of a translational component and the amount of deformation of a non-translational component related to the deformation from the reference state to the defined state.

8. The method for constructing an inference model according to claim 7, wherein: The object includes the character's head. The information representing the deformation amount of the translation component of the second feature value is estimated based on the amount of movement between the distribution of the control points in the reference state and the distribution of the control points in the definition state, for at least one control point set for a third category of components constituting the character's head. The information representing the deformation amount of the non-translational component of the second feature value is estimated based on the difference in movement amount between the distribution of the control points in the reference state and the distribution of the control points in the defined state for multiple control points set for the fourth category of components constituting the character's head.

9. The method for constructing an inference model according to claim 8, wherein: The fourth category of components is the components that express the concavity and convexity of the character's head. The deformation amount of the non-translational component is estimated based on the difference between the movement amounts of the concave portion and the convex portion of the character.

10. The method for constructing an inference model according to claim 8, wherein: The third category of components is a component representing a character's face.

11. An inference model construction device, comprising constructing an inference model for a representation model that achieves a depiction representation of an object corresponding to a state different from the reference state by deforming each component of a two-dimensional image of the object corresponding to the reference state, the inference model inferring the deformation of each component of the two-dimensional image in a defined state different from the reference state. It is characterized by: The representation model is configured to be defined by defining the deformation of at least one of the components of the two-dimensional image in the defined state, and to be capable of achieving a depiction representation corresponding to a state between the reference state and the defined state. The deformation of each component in the expression model is controlled by the distribution of control points set for the component. The derivation model construction device comprises: an acquiring unit configured to acquire, for the defined performance model, distributions of the control points in the reference state and the defined state; an extracting unit configured to extract a first feature value based on the distribution of the control points in the reference state acquired by the acquiring unit; as well as A learning unit uses the first feature extracted by the extraction unit as a label, performs machine learning on the distribution of control points in the defined state obtained by the acquisition unit, and constructs an inference model based on the learning results obtained by learning multiple defined performance models. 12 . A recording medium having a program recorded thereon for causing a computer to execute each step of the derivation model construction method according to claim 1 .

13. A recording medium having a program recorded thereon for causing a computer to construct a representation model using the inference model constructed by the inference model construction method according to claim 1, wherein the representation model deforms each component of a two-dimensional image corresponding to a reference state of the object to achieve a depiction representation of the object corresponding to a state different from the reference state. It is characterized by: The deformation of each component in the expression model is controlled by the distribution of control points set for the component. The program causes the computer to execute the following processing: Input processing, obtaining the distribution of the control points in the reference state for the object to be constituted; a first determination process of determining the first feature value based on the information obtained through the input process; a derivation process of deducing, using the derivation model, a distribution of the control points in the defined state of the object serving as the constituent object based on the first feature amount determined by the first determination process; as well as The output process constructs the representation model of the object to be constructed based on the derivation result of the derivation process and outputs the representation model.

14. A recording medium having a program recorded thereon for causing a computer to construct a representation model using the inference model constructed by the inference model construction method according to claim 6, wherein the representation model deforms each component of a two-dimensional image corresponding to a reference state of the object to achieve a depiction representation of the object corresponding to a state different from the reference state. It is characterized by: The deformation of each component in the expression model is controlled by the distribution of control points set for the component. The program causes the computer to execute the following processing: Input processing, obtaining the distribution of the control points in the reference state for the object to be constituted; a first determination process of determining the first feature value based on the information obtained through the input process; a second determining process of determining the second feature amount; a derivation process of derivation, using the derivation model, of derivation of a distribution of control points in the defined state of the object serving as the constituent object based on the first feature quantity determined by the first determination process and the second feature quantity determined by the second determination process; as well as The output process constructs the representation model of the object to be constructed based on the derivation result of the derivation process and outputs the representation model.

15. The recording medium according to claim 13 or 14, wherein The program further causes the computer to execute the following processing: a display control process for causing a display unit to display a depiction of the object serving as the constituent object corresponding to the defined state based on the representation model output by the output process; Accepting processing for accepting adjustment of a deformation amount of at least one of a translational component and a non-translational component of a deformation of the depiction corresponding to the defined state relative to the reference state; as well as a change process for changing the distribution of the control points in the defined state of the output representation model based on the adjusted deformation amounts of the translational component and the non-translational component when the adjustment is accepted by the acceptance process; When the adjustment is accepted by the acceptance process, the display unit is caused to display the drawn representation of the object serving as the constituent object corresponding to the defined state based on the representation model changed by the change process.

16. The recording medium according to claim 15, wherein The program further causes the computer to execute a separation process in which the distribution of the control points in the defined state derived by the derivation process is separated into a distribution of a translation component and a distribution of a non-translation component. In the change processing, the distribution of the control points in the defined state of the output expression model is changed to the distribution of the control points obtained by synthesizing the distribution of the translation component after the deformation amount of the translation component and the non-translation component after the adjustment and the distribution of the non-translation component.

17. The recording medium according to claim 16, wherein A constraint condition defining a configuration relationship for at least a portion of the components of the object constituting the object is defined. In the changing process, the arrangement position of the at least one component is changed so that the arrangement relationship of the at least one component is maintained before and after the adjustment.

18. A construction device for constructing a representation model using the inference model constructed by the inference model construction method according to claim 1, wherein the representation model is configured to achieve a depiction representation of an object corresponding to a state different from the reference state by deforming each component of a two-dimensional image of the object corresponding to the reference state. It is characterized by: The deformation of each component in the expression model is controlled by the distribution of control points set for the component. The constituent device comprises: an input unit for acquiring, for an object serving as a constituent object, a distribution of the control points in the reference state; a determination unit configured to determine the first feature value based on the information acquired by the input unit; a derivation unit that deduces a distribution of the control points in the defined state of the object serving as a constituent object based on the first feature amount determined by the determination unit using the derivation model; as well as An output unit constructs the representation model of the object as a composition target based on the deduction result of the derivation unit and outputs the representation model.

19. A construction device for constructing a representation model using the inference model constructed by the inference model construction method according to claim 6, wherein the representation model realizes a depiction representation of an object corresponding to a state different from the reference state by deforming each component of a two-dimensional image corresponding to the reference state of the object, It is characterized by: The deformation of each component in the expression model is controlled by the distribution of control points set for the component. The constituent device comprises: an input unit for acquiring, for an object serving as a constituent object, a distribution of the control points in the reference state; a first determining unit that determines the first feature value based on the information acquired by the input unit; a second determining unit configured to determine the second feature value; a derivation unit that deduces a distribution of the control points in the defined state of the object serving as the constituent object based on the first feature amount determined by the first determination unit and the second feature amount determined by the second determination unit using the derivation model; as well as An output unit constructs the representation model of the object as a composition target based on the deduction result of the derivation unit and outputs the representation model.

20. A method for constructing a representation model using the inference model constructed by the inference model construction method according to claim 1, wherein the representation model is configured to achieve a depiction representation of the object corresponding to a state different from the reference state by deforming each component of a two-dimensional image of the object corresponding to the reference state, It is characterized by: The deformation of each component in the expression model is controlled by the distribution of control points set for the component. The construction method comprises the following steps: An input step of obtaining, for an object to be constituted, a distribution of the control points in the reference state; a determining step of determining the first feature value based on the information acquired in the inputting step; a derivation step of deducing a distribution of the control points in the defined state of the object serving as the constituent object based on the first feature quantity determined in the determination step using the derivation model; as well as The output step constructs the representation model of the object to be constructed based on the derivation result in the derivation step and outputs the representation model.

21. A method for constructing a representation model using the inference model constructed by the inference model construction method according to claim 6, wherein the representation model deforms each component of a two-dimensional image of the object corresponding to a reference state to achieve a depiction representation of the object corresponding to a state different from the reference state, It is characterized by: The deformation of each component in the expression model is controlled by the distribution of control points set for the component. The construction method comprises the following steps: An input step of obtaining, for an object to be constituted, a distribution of the control points in the reference state; a first determining step of determining the first feature value based on the information acquired in the inputting step; a second determining step of determining the second feature value; a derivation step of deducing, using the derivation model, a distribution of the control points in the defined state of the object serving as the constituent object based on the first feature quantity determined in the first determination step and the second feature quantity determined in the second determination step; as well as The output step constructs the representation model of the object to be constructed based on the derivation result in the derivation step and outputs the representation model.

Citation Information

Patent Citations

  • Data structure for image formation and method of forming image

    JP2009104570A

  • Training data set generation apparatus and method for machine learning

    US20200151963A1