Information processing system, information processing method, and computer program

WO2026203824A1PCT designated stage Publication Date: 2026-10-01SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/003714
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-02-03
Publication Date
2026-10-01

Smart Images

  • Figure JP2026003714_01102026_PF_FP_ABST
    Figure JP2026003714_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides an information processing system that performs processing related to facial expression parameter value generation. This information processing system is provided with: an acquisition unit that acquires an input image comprising an illustration depicting a face of a character; an extraction unit that extracts facial features from the input image; and a conversion unit that converts the extracted features into a facial expression parameter value that includes position information for at least one facial feature point of an avatar. The extraction unit extracts different types of facial features on the basis of mutually different methods or algorithms. The conversion unit converts each type of feature into a facial expression parameter value, and consolidates the converted values to obtain the facial expression parameter value.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing System, Information Processing Method, and Computer Program

[0001] The technology disclosed in the present specification (hereinafter referred to as "the present disclosure") relates to an information processing system, an information processing method, and a computer program that perform processing related to facial expression generation for characters used in avatars, agents, and the like.

[0002] Dialogue systems based on large language models like ChatGPT have been developing, and with the incorporation of such technologies, the potential of agent systems that allow conversation with characters or digital humans displayed on screens is also increasing. By creating an agent system using characters from animations and manga, users can freely converse with the characters, which can provide new experiential value. When creating an agent system using characters from animations and manga (hereinafter also referred to as "anime-style"), in order to realize natural expression of anime-style facial expressions, it is necessary to prepare a large number of facial expression parameter values for expressing characteristic and exaggerated expressions compared to realistic human facial expressions.

[0003] Japanese Unexamined Patent Application Publication No. 2012-185624

[0004] An object of the present disclosure is to provide an information processing system, an information processing method, and a computer program that perform processing related to generation of facial expression parameter values used for expressing facial expressions of a character.

[0005] The present disclosure has been made in consideration of the above problem, and a first aspect thereof is an information processing system comprising: an acquisition unit that acquires an input image including an illustration in which a face of a character is drawn; an extraction unit that extracts facial feature amounts from the input image; and a conversion unit that converts the extracted feature amounts into facial expression parameter values including position information of at least one facial feature point of an avatar.

[0006] However, the term "system" as used here refers to a logical collection of multiple devices (or functional modules that perform specific functions), without regard to whether each device or functional module resides within a single enclosure. In other words, both a single device consisting of multiple components or functional modules, and a collection of multiple devices, qualify as a "system."

[0007] The extraction unit includes a plurality of extraction units that extract different types of facial feature quantities based on different methods or algorithms, and the conversion unit includes a plurality of conversion units that convert the feature quantities of the type corresponding to each of the plurality of extraction units into facial expression parameter values. The information processing system relating to the first aspect further includes an integration unit that integrates the facial expression parameter values ​​output from the plurality of conversion units.

[0008] Specifically, the plurality of extraction units include a tag extraction unit that extracts tags related to facial expressions from the input image, a face feature point extraction unit that extracts positional information of facial feature points from the input image, and an image analysis unit that generates answers to questions regarding the input image. The plurality of conversion units include a tag-facial expression parameter conversion unit that converts the tags extracted by the tag extraction unit into facial expression parameter values, a feature point-facial expression parameter conversion unit that converts the positional information of facial feature points extracted by the face feature point extraction unit into facial expression parameter values, and an answer-facial expression parameter conversion unit that converts the answers generated by the image analysis unit into facial expression parameter values.

[0009] Furthermore, a second aspect of this disclosure is an information processing method comprising: an acquisition step of acquiring an input image consisting of an illustration depicting a character's face; an extraction step of extracting facial feature quantities from the input image; and a conversion step of converting the extracted feature quantities into an expression parameter value that includes positional information of at least one facial feature point of the avatar.

[0010] Furthermore, a third aspect of this disclosure is a computer program written in a computer-readable format to cause a computer to function as: an acquisition unit that acquires an input image consisting of an illustration depicting a character's face; an extraction unit that extracts facial feature quantities from the input image; and a conversion unit that converts the extracted feature quantities into facial parameter values ​​that include positional information of at least one facial feature point of an avatar.

[0011] The computer program relating to the third aspect of this disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes via a storage medium, communication medium, such as an optical disk, magnetic disk, semiconductor memory, or a network, in a computer-readable format. By installing the computer program relating to the third aspect of this disclosure onto a computer via any of these media, collaborative effects can be achieved on the computer, similar to the effects of the information processing system relating to the first aspect of this disclosure.

[0012] Figure 1 shows the functional configuration of the information processing system 10. Figure 2 shows the functional configuration of the information processing system 20. Figure 3 shows the functional configuration of the information processing system 100. Figure 4 shows the execution procedure for extracting facial expression parameter values ​​in the information processing system 100. Figure 5 shows how metadata is associated with the facial expression parameter values ​​that the information processing system ultimately outputs. Figure 6 shows an example of the hardware configuration of the information processing device. Figure 7 shows an example of a character's face with distinctive or exaggerated expressions. Figure 8 shows an example of visualizing the results of facial feature point extraction superimposed on an input image. Figure 9 shows an example of applying facial expression parameter values ​​to an avatar.

[0013] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.

[0014] A. Overview B. Basic Configuration C. Examples C-1. System Configuration C-2. System Operation C-3. Effects D. System Details D-1. Input Image D-2. Tag Extraction D-3. Facial Feature Extraction D-4. VLM Image Analysis D-5. Tag / Expression Parameter Conversion D-6. Feature Point / Expression Parameter Conversion D-7. Response / Expression Parameter Conversion D-8. Information Integration D-9. Expression Parameter Values ​​D-10. Input Assistance Information D-11. Colorization of Black and White Images E. Examples of Using Expression Parameter Values ​​F. Hardware Configuration of Information Processing Device G. Effects

[0015] A. Overview: In an agent system using anime-style characters, expressing a wide range of facial expressions for the avatar requires preparing numerous facial expression parameter values. Methods such as manually preparing these parameters or utilizing existing facial expression datasets are readily conceivable.

[0016] However, when preparing the facial expression parameter values ​​manually, the problem is that the avatar's facial expression parameter values ​​consist of dozens of parameters that change the shape of each part of the face, so creating each one manually is extremely time-consuming and costly.

[0017] Furthermore, when using existing facial expression datasets, many of the existing facial expression parameter values ​​are based on data obtained from real human faces. For example, a three-dimensional face model data generation device has been proposed that extracts the region corresponding to the face and the coordinates of control points necessary to express facial expressions as facial expression data from captured human footage (see Patent Document 1). However, when facial expression data obtained from real human faces is applied to anime-style avatars, problems arise such as the facial expression changes appearing minimal and the inability to express the exaggerated expressions characteristic of anime-style characters. In order to achieve natural expression of anime-style facial expressions, it is necessary to prepare a large number of facial expression parameter values ​​that are characteristic of facial expressions and exaggerated expressions compared to realistic human facial expressions.

[0018] Therefore, this disclosure proposes a technology for automatically creating facial expression parameter values, including exaggerated expressions, using illustrations and the like. With illustrations, it is relatively easy to collect many character faces with distinctive or exaggerated expressions compared to realistic human expressions, such as shocked faces (vertical lines on the forehead when shocked), sweat, blushing (a face with flushed cheeks), surprised (OO) eyes, and inequality (><) eyes, as exemplified in Figure 7. Then, in this disclosure, features are extracted from the faces of characters with various expressions depicted in these illustrations, for example, using a visual model, and then converted into facial expression parameter values. According to this disclosure, by extracting expressions from the faces of characters depicted in illustrations such as anime and manga, and converting them into facial expression parameter values ​​that can be applied to express the expressions of avatars, it is possible to automatically create facial expression parameter values ​​that can be applied to a large number of anime-style avatars.

[0019] The "expression parameter values" discussed in this disclosure are assumed to be a collection of multiple numerical values, including the positional information of at least one facial feature point of the avatar, but other representation methods are also acceptable. For example, like blend shapes, each value constituting the expression parameter value represents the degree to which a change in some facial part is reflected, and the vertex positions of the facial parts of the character model are changed according to the value. When reproducing exaggerated expressions, the exaggerated expression must already be implemented in the avatar, and its presence or absence must be changeable using the expression parameter values.

[0020] B. Basic Configuration First, the basic configuration of the information processing system relating to this disclosure will be explained. Figure 1 shows a schematic representation of the functional configuration of the information processing system 10 relating to this disclosure. The illustrated information processing system 10 comprises a feature extraction unit 11 and a feature / expression parameter conversion unit 12.

[0021] The input image to the information processing system 10 is an illustration depicting a character's face, and it is preferable if it also includes corresponding dialogue text, similar to a comic strip.

[0022] The feature extraction unit 11 extracts the facial features of the character from the input image. The feature extraction unit 11 basically uses a pre-trained visual model to extract the facial features of the character, but of course, other methods may also be used to extract the facial features of the character.

[0023] The feature / expression parameter conversion unit 12 converts the character's facial features extracted by the feature extraction unit 11 into expression parameter values ​​that can be applied to express the avatar's facial expressions. The expression parameter values ​​generated by the feature / expression parameter conversion unit 12 are a collection of multiple numerical values ​​that include the position information of at least one facial feature point of the avatar. The position information of the facial feature point is geometric data of the facial parts. Facial parts include eyes, eyebrows, and mouths, and the geometric data includes coordinate information of specific points such as the corners of the eyes and the edges of the mouth, as well as data indicating the movement and deformation of the facial parts. The expression parameter values ​​are in the form of, for example, blend shapes, where multiple numerical values ​​each represent the degree to which a change in some facial part is reflected, and the vertex positions of the facial parts of the character model are changed according to the values. The feature / expression parameter conversion unit 12 may convert features into expression parameter values ​​based on rules obtained from, for example, empirical rules, or it may convert features into expression parameter values ​​using a trained model.

[0024] The facial expression parameter values ​​generated by the feature / facial expression parameter conversion unit 12 may include text data such as dialogue included in the input image illustration or tags indicating the corresponding facial expression, or such text data may be stored in association with the facial expression parameter values.

[0025] Figure 2 shows the functional configuration of the modified information processing system 20. It is common to the information processing system 10 shown in Figure 1 in that it extracts facial features from an input image consisting of an illustration of a character's face and converts the features into facial expression parameter values. The information processing system 20 shown in Figure 2 is configured to output one facial expression parameter value (such as a blend shape) from one input image (an illustration of a character's face). The information processing system 20 consists of N feature extraction units, namely the first feature extraction unit 21-1, the second feature extraction unit 21-2, ..., and the Nth feature extraction unit 21-N, and N feature / facial expression parameter conversion units 22, namely the first feature / facial expression parameter conversion unit 22-1, the second feature / facial expression parameter conversion unit 22-2, ..., and the Nth feature / facial expression parameter conversion unit 22-N, which are provided in correspondence with each feature extraction unit.

[0026] The first feature extraction unit 21-1, the second feature extraction unit 21-2, ..., and the Nth feature extraction unit 21-N each extract different types of facial features from the same input image using different visual models or following different algorithms. That is, the first feature extraction unit 21-1 extracts the first type of facial features, the second feature extraction unit 21-2 extracts the second type of facial features, ..., and the Nth feature extraction unit 21-N extracts the Nth type of facial features.

[0027] The first feature / expression parameter conversion unit 22-1, the second feature / expression parameter conversion unit 22-2, ..., the Nth feature / expression parameter conversion unit 22-N each convert the character's facial features output from the corresponding feature extraction unit into facial expression parameter values ​​based on their respective rules. That is, the first feature / expression parameter conversion unit 22-1 converts the first type of facial feature into a first facial expression parameter value based on the first rule, the second feature / expression parameter conversion unit 22-2 converts the second type of facial feature into a second facial expression parameter value based on the second rule, ..., the Nth feature / expression parameter conversion unit 22-N converts the Nth type of facial feature into an Nth facial expression parameter value based on the Nth rule.

[0028] The integration unit 23 then integrates the first, second, ..., and Nth facial expression parameter values ​​output from the first feature / facial expression parameter conversion unit 22-1, the second feature / facial expression parameter conversion unit 22-2, ..., and Nth feature / facial expression parameter conversion unit 22-N into a single facial expression parameter value. The integration unit 23 may simply take the arithmetic mean of these multiple facial expression parameter values ​​to integrate them, or it may preferentially adopt the facial expression parameter value obtained from the facial feature with the higher priority based on the priority order among the different types of facial features, or it may perform a weighted average based on priority to integrate them into a single facial expression parameter value.

[0029] The facial expression parameter values ​​integrated by the integration unit 23 may include text data such as dialogue included in the input image illustration or tags indicating the corresponding facial expression, or such text data may be stored in association with the facial expression parameter values.

[0030] By extracting features from the character's face in the illustration using multiple methods and integrating these results, it is expected that more accurate facial expression parameter values ​​can be obtained than when relying on features obtained from a single method. Furthermore, even if there are facial feature points for which sufficient information cannot be obtained using one feature extraction method, this can be supplemented using information obtained from another method.

[0031] C. Example C-1. Example Figure 2 shows a basic system configuration according to the present disclosure, in which different types of facial features are extracted from the same input image using different visual models or according to different algorithms, and the expression parameter values ​​obtained by converting each facial feature are integrated into one. Figure 3 shows an information processing system 100 as an embodiment of the present disclosure, which is configured to extract three types of facial features—tags, facial feature points, and expressions or emotions—from an input image (an illustration of a character's face), convert each type of facial feature into an expression parameter value, and then integrate these three expression parameter values.

[0032] The information processing system 100 shown in Figure 3 includes a tag extraction unit 111, a facial feature point extraction unit 121, and a VLM (Visual Language Model) image analysis unit 131 as feature extraction units. It also includes three functional modules corresponding to each feature extraction unit: a tag / facial expression parameter conversion unit 112, a feature point / facial expression parameter conversion unit 122, and a response / facial expression parameter conversion unit 132. Furthermore, it includes an information integration unit 140 that integrates the outputs of these functional modules, enabling the extraction of facial expression parameter values ​​for an anime-style avatar from an input image.

[0033] The input to the information processing system 100 includes an illustration of a character's face and, as an optional item, text related to the face image. The information processing system 100 is configured to use illustrations obtained from comics and the like, and allows the inclusion of dialogue corresponding to the character's face as an optional input item.

[0034] The tag extraction unit 111 extracts tags related to facial expressions from the input image. The tag-to-facial expression parameter conversion unit 112 then converts the tags extracted from the input image by the tag extraction unit 111 into facial expression parameter values ​​that can be applied to the avatar.

[0035] The facial feature point extraction unit 121 extracts positional information of facial feature points from the input image. Then, the feature point / expression parameter conversion unit 122 converts the facial feature points extracted from the input image by the facial feature point extraction unit 121 into expression parameter values ​​that can be applied to the avatar.

[0036] The VLM image analysis unit 131 analyzes the input image and generates answers to questions about the facial expressions and emotions of the character contained in the input image. The answer / expression parameter conversion unit 132 then converts the answers to the expression and emotion questions extracted by the VLM image analysis unit 131 from the input image into expression parameter values ​​that can be applied to the avatar. By further inputting text related to the input image (for example, the character's lines), it is expected that the accuracy of the answers from the VLM image analysis unit 131 will improve.

[0037] The information integration unit 140 integrates the expression parameter values ​​output from the tag / expression parameter conversion unit 112, the feature point / expression parameter conversion unit 122, and the response / expression parameter conversion unit 132 to determine all expression parameter values ​​and output them as expression parameter values ​​that can be applied to the avatar. The output of the information processing system 100 is an expression parameter value consisting of several dozen numerical values, specifically a blend shape that includes the positional information of several dozen facial feature points. Furthermore, if the information integration unit 140 determines that it cannot extract expressions from the input image (illustration, etc.) (i.e., cannot output expression parameter values), it may stop processing that input image.

[0038] Furthermore, if the input image, such as an illustration, contains text information such as lines spoken by the character in the illustration, past lines, or the situation, the information processing system 100 may input such text information along with the illustration. In addition, the information processing system 100 may input not only the illustration of the face, but also the character's whole body or other images representing the context at the same time as the illustration. In such cases, the VLM image analysis unit 131 can further consider this text information and contextual images to extract the facial expressions and emotions of the character depicted in the illustration with higher accuracy.

[0039] Furthermore, the information processing system 100 may also include a coloring unit 150 that colors the input image if it is black and white, converting it into a color image. The facial feature point extraction unit 121 can extract facial feature points with higher accuracy by using a color image colored by the coloring unit 150 rather than a black and white input image.

[0040] The coloring unit 150 may use the tags extracted from the input image by the tag extraction unit 111 to color the black and white image with higher accuracy. If the coloring unit 150 uses an image generation AI, the tags extracted from the input image by the tag extraction unit 111 should be input to the image generation AI as a prompt.

[0041] The facial expression parameter values generated by the information processing system 100 from various input images are stored in the facial expression parameter value database 160. The facial expression parameter values stored in the facial expression parameter value database 160 can later be applied to express the facial expression of an avatar. The facial expression parameter values stored in the facial expression parameter value database 160 are stored with text information input together with the input image associated as metadata (see FIG. 5). Alternatively, metadata may be included in the facial expression parameter values and stored in the facial expression parameter value database 160. Specifically, the text information input together with the input image is dialogue included in the illustration that is the input image. In addition, tags extracted from the input image by the tag extraction unit 111, and answer text generated from the input image by the VLM image analysis unit 131 (information such as facial expressions and emotions obtained through image analysis) may also be included in the metadata and stored.

[0042] C-2. System Operation Next, the operation of the information processing system 100 will be described. FIG. 4 shows an execution procedure of processing for extracting facial expression parameter values from a face image and dialogue corresponding to the face image in the information processing system 100.

[0043] First, the information processing system 100 inputs an illustration image from which it is desired to extract the character's facial expression (step S401).

[0044] The tag extraction unit 111 extracts a tag from the input image (step S402). Then, the tag-facial expression parameter conversion unit 112 converts the tag into a facial expression parameter value that can be applied to the avatar (step S403).

[0045] Further, the facial feature point extraction unit 121 extracts facial feature points from the input image (step S404), and the feature point-facial expression parameter conversion unit 122 predicts the position and shape of each part from the facial feature points, and converts the result into a facial expression parameter value that can be applied to the avatar (step S405).

[0046] Furthermore, the VLM image analysis unit 131 extracts facial expressions and emotions from the input image (step S406), and the response / facial expression parameter conversion unit 132 converts the facial expressions and emotions of the face image output from the VLM image analysis unit 131 into facial expression parameter values applicable to an avatar (step S407).

[0047] The facial expression parameter value acquisition processing performed by each of the tag / facial expression parameter conversion unit 112, the feature point / facial expression parameter conversion unit 122, and the response / facial expression parameter conversion unit 132 may be performed chronologically or concurrently in parallel. Furthermore, any order may be adopted when performed chronologically.

[0048] When text information such as lines uttered by the character in the input image, past lines, and situations exists in the input image, the information processing system 100 also inputs such text information together (step S410). Furthermore, the information processing system 100 may be configured to simultaneously input not only the face image but also an image representing the whole body of the character and other contexts. In such a case, in step S406, the VLM image analysis unit 131 can further consider these pieces of text information and the image representing the context to extract facial expressions and emotions with higher accuracy.

[0049] Furthermore, the coloring unit 150 may be configured to color a monochrome input image and convert it into a color image (step S411). In step S411, the coloring unit 150 can color the monochrome image with higher accuracy by using the tags extracted from the input image by the tag extraction unit 111 in step S402. Then, in step S404, the facial feature point extraction unit 121 can extract facial feature points with higher accuracy by using the color image colored by the coloring unit 150 rather than the monochrome input image.

[0050] Then, the information integration unit 140 integrates all facial expression parameter values output from each module of the tag / facial expression parameter conversion unit 112, the feature point / facial expression parameter conversion unit 122, and the response / facial expression parameter conversion unit 132, outputs the integrated values as facial expression parameter values relating to the facial expression of the avatar (step S408), and ends the processing of the illustration image input in step S401.

[0051] The facial expression parameter values ​​output from the information processing system 100 are stored in the facial expression parameter value database 160. The facial expression parameter values ​​stored in the facial expression parameter value database 160 can later be applied to display the facial expressions of the avatar. The text information entered along with the input image may be linked to and stored as metadata for the facial expression parameter values ​​stored in the facial expression parameter value database 160 (see Figure 5). Alternatively, metadata may be included in the facial expression parameter values ​​and stored in the facial expression parameter value database 160.

[0052] Furthermore, if the information integration unit 140 determines that it cannot extract facial expressions from the input image (such as an illustration), it will stop processing that input image.

[0053] C-3. Effects: The effects and benefits brought about by the configuration of the information processing system 100 shown in Figure 3 and the system operation shown in Figure 4 will be explained.

[0054] Tag extraction and conversion of tags to facial expression parameter values: In step S402, the tag extraction unit 111 extracts tags related to facial expressions. Tags related to facial expressions include, for example, "smile", "open_mouth", and "closed_eyes". The tags extracted by the tag extraction unit 111 also include illustrative exaggerations. For example, "blush" and ">_<" (eyes in the shape of inequality signs (>_<)). Then, in step S403, the tag-to-facial expression parameter conversion unit 112 converts tags into facial expression parameter values ​​that can be applied to the avatar, based on the rules for converting tags to facial expression parameter values.

[0055] The tag extraction unit 111 utilizes a pre-trained model learned from illustrations, making it easier to extract characteristic elements from illustrations compared to the face feature point extraction unit 121 and the VLM image analysis unit 131. While the face feature point extraction unit 121 may fail to extract feature points properly due to exaggerated expressions, the tag extraction unit 111 can extract tags relatively stably. Compared to the VLM image analysis unit 131, the tag-to-expression parameter conversion unit 112 can extract tags with high accuracy. However, it is not possible to extract expression elements for which no tags exist.

[0056] Regarding facial feature point extraction and conversion of feature points to expression parameter values: In step S404, the facial feature point extraction unit 121 obtains positional information of each feature point from the input facial image, such as the start and end points of the eyebrows, the positions of the inner corner of the eye, the outer corner of the eye, the upper eyelid, the lower eyelid, and the corners of the mouth, the upper lip, and the lower lip. Then, in step S405, the feature point / expression parameter conversion unit 122 predicts the position and shape of each part based on the positional information of the facial feature points obtained in step S404, and obtains expression parameter values ​​that can be applied to the avatar based on the rules for converting each feature point to an expression parameter value.

[0057] The facial feature point extraction unit 121 has the effect of being able to obtain the position, angle, and shape of each part of the face by using a face detection model learned from illustrations. On the other hand, it is difficult to obtain the position, angle, and shape of each part of the face by extracting tags with the tag extraction unit 111 or by extracting expressions and emotions with the VLM image analysis unit 131.

[0058] Regarding VLM image analysis and conversion of facial expression parameter values ​​from image analysis results: In step S406, an answer is obtained by inputting an image and a question into the VLM image analysis unit 131 (details of the question and answer will be described later). In step S407, the answer / facial expression parameter conversion unit 132 can obtain facial expression parameter values ​​that can be applied to the avatar based on the rules for converting the obtained answer into facial expression parameter values. If the information integration unit 140 determines that the predictions of the tag / facial expression parameter conversion unit 112 and the feature point / facial expression parameter conversion unit 122 are insufficient, it is also possible to input an additional question into the VLM image analysis unit 131.

[0059] Regarding the integration of conversion results: In step S408, the information integration unit 140 integrates the expression parameter values ​​output from each module: the tag / expression parameter conversion unit 112, the feature point / expression parameter conversion unit 122, and the response / expression parameter conversion unit 132, to determine all expression parameter values.

[0060] If multiple modules predict the same facial expression parameter values, the information integration unit 140 determines priority and adopts the output of one of the modules. Furthermore, if a module makes an insufficient prediction, the information integration unit 140 takes measures such as having that module re-predict the facial expression parameter values, utilizing the prediction results of other modules, or applying default facial expression parameter values. For example, if the information integration unit 140 determines that the predictions of the tag-facial expression parameter conversion unit 112 and the feature point-facial expression parameter conversion unit 122 are insufficient, it can also input additional question text into the VLM image analysis unit 131.

[0061] Furthermore, if the information integration unit 140 determines that it cannot extract facial expressions from the input image (such as an illustration), it may stop processing that input image.

[0062] Regarding the input of supplementary information: If there is text information related to the input image, such as lines spoken by the character in the input image, past lines, or the situation, then in step S410, this text information is also input. By further considering this text information and the image representing the context, the VLM image analysis unit 131 can generate answers to questions about the facial expressions and emotions of the character in the input image with higher accuracy.

[0063] Regarding the coloring of the black and white input image: In step S411, the coloring unit 150 colors the black and white input image and converts it into a color image. As a result, in step S404, the face feature point extraction unit 121 can extract face feature points with higher accuracy. When the coloring unit 150 infers a color image from the black and white input image, it may also use the tags extracted from the input image by the tag extraction unit 111.

[0064] The effects described herein are illustrative only, and the effects brought about by this disclosure are not limited to those described herein. Furthermore, this disclosure may produce additional effects beyond those described above. Further objectives, features, and advantages of this disclosure will become apparent from a more detailed description based on the embodiments and accompanying drawings described herein.

[0065] D. System Details Section D describes the details of each module and other components of the information processing system 100 according to one embodiment of the present disclosure.

[0066] D-1. The input image information processing system 100 receives an illustration image containing a face from which facial expressions are to be extracted. Basically, it assumes an image where the face is prominently depicted in the center. For example, an image that has been edited beforehand, such as by cropping, to ensure the face is prominently displayed in the center. Alternatively, if the image is obtained from a dataset, the system may automatically extract the face portion if it has annotations indicating the face area. Alternatively, an existing system for extracting the faces of characters in illustrations may be used.

[0067] In this embodiment, live-action images are not assumed as input images. Various techniques already exist for extracting facial expressions from live-action images of humans. In contrast, the input image to the information processing system 100 may be a generated image such as a comic strip panel, an anime scene, or a computer-generated image.

[0068] There are no explicit restrictions on the style of the illustration used as input image, but the accuracy of extracting tags and facial features will vary depending on the style.

[0069] The input image can also be an image frame extracted from a video.

[0070] Even if the input image has been colored by the coloring unit 150, it may be used as a black and white image in the face feature point extraction unit 121, etc.

[0071] The input image format is not particularly limited. For example, image formats such as JPEG and PNG are acceptable.

[0072] There are no restrictions on the resolution or aspect ratio of the input image. However, it must have a resolution sufficient to recognize facial expressions.

[0073] By extracting each frame from the video and inputting them as a series of images, the video can be used as input images for the information processing system 100.

[0074] Although not shown in Figure 3, the information processing system 100 may further include a preprocessing unit that preprocesses the input image to satisfy the above requirements.

[0075] D-2. Tag Extraction The tag extraction unit 111 extracts text information (tags) about elements within the input image, which is an illustration. In this embodiment, the tag extraction unit 111 is expected to primarily extract elements related to facial expressions.

[0076] The tag extraction unit 111 can extract tags from an illustration image, for example, using Stable Diffusion Tagger. Stable Diffusion Tagger is an extension of the Stable Diffusion model, which is one of the visual models, and is known to be able to analyze the content of an image and extract features and attributes as tags. If existing technologies such as Stable Diffusion Tagger are used, only fixed, known tags will be output, and unknown tags will not be output. Of course, if known tags can be extracted from the illustration, technologies other than Stable Diffusion Tagger may be applied to the tag extraction unit 111.

[0077] For example, if an illustration of a smiling face is input to the tag extraction unit 111, known tags such as "smile" can be obtained. Other examples of tags that can be output include "closed_eyes", "wide-eyeed", "one_eye_closed", "closed_mouth", "open_mouth", "v-shaped_eyebrows", "sweatdrop", "brush", "anger_vein", "frown", "o_o", "^_^", and ">_<". Note that the tags extracted by the tag extraction unit 111 may be in languages ​​other than English.

[0078] For each tag output from visual models such as the Stable Diffusion Tagger, a numerical value such as a probability may also be output. The tag extraction unit 111 can ignore elements with low probabilities according to these probabilities.

[0079] The tag extraction unit 111 can also use a model trained on illustrations. By using such a model, it is possible to extract elements characteristic of illustrations, such as exaggerated expressions like sweat and blushing. Furthermore, by using a model trained on various illustrations, tag extraction can be performed in response to various styles of illustrations.

[0080] The faces in the images input to the tag extraction unit 111 may be facing forward or in profile. Furthermore, the images may be either black and white or color. If the input image is black and white, it can be converted to a color image by the coloring unit 150 (as described above).

[0081] D-3. Face Feature Extraction The face feature point extraction unit 121 is a functional module for extracting the positions of feature points such as eyebrows, eyes, mouth, and contours from the face of a character depicted in an illustration. When an image consisting of an illustration is input, it outputs the pixel position of each part within the image. Figure 8 shows an example in which the output information of the face feature point extraction unit 121 is superimposed on the input image and visualized.

[0082] The facial feature point extraction unit 121 is intended to use an existing visual model, such as AnimeFaceDetector, which is specialized for recognizing the faces of anime characters. AnimeFaceDetector is widely used to detect faces and extract feature points from anime images. Since AnimeFaceDetector improves the recognition accuracy of color images, if the input image is black and white, it uses a color image colored by the coloring unit 150. Of course, the facial feature point extraction unit 121 may also perform facial feature point extraction on a black and white image.

[0083] If it is possible to output the positions of each part from the image, other technologies (visual models) besides AnimeFaceDetector may be used in the facial feature point extraction unit 121. There are many systems that analyze realistic human face images. In contrast, it should be noted that the facial feature point extraction unit 121 is designed specifically for extracting facial feature points from illustrations.

[0084] For each position output from models such as AnimeFaceDetector, a numerical value such as a probability may also be output. The face feature point extraction unit 121 can ignore elements with low probabilities according to these probabilities. In this case, the output results of other functional modules can be used to compensate.

[0085] D-4. VLM Image Analysis The VLM image analysis unit 131 is a functional module that outputs a text response when a question and an image such as an illustration are input in natural language. VLM is a visual language model that can understand by integrating visual and text information, and can answer the question in text when an image and a text question about that image are input. The natural language of the question can be in English, Japanese, or any other language, but the accuracy may differ depending on the language. Examples of natural language questions and answers are given below.

[0086] Example 1: Does the character in the image have a smiling mouth? Answer in yes / no. Example answer: "Yes" Example 2: Which eyes is the character winning? Answer with one word, right or left. Example answer: "right" Example 3: Is there more than one sweat drop? Answer in yes / no. Example answer: "Yes" Example 4: Does the character in the image have diagonal lines on their cheeks to show their Embarassment? Answer in yes / no. The character says "{dialogue}" Example answer: "Yes"

[0087] As shown in Example 4, context can be presented by including dialogue in the question. Therefore, the VLM image analysis unit 131 can improve the accuracy of image analysis by utilizing the context.

[0088] The question sentences exemplified above can be provided to the VLM image analysis unit 131 as input prompts. Therefore, by changing the input prompts, it is possible to add or modify the question sentences provided to the VLM image analysis unit 131. The information integration unit 140 may also add a question sentence to the input prompt if it determines that the predictions of the tag / expression parameter conversion unit 112 and the feature point / expression parameter conversion unit 122 are insufficient. The question sentences may be created manually or automatically generated using a generative model such as a large-scale language model.

[0089] The VLM image analysis unit 131 is designed to use a large-scale language model capable of image input, such as GPT-4o.

[0090] The VLM image analysis unit 131 returns an answer selected from a predetermined set of options in response to a question, for processing by the subsequent answer / expression parameter conversion unit 132. For example, the question is adjusted to return either "yes" or "no". It is also possible to provide multiple options for the question. The adjustment of the question can be done through internal processing of the VLM image analysis unit 131, or by changing the input prompt given to the VLM image analysis unit 131.

[0091] If the accuracy of the VLM image analysis unit 131 is sufficiently high, the tag extraction unit 111 and the face feature extraction unit 122 are unnecessary. In other words, if questions corresponding to tag extraction and face feature extraction are prepared, the VLM image analysis unit 131 alone can process both tag extraction and face feature extraction.

[0092] D-5. Tag and Expression Parameter Conversion The tag and expression parameter conversion unit 112 converts the tags obtained by the tag extraction unit 111 into 3D character expression parameter values ​​that can be applied to the avatar. The tag and expression parameter conversion unit 112 takes zero or more tags as input from the tag extraction unit 111 and outputs expression parameter values ​​consisting of several tens of decimal values.

[0093] The tag / expression parameter conversion unit 112 pre-determines the processing for each tag and performs the processing when a tag exists.

[0094] For example, if the tag "closed_eyes" is obtained from a certain illustration, the tag / expression parameter conversion unit 112 sets the value of the parameter for closing the eyelids of the avatar to which it is applied to the maximum value, so that the pupils are hidden. In other cases, the tag / expression parameter conversion unit 112 itself does not determine the eyelid parameters, but instead has the feature point / expression parameter conversion unit 122 or the response / expression parameter conversion unit 132 determine the eyelid parameters. In this way, the text information of the tags can be changed to the behavior of the avatar.

[0095] Furthermore, for example, if a tag like ">_<" (eyes in the shape of ><) is obtained for a certain illustration, the tag / expression parameter conversion unit 112 sets the value of the parameter that enables ><-shaped eyes (inequality eyes) for the avatar to be applied to the tag to its maximum value, making the avatar's eyes inequality eyes. Even if this tag is not obtained, the avatar's eyes may be changed by other tags (such as surprised (OO) eyes). If there are no other tags related to eyes, the tag / expression parameter conversion unit 112 itself does not determine the eye parameters, but instead lets the feature point / expression parameter conversion unit 122 or the response / expression parameter conversion unit 132 determine the eye parameters. In this way, the text information of the tags can be changed to the avatar's behavior.

[0096] Each tag obtained by the tag extraction unit 111 may also have a numerical value, such as a probability. In this case, the tag-to-expression parameter conversion unit 112 can also decide whether or not to utilize each tag based on that numerical value.

[0097] The tag extraction unit 111 may obtain conflicting tags. By determining a priority among the tags, the tag-expression parameter conversion unit 112 prevents the output information from becoming corrupted. The priority can be determined in advance, or it can be determined by a numerical value such as probability.

[0098] In some cases, the tag extraction unit 111 may input zero tags for an input image. In this case, the tag-to-expression parameter conversion unit 112 will not output anything. Also, if an unknown tag is input from the tag extraction unit 111, the tag-to-expression parameter conversion unit 112 will not process that tag.

[0099] The tag-expression parameter conversion unit 112 may use a "tag-expression parameter conversion dictionary" that stores rules for converting tags into expression parameter values ​​for avatars to convert the tags obtained by the tag extraction unit 111 into expression parameter values. Table 1 below shows an example of a tag-expression parameter conversion dictionary. This tag-expression parameter conversion dictionary uses known tags output from the tag extraction unit 111 as headwords and shows conversion rules such as the avatar's facial feature points to be manipulated, their geometric information (position and displacement), and priority corresponding to each tag.

[0100]

[0101] The conversion rules stored in the tag-expression parameter conversion dictionary may be derived from empirical rules or generated using a pre-trained model learned from illustrations. Alternatively, the tag-expression parameter conversion unit 112 may use a pre-trained model learned from illustrations, rather than the tag-expression parameter conversion dictionary described above, to convert tags to expression parameter values.

[0102] D-6. Feature Point / Expression Parameter Conversion The feature point / expression parameter conversion unit 122 converts the positions of each feature point of the character's face drawn in the illustration, output from the face feature point extraction unit 121, into 3D character expression parameter values ​​that can be applied to the avatar.

[0103] The feature point / expression parameter conversion unit 122 pre-determines the processing for each vertex position and then performs the processing. For example, it determines whether the eyebrows are raised or drooping based on the height information of the eyebrow tail and eyebrow head, and determines specific expression parameter values ​​related to the eyebrows.

[0104] Each position output from the facial feature point extraction unit 121 may also have a numerical value, such as a probability. In this case, the feature point / expression parameter conversion unit 122 can decide whether or not to utilize that vertex information based on that numerical value. If it does not utilize it, it will prioritize the results of the tag / expression parameter conversion unit 112 or the response / expression parameter conversion unit 132, or apply default values ​​to the expression parameter values.

[0105] The feature point / expression parameter conversion unit 122 may use a "feature point / expression parameter conversion dictionary" that stores rules for converting the positional information of facial feature points extracted from the character's face depicted in the illustration into expression parameter values ​​for the avatar, to convert the positional information of facial feature points of the character's face obtained by the facial feature point extraction unit 121 into expression parameter values. Table 2 below shows an example of a feature point / expression parameter conversion dictionary. This feature point / expression parameter conversion dictionary uses facial feature points extracted from the character's face image as headwords and stores the correspondence between the positional information of the character's facial feature points and the facial feature points of the avatar and their geometric information (position and displacement) as conversion rules.

[0106]

[0107] The conversion rules stored in the feature point / expression parameter conversion dictionary may be derived from empirical rules or generated using a pre-trained model learned from illustrations. Alternatively, the feature point / expression parameter conversion unit 122 may use a pre-trained model learned from illustrations, rather than the feature point / expression parameter conversion dictionary described above, to convert the feature points of a character's face into expression parameter values.

[0108] D-7. Answer / Expression Parameter Conversion The Answer / Expression Parameter Conversion Unit 132 converts the answers to the questions about the input image, output from the VLM Image Analysis Unit 131, into 3D character expression parameter values ​​that can be applied to the avatar.

[0109] The response / facial expression parameter conversion unit 132 pre-determines the processing for each answer to each question and then performs the processing. In addition, if an unexpected answer is input from the VLM image analysis unit 131, the response / facial expression parameter conversion unit 132 stops processing.

[0110] The response / expression parameter conversion unit 132 may convert the responses obtained by the VLM image analysis unit 131 into expression parameter values ​​using a "response / expression parameter conversion dictionary" which stores questions about the character's face depicted in the illustration and rules for converting each answer to those questions into expression parameter values ​​for the avatar. Table 3 below shows an example of a response / expression parameter conversion dictionary. This response / expression parameter conversion dictionary uses questions about the input image as headwords and stores the correspondence between the avatar's facial feature points and their geometric information (position and displacement) for each answer to the question as conversion rules.

[0111]

[0112] The conversion rules stored in the response / expression parameter conversion dictionary may be derived from empirical rules or generated using a pre-trained model learned from illustrations. Alternatively, the response / expression parameter conversion unit 132 may use a pre-trained model learned from illustrations, rather than the response / expression parameter conversion dictionary described above, to convert the responses to questions about the input image into expression parameter values.

[0113] D-8. Information Integration The information integration unit 140 integrates the outputs of the tag / expression parameter conversion unit 112, the feature point / expression parameter conversion unit 122, and the response / expression parameter conversion unit 132 to finally determine the expression parameter values ​​of the 3D character that can be applied to the avatar.

[0114] The correspondence status and accuracy of each module—the tag / expression parameter conversion unit 112, the feature point / expression parameter conversion unit 122, and the response / expression parameter conversion unit 132—differ for each body part and element. Therefore, the information integration unit 140 predetermines the priority of each module and determines the final value of any conflicting expression parameter values ​​between modules. Furthermore, if there are any expression parameter values ​​that cannot be predicted by any module, the information integration unit 140 applies a default value.

[0115] The information integration unit 140 can terminate processing if it determines that it cannot determine the facial expression parameter values ​​from the information extracted by the tag extraction unit 111, the facial feature point extraction unit 121, and the VLM image analysis unit 131. This prevents a decrease in the accuracy of the facial expression parameter values ​​output from the information processing system 100.

[0116] Furthermore, if the information integration unit 140 determines that it has been able to determine the facial expression parameter values ​​to some extent from the information extracted by some modules, such as the tag extraction unit 111 and the facial feature point extraction unit 121, it may skip some of the processing of the remaining modules, such as the VLM image analysis unit 131. Also, if the information integration unit 140 determines that the predictions of the tag-facial expression parameter conversion unit 112 and the feature point-facial expression parameter conversion unit 122 are insufficient, it may input additional question text into the VLM image analysis unit 131.

[0117] D-9. The facial expression parameter values ​​that the facial expression parameter value information processing system 100 generates from the input image (illustration) and which are applied to generating the avatar's facial expressions are stored in the facial expression parameter value database 160. Each facial expression parameter value stored in the facial expression parameter value database 160 is associated with and stored as metadata the text information entered along with the input image (see Figure 5). Specifically, the text information entered along with the input image is the dialogue contained in the illustration which is the input image. In addition, tags extracted from the input image by the tag extraction unit 111 and the response text generated from the input image by the VLM image analysis unit 131 (information such as facial expressions and emotions obtained through image analysis) may also be included and stored as metadata.

[0118] By applying the facial expression parameter values ​​finally determined by the information integration unit 140 to a 3D character such as an avatar, it is expected that the avatar will display the same facial expression as the character depicted in the illustration input to the information processing system 100.

[0119] D-10. Input Auxiliary Information: If the input image contains text information such as lines spoken by the character in the image, past lines, or the situation (speaker information, background information, character personality, etc.), the information processing system 100 may also input such text information as auxiliary information. In addition to the face image, images of the character's whole body or other images representing the context may also be input as auxiliary information along with the image.

[0120] Based on this auxiliary information, the VLM image analysis unit 131 can extract expressions and emotions from the faces of characters depicted in the input image, which is an illustration, with higher accuracy. In short, the auxiliary information is information that improves the accuracy of the expression parameter values. In the embodiments shown in Figures 3 and 4, only the VLM image analysis unit 131 is shown to use the input auxiliary information, but the tag extraction unit 111 and the face feature point extraction unit 121 may also use the input auxiliary information.

[0121] D-11. The black and white image coloring information processing system 100 may further include a coloring unit 150 that colors an input image when it is black and white and converts it into a color image. The coloring unit 150 may use tags extracted from the input image by the tag extraction unit 111 to color the black and white image with higher accuracy.

[0122] The facial feature point extraction unit 121 can extract facial feature points with higher accuracy by using a color image colored by the coloring unit 150 rather than a black and white input image. Of course, the tag extraction unit 111 and the VLM image analysis unit 131 also benefit from using a color image colored by the coloring unit 150 if the output accuracy is higher with a color image than with a black and white image. In the embodiments shown in Figures 3 and 4, only the facial feature point extraction unit 121 is shown to use a color image colored by the coloring unit 150, but the tag extraction unit 111 and the VLM image analysis unit 131 may also use a color image colored by the coloring unit 150.

[0123] The coloring unit 150 is intended to apply, for example, Image-to-Image technology using generation AI. Of course, if it is possible to output an image with natural coloring applied to the input image without changing anything other than color, such as facial expression, when a black and white image is input, other technologies may be applied to the coloring unit 150.

[0124] For example, Stable Diffusion, one of the image generation AIs, has ControlNet, an extension that allows for specifying detailed conditions during image generation. By applying this ControlNet to the coloring unit 150, a black and white image can be colored.

[0125] Furthermore, if text can be entered simultaneously during coloring to instruct the coloring content, the accuracy of coloring can be improved by using some of the tags extracted by the tag extraction unit 111. When the coloring unit 150 uses image generation AI, the tags extracted from the input image by the tag extraction unit 111 should be input to the image generation AI as prompts. For example, by inputting the "open_mouth" tag extracted by the tag extraction unit 111 to the coloring unit 150 at the same time as the image, it is possible to reduce the likelihood of incorrect coloring results where the mouth is closed.

[0126] E. Examples of using facial expression parameter values ​​The facial expression parameter values ​​generated by the information processing system 100 from various input images are stored in the facial expression parameter value database 160 and can be applied to display the facial expressions of the avatar. In addition, the text information entered along with the input image is linked to and stored as metadata for the facial expression parameter values ​​stored in the facial expression parameter value database 160 (see Figure 5). By applying the facial expression parameter values ​​stored in the facial expression parameter value database 160 to the avatar, it is expected that the avatar will display the same facial expression as the character depicted in the original illustration.

[0127] Figure 9 shows an example of a process that applies facial expression parameter values ​​to an avatar to express characteristic facial expressions and exaggerations.

[0128] Avatar 901 is accompanied by supplementary information 902, which mainly consists of text information such as Avatar 901's lines, situations, and past conversations in the virtual space in which Avatar 901 appears.

[0129] On the other hand, metadata is associated with each facial expression parameter value stored in the facial expression parameter value database 160. This metadata mainly consists of the character's dialogue included in the original illustration. In addition, the metadata may also include tags extracted from the illustration, and text information such as answers to questions about facial expressions and emotions generated by the VLM image analysis unit 131 from the input image. Figure 9 shows one facial expression parameter value (blend shape) 911 retrieved from the facial expression parameter value database 160, and the metadata 912 associated with it.

[0130] By retrieving the expression parameter values ​​most appropriate to the current state of avatar 901 from the expression parameter value database 160 and applying them to avatar 901, it is possible to display appropriate and characteristic expressions or exaggerated expressions on avatar 901.

[0131] The method for finding the most appropriate facial expression parameter values ​​for Avatar 901's current state will be explained with reference to Figure 9.

[0132] First, the accompanying information 902 of the avatar 901 is input to the visual language model 921 to generate a feature vector 922 that quantifies the characteristics of the avatar 901's current facial expressions and emotions. The visual language model (VLM) 921 is a large-scale language model, such as GPT-4o.

[0133] Meanwhile, metadata 912 linked to the facial expression parameter value (blend shape) 911 extracted from the facial expression parameter value database 160 is similarly input into the visual language model 931 to generate a feature vector 932 that quantifies the facial expression and emotional characteristics corresponding to this facial expression parameter value 911.

[0134] Next, the system searches the expression parameter value database 160 for the expression parameter value 911 that is closest to the expression and emotion of the avatar 901. Specifically, it finds the expression parameter value 911 that corresponds to the accompanying information 912, which is the feature vector 932 most similar to the feature vector 922 generated from the metadata 902 of the avatar 901. For example, the similarity between vectors can be expressed using cosine similarity.

[0135] Then, by applying the facial expression parameter value 911, which is closest to the facial expression and emotions of avatar 901, it is possible to generate an avatar face 941 that expresses appropriate and characteristic facial expressions and exaggerations.

[0136] F. Hardware Configuration Diagram 6 of the Information Processing Device shows an example of the hardware configuration of the information processing device 2000 applicable to this disclosure. The information processing device 2000 is composed of, for example, a personal computer (PC), but some functions may be composed of information terminals such as tablets and smartphones. An information processing system 100 according to this disclosure can be constructed using the information processing device 2000. The information processing system 200 according to this disclosure can be constructed using one information processing device 2000, or it can be constructed by linking multiple information processing devices 2000. Furthermore, an agent system that utilizes characters from anime and manga can also be constructed using the information processing device 2000.

[0137] This information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013.

[0138] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which consists of a CPU bus and other components. The CPU 2001 controls the overall operation of the information processing device 2000 according to various programs. The ROM 2002 non-volatilely stores programs (such as the basic input / output system) and arithmetic parameters used by the CPU 2001. The RAM 2003 is used to load programs to be executed by the CPU 2001 and to temporarily store parameters such as work data that change as needed during program execution. Through the collaborative operation of the ROM 2002 and RAM 2003, the CPU 2001 can execute various application programs in the execution environment provided by the operating system (OS) to realize various functions and services.

[0139] If the information processing device 2000 is a PC, the OS is, for example, Microsoft's Windows® or Unix® or its successor OS. The programs loaded into RAM 2003 and executed on CPU 2001 are the OS and various application programs. For example, in the information processing system 100 according to this disclosure, each functional module that performs processes such as tag extraction, facial feature point extraction, VLM image analysis, grayscale image coloring, tag-to-expression parameter conversion, feature point-to-expression parameter conversion, response-to-expression parameter conversion, and information integration are executed on the information processing device 2000.

[0140] Furthermore, when performing computationally intensive processes such as training AI (Artificial Intelligence) models on the information processing device 2000, it is desirable that the CPU 2001 be a multi-core CPU (e.g., Apple M1 Max), or that the information processing device 2000 be equipped with additional multi-core processors such as a GPU or GPGPU (General-purpose computing on graphics processing units) (e.g., NVIDIA's "RTX A6000") in addition to the CPU 2001. However, for convenience, these will be collectively referred to simply as the CPU 2001 below.

[0141] The host bus 2004 is connected to the expansion bus 2006 via the bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 is not limited to a configuration in which circuit components are hierarchically separated by the host bus 2004, the bridge 2005 and the expansion bus 2006, but may also be configured in which almost all circuit components are interconnected by a single bus (not shown).

[0142] The interface unit 2007 connects peripheral devices such as the input unit 2008, output unit 2009, storage unit 2010, drive 2011, and communication unit 2013 in accordance with the expansion bus 2006 standard. However, not all peripheral devices shown in Figure 6 are necessarily required, and the information processing device 2000 may include additional peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some peripheral devices may be externally connected to the main body of the information processing device 2000.

[0143] The input unit 2008 consists of an input control circuit that generates an input signal based on user input and outputs it to the CPU 2001. If the information processing device 2000 is a PC, the input unit 2008 may include a keyboard, mouse, touch panel, camera, and microphone. The output unit 2009 may include, for example, a liquid crystal display (LCD) device, an organic EL (Electro-Luminescence) display device, and an LED (Light Emitting Diode) display device, as well as an audio output device such as a speaker.

[0144] The storage unit 2010 stores files such as programs (applications, OS, etc.) and various data executed by the CPU 2001. The storage unit 2010 is composed of, for example, a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device.

[0145] The removable storage medium 2012 is a storage medium configured in a cartridge format, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 2012. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or storage unit 2010, and writes data on the RAM 2003 or storage unit 2010 to the removable storage medium 2012.

[0146] The communication unit 2013 is a device that performs wireless communication such as Wi-Fi®, Bluetooth®, and cellular communication networks such as 4G and 5G. Furthermore, the communication unit 2013 may be equipped with terminals such as USB (Universal Serial Bus) and HDMI® (High-Definition Multimedia Interface), and may also have the functionality to perform HDMI® communication with USB devices such as scanners and printers, and displays. Programs executed on the information processing device 2000 are installed externally, for example, through the communication unit 2013.

[0147] G. Effects The effects brought about by the information processing system 100 related to this disclosure are summarized below.

[0148] (1) The information processing system 100 relating to this disclosure is configured to obtain 3D character expression parameter values ​​that can be applied to generating avatar expressions by combining multiple visual models, inferring expression elements from illustrations according to the characteristics of each model, and integrating these results. Therefore, according to the information processing system 100 relating to this disclosure, a large number of expression parameter values ​​applicable to 3D characters can be automatically obtained from the faces of characters depicted in illustrations. Furthermore, by processing illustrations of various styles, a wide range of expression parameter values ​​can be obtained.

[0149] (2) When the input image is a grayscale image, the accuracy of the output may decrease in some visual language models. In response to this, the information processing system 100 according to this disclosure can improve accuracy by applying color to the input grayscale image using a coloring AI model before inputting it to the visual language model. For example, the accuracy of the facial expression parameter values ​​output by the facial feature point extraction unit can be improved. In addition, when applying color, the accuracy of coloring can be improved by providing some of the tags output from the tag extraction unit to the coloring AI model as auxiliary information. For example, by inputting the open_mouth tag output from the tag extraction unit to the coloring AI model, it is possible to reduce the occurrence of incorrect coloring results where the mouth is closed.

[0150] (3) If the input image contains text information such as lines spoken by the character in the image, past lines, or the situation (speaker information, background information, character personality, etc.), the information processing system may also input such text information as auxiliary information. In such cases, the VLM image analysis unit can use the context to improve the accuracy of the image analysis.

[0151] The present disclosure has been described in detail above with reference to specific embodiments. However, this disclosure should not be construed as being limited to the embodiments described above, and it will be obvious that those skilled in the art can modify or substitute these embodiments without departing from the gist of the disclosure. Furthermore, the effects described herein are merely illustrative, and the effects brought about by this disclosure are not limited, and there may be additional effects not described herein.

[0152] According to this disclosure, a large number of facial expression parameter values ​​applicable to avatars can be prepared at low cost. For example, in a system that interacts with anime-style characters such as avatars or agents, facial expression parameter values ​​generated using this disclosure can be used when the character expresses a variety of facial expressions that are appropriate to the context, such as dialogue or situation.

[0153] This disclosure allows for the creation of numerous facial expression parameter values, which include parameters related to the facial image and data that allows for the inference of emotions (such as the current line of dialogue and the preceding line of dialogue). Therefore, when an avatar speaks a line of dialogue, by applying the facial expression parameter value associated with the emotion that is closest to the spoken line (for example, the emotion with the highest cosine similarity), the avatar's facial expression can be made to match the spoken line.

[0154] In short, this disclosure has been explained in the form of examples, and the contents of this specification should not be interpreted restrictively. The claims should be considered in order to determine the gist of this disclosure.

[0155] The series of processes described herein can be executed by hardware, software, or a combination of hardware and software. When executing the processes by software, a program recording the processing sequence related to the implementation of this disclosure is installed and executed in the memory of a computer embedded in dedicated hardware. It is also possible to install the program on a general-purpose computer capable of executing various processes and execute the processes related to the implementation of this disclosure.

[0156] The program can be pre-stored on a storage medium installed in a computer, such as an HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable storage medium such as a flexible disk, CD-ROM (Compact Disc Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), BD (Blu-Ray Disc®), magnetic disk, or USB (Universal Serial Bus) memory. Using such a removable storage medium, the program related to the implementation of this disclosure can be provided as so-called packaged software.

[0157] Alternatively, the program may be transferred to a computer wirelessly or via a wired connection from a download site to a network such as a cellular network (WAN, Wide Area Network), a LAN (Local Area Network), or the Internet. The computer can then receive the program transferred in this way and install it into a large-capacity storage device such as an HDD or SSD.

[0158] Furthermore, this disclosure may also take the following form.

[0159] (1) An information processing system comprising: an acquisition unit that acquires an input image consisting of an illustration of a character's face; an extraction unit that extracts facial feature quantities from the input image; and a conversion unit that converts the extracted feature quantities into an expression parameter value that includes positional information of at least one facial feature point of the avatar.

[0160] (2) The information processing system according to (1) above, wherein the extraction unit includes a plurality of extraction units that extract different types of facial feature quantities based on different methods or algorithms, the conversion unit includes a plurality of conversion units that convert the feature quantities of the type corresponding to each of the plurality of extraction units into facial expression parameter values, and the system further comprises an integration unit that integrates the facial expression parameter values ​​output from the plurality of conversion units.

[0161] (3) The information processing system as described in (2) above, wherein the plurality of extraction units include a tag extraction unit that extracts tags related to facial expressions from the input image, a face feature point extraction unit that extracts positional information of face feature points from the input image, and an image analysis unit that generates answers to questions regarding the input image, and the plurality of conversion units include a tag-facial expression parameter conversion unit that converts the tags extracted by the tag extraction unit into facial expression parameter values, a feature point-facial expression parameter conversion unit that converts the positional information of face feature points extracted by the face feature point extraction unit into facial expression parameter values, and an answer-facial expression parameter conversion unit that converts the answers generated by the image analysis unit into facial expression parameter values.

[0162] (4) The information processing system as described in (3) above, wherein the tag extraction unit extracts tags related to facial expressions from the input image using a trained model trained on illustrations, and the tag-to-facial-parameter conversion unit converts the tags extracted by the tag extraction unit into facial-parameter values ​​based on rules for converting tags into facial-parameter values.

[0163] (5) The information processing system according to either (3) or (4) above, wherein the facial feature point extraction unit detects positional information of facial feature points from the input image using a face detection model learned from illustrations, and the feature point / expression parameter conversion unit predicts the position and shape of each part based on the positional information of facial feature points extracted by the facial feature point extraction unit, and obtains expression parameter values ​​based on the rules for converting the position and shape of each part into expression parameter values.

[0164] (6) The information processing system according to any one of (3) to (5) above, wherein the image analysis unit generates answers to questions about the input image using a visual language model, and the answer / expression parameter conversion unit converts the answers to questions about the input image generated by the image analysis unit into expression parameter values ​​based on rules for converting answers into expression parameters.

[0165] (7) The information processing system according to any one of (3) to (6) above, wherein the integration unit integrates the expression parameter values ​​output from each of the tag / expression parameter conversion unit, the feature point / expression parameter conversion unit, and the response / expression parameter conversion unit to determine all expression parameter values.

[0166] (8) The information processing system as described in (7) above, wherein the integration unit determines priority and adopts one of the outputs if the same facial expression parameter value is output from two or more of the tag / facial expression parameter conversion unit, the feature point / facial expression parameter conversion unit, and the response / facial expression parameter conversion unit.

[0167] (9) The information processing system according to either (7) or (8) above, wherein if the output of an expression parameter value from any of the expression parameter conversion units among the tag / expression parameter conversion unit, the feature point / expression parameter conversion unit, and the response / expression parameter conversion unit is insufficient, the integration unit prompts the corresponding expression parameter conversion unit to re-output an expression parameter value, and applies an expression parameter value output from another expression parameter conversion unit or a default expression parameter value.

[0168] (10) The information processing system according to any one of (7) to (9) above, wherein the integration unit determines that the output of facial expression parameter values ​​from the tag / facial expression parameter conversion unit and the feature point / facial expression parameter conversion unit is insufficient, and additionally inputs a question to the image analysis unit.

[0169] (11) The information processing system according to any one of (3) to (10) above, wherein the image analysis unit further considers text information related to the input image to generate an answer to a question about the input image.

[0170] (12) The information processing system according to any one of (3) to (11) above, further comprising a coloring unit that colors the black and white input image to generate a color image, wherein the face feature point extraction unit extracts positional information of face feature points from the color image generated by the coloring unit.

[0171] (13) The information processing system according to (12) above, wherein the coloring unit uses the tags extracted from the input image by the tag extraction unit to color the black and white input image.

[0172] (13-1) The coloring unit uses ControlNet, an extended function of Stable Diffusion, or other image generation AI to color the input image, as described in either (12) or (13) above.

[0173] (14) The information processing system according to any one of (3) to (13) above, wherein the tag extraction unit extracts known tags from the input image using a visual model trained on Stable Diffusion Tagger or other illustrations.

[0174] (15) The information processing system according to any one of (3) to (14) above, wherein the facial feature point extraction unit extracts positional information of facial feature points from the input image using a visual model trained on facial images of AnimeFaceDetector or other illustrations.

[0175] (16) The information processing system according to any one of (3) to (15) above, wherein the image analysis unit generates answers to questions regarding the input image using a large-scale language model or other visual model capable of image input.

[0176] (17) The information processing system according to any one of (3) to (16) above, wherein the tag-expression parameter conversion unit converts the tags extracted by the tag extraction unit into expression parameter values ​​based on a predetermined process for each tag.

[0177] (17) The information processing system described in (17) above, wherein the tag-expression parameter conversion unit converts the tags into expression parameter values ​​by utilizing the probability or other numerical values ​​of the tags extracted by the tag extraction unit.

[0178] (17-2) The information processing system according to (17) above, wherein the tag / expression parameter conversion unit processes the competing tags extracted by the tag extraction unit based on the predetermined priority of each tag.

[0179] (17-3) The information processing system described in (17) above, wherein the tag / expression parameter conversion unit does nothing if the tag input from the tag extraction unit is 0.

[0180] (17-4) The information processing system described in (17) above, wherein the tag / expression parameter conversion unit does not perform any processing when an unknown tag is input from the tag extraction unit.

[0181] (18) The information processing system according to any one of (3) to (17) above, wherein the feature point / expression parameter conversion unit converts the position information of the face feature points extracted by the face feature point extraction unit into expression parameter values ​​based on processing predetermined for each face feature point.

[0182] (18-1) The information processing system described in (18) above, wherein the feature point / expression parameter conversion unit converts the probability or other numerical value of the facial feature points extracted by the facial feature point extraction unit into an expression parameter value.

[0183] (19) An information processing method comprising: an acquisition step of acquiring an input image consisting of an illustration of a character's face; an extraction step of extracting facial feature quantities from the input image; and a conversion step of converting the extracted feature quantities into an expression parameter value that includes position information of at least one facial feature point of the avatar.

[0184] (20) A computer program written in a computer-readable format to cause a computer to function as: an acquisition unit that acquires an input image consisting of an illustration of a character's face; an extraction unit that extracts facial feature quantities from the input image; and a conversion unit that converts the extracted feature quantities into facial parameter values ​​that include positional information of at least one facial feature point of an avatar.

[0185] 10... Information processing system, 11... Feature extraction unit, 12... Feature / expression parameter conversion unit, 20... Information processing system, 21... Feature extraction unit, 22... Feature / parameter conversion unit, 23... Integration unit, 100... Information processing system, 111... Tag extraction unit, 112... Tag / expression parameter conversion unit, 121... Face feature point extraction unit, 122... Feature point / expression parameter conversion unit, 131... VLM image analysis unit, 132... Response / expression parameter conversion unit, 140... Information integration unit, 150... Coloring unit, 160... Expression parameter value database, 2000... Information processing device, 2001... CPU, 2002... ROM, 2003... RAM, 2004... Host bus, 2005... Bridge, 2006... Expansion bus, 2007... Interface unit, 2008... Input unit, 2009... Output unit, 2010... Storage unit 2011…Drive, 2012…Removable storage medium, 2013…Communications department

Claims

1. An information processing system comprising: an acquisition unit that acquires an input image consisting of an illustration depicting a character's face; an extraction unit that extracts facial feature quantities from the input image; and a conversion unit that converts the extracted feature quantities into facial parameter values ​​that include positional information of at least one facial feature point of an avatar.

2. The information processing system according to claim 1, wherein the extraction unit includes a plurality of extraction units that extract different types of facial feature quantities based on different methods or algorithms, the conversion unit includes a plurality of conversion units that convert the feature quantities of the type corresponding to each of the plurality of extraction units into facial expression parameter values, and the system further comprises an integration unit that integrates the facial expression parameter values ​​output from the plurality of conversion units.

3. The information processing system according to claim 2, wherein the plurality of extraction units include a tag extraction unit that extracts tags related to facial expressions from the input image, a face feature point extraction unit that extracts positional information of face feature points from the input image, and an image analysis unit that generates answers to questions regarding the input image, and the plurality of conversion units include a tag-facial expression parameter conversion unit that converts the tags extracted by the tag extraction unit into facial expression parameter values, a feature point-facial expression parameter conversion unit that converts the positional information of face feature points extracted by the face feature point extraction unit into facial expression parameter values, and an answer-facial expression parameter conversion unit that converts the answers generated by the image analysis unit into facial expression parameter values.

4. The information processing system according to claim 3, wherein the tag extraction unit extracts tags related to facial expressions from the input image using a trained model learned from illustrations, and the tag-to-facial expression parameter conversion unit converts the tags extracted by the tag extraction unit into facial expression parameter values ​​based on rules for converting tags into facial expression parameter values.

5. The information processing system according to claim 3, wherein the facial feature point extraction unit detects positional information of facial feature points from the input image using a face detection model learned from illustrations, and the feature point / expression parameter conversion unit predicts the position and shape of each part based on the positional information of facial feature points extracted by the facial feature point extraction unit and obtains expression parameter values ​​based on rules for converting the position and shape of each part into expression parameter values.

6. The information processing system according to claim 3, wherein the image analysis unit generates answers to questions about the input image using a visual language model, and the answer / expression parameter conversion unit converts the answers to questions about the input image generated by the image analysis unit into expression parameter values ​​based on rules for converting answers into expression parameters.

7. The information processing system according to claim 3, wherein the integration unit integrates the expression parameter values ​​output from each of the tag / expression parameter conversion unit, the feature point / expression parameter conversion unit, and the response / expression parameter conversion unit to determine all expression parameter values.

8. The information processing system according to claim 7, wherein if the same facial expression parameter value is output from two or more of the tag / facial expression parameter conversion unit, the feature point / facial expression parameter conversion unit, and the response / facial expression parameter conversion unit, the integration unit determines the priority and adopts one of the outputs.

9. The information processing system according to claim 7, wherein if the output of an expression parameter value from any of the expression parameter conversion units among the tag / expression parameter conversion unit, the feature point / expression parameter conversion unit, and the response / expression parameter conversion unit is insufficient, the integration unit prompts the corresponding expression parameter conversion unit to re-output an expression parameter value, and applies an expression parameter value output from another expression parameter conversion unit or a default expression parameter value.

10. The information processing system according to claim 7, wherein if the integration unit determines that the output of facial expression parameter values ​​from the tag / facial expression parameter conversion unit and the feature point / facial expression parameter conversion unit is insufficient, it inputs an additional question to the image analysis unit.

11. The information processing system according to claim 3, wherein the image analysis unit further considers text information related to the input image to generate an answer to a question about the input image.

12. The information processing system according to claim 3, further comprising a coloring unit that colors the aforementioned black and white input image to generate a color image, wherein the facial feature point extraction unit extracts positional information of facial feature points from the color image generated by the coloring unit.

13. The information processing system according to claim 12, wherein the coloring unit uses the tags extracted from the input image by the tag extraction unit to color the black and white input image.

14. The information processing system according to claim 3, wherein the tag extraction unit extracts known tags from the input image using a visual model trained on Stable Diffusion Tagger or other illustrations.

15. The information processing system according to claim 3, wherein the facial feature point extraction unit extracts positional information of facial feature points from the input image using a visual model trained on facial images of AnimeFaceDetector or other illustrations.

16. The information processing system according to claim 3, wherein the image analysis unit generates answers to questions regarding the input image using a large-scale language model or other visual model capable of image input.

17. The information processing system according to claim 3, wherein the tag-expression parameter conversion unit converts the tags extracted by the tag extraction unit into expression parameter values ​​based on a predetermined process for each tag.

18. The information processing system according to claim 3, wherein the feature point / expression parameter conversion unit converts the position information of the facial feature points extracted by the facial feature point extraction unit into an expression parameter value based on a predetermined process for each facial feature point.

19. An information processing method comprising: an acquisition step of acquiring an input image consisting of an illustration depicting a character's face; an extraction step of extracting facial feature quantities from the input image; and a conversion step of converting the extracted feature quantities into an expression parameter value that includes positional information of at least one facial feature point of the avatar.

20. A computer program written in a computer-readable format to cause a computer to function as: an acquisition unit that acquires an input image consisting of an illustration depicting a character's face; an extraction unit that extracts facial feature quantities from the input image; and a conversion unit that converts the extracted feature quantities into facial parameter values ​​that include positional information of at least one facial feature point of an avatar.