Information processing device, information processing method, program, data generation device, data generation method, and learning system
By analyzing attribute distributions and adjusting training data based on user input and CG image generation, the system addresses inaccuracies in existing AI model training, enhancing learning accuracy and efficiency.
Patent Information
- Application Number
- PCT/JP2025/009570
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2025-03-13
- Publication Date
- 2025-10-02
AI Technical Summary
Existing AI model training methods, such as those described in Patent Document 1, may not accurately estimate data biases based on inference results, leading to suboptimal training accuracy, and do not account for user-specific preferences in data adjustments.
An information processing device and system that analyzes attribute distributions using a turbulence evaluation value for convolution filter weights, presents results to users, and adjusts training data based on user input, while generating training data using CG images to ensure a desired attribute distribution, thereby improving learning accuracy and efficiency.
This approach allows for precise evaluation of AI model correctness and tailored data adjustments, enhancing learning accuracy while reducing the need for capturing real-life images, thus improving training efficiency and privacy compliance.
Smart Images

Figure JP2025009570_02102025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, program, data generation device, data generation method, and learning system
[0001] The present technology relates to an information processing device, an information processing method, a program, a data generation device, a data generation method, and a learning system, and in particular to a technology related to learning of an AI (Artificial Intelligence) model.
[0002] There are various techniques for performing inference processing using AI (artificial intelligence) models. In training an AI model, using appropriate training data is considered important for improving performance. For example, Patent Document 1 below discloses a technique for estimating bias in attributes of training data based on annotation information (attribute information) of the training data and the inference results of the AI model, and adjusting the training data based on the estimated bias information. Specifically, Patent Document 1 below discloses a technique for an AI model that performs object detection processing targeting people, in which if there are many false detections of people in their 30s, it is estimated that there is insufficient training data with people in their 30s as subjects, and the missing data is added to the training data.
[0003] Japanese Patent Application Laid-Open No. 2021-111101
[0004] However, the accuracy of an inference result does not necessarily indicate the accuracy of the operation of an AI model. For example, in a well-known example, an AI model that performs object detection processing targeting a specific subject may obtain a correct detection result even though it references an area in an input image that is different from the image area of the subject. For this reason, estimating the bias in attributes of training data based on the inference result as in Patent Document 1, specifically estimating attributes that are lacking (or considered excessive) in training, may not accurately estimate the bias in attributes of training data. As a result, the technology in Patent Document 1 may not improve training accuracy.
[0005] Furthermore, in Patent Document 1, adjustment of learning data is performed automatically based on annotation information and inference results, but it is conceivable that different users may have different preferences for which attributes of insufficient or excess data they would like to prioritize.
[0006] This technology was developed in consideration of the above circumstances, and aims to realize a method for adjusting the learning data used in AI model training that improves learning accuracy while also addressing user requests.
[0007] The information processing device according to the present technology includes an analysis unit that performs an analysis of an attribute distribution, which is a statistical data distribution of attributes of training data used for training an AI model, based on a turbulence evaluation value indicating the degree of turbulence of weight coefficients of a convolution filter when the training data is input to the AI model, a presentation processing unit that processes to present the analysis results by the analysis unit to a user, and an adjustment control unit that controls adjustment of the training data based on a target distribution of the attribute distribution set based on user operation related to the analysis results presented by the presentation processing unit. By using the turbulence evaluation value indicating the degree of turbulence of the weight coefficients of the convolution filter, in other words, the responsiveness of the convolution filter, it is possible to appropriately evaluate the correctness of the operation of the AI model. The presentation processing unit and the adjustment control unit then present the analysis results of the attribute distribution based on the turbulence evaluation value that can appropriately evaluate the correctness of the operation of the AI model, and adjust the training data based on user operation related to the presented analysis results.
[0008] In addition, a data generation device according to the present technology includes a head model generation unit that generates a three-dimensional head model by three-dimensionalizing a generated facial image, a human body model generation unit that generates a three-dimensional human body model by integrating the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body, and an imaging processing unit that two-dimensionalizes the three-dimensional human body model to generate training data using CG images. Generating training data using CG images eliminates the need to capture images of each attribute to obtain training data with a desired attribute distribution, thereby improving the efficiency of data adjustment work for the training data. In this case, by using the generated facial image as the original data for the head model, the subject of the training data can be a person who is highly realistic but does not actually exist.
[0009] Furthermore, the learning system according to the present technology includes a data generation unit having a head model generation unit that three-dimensionalizes a generated facial image to generate a three-dimensional head model, a human body model generation unit that integrates the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body to generate a three-dimensional human body model, and an imaging processing unit that two-dimensionalizes the three-dimensional human body model to generate learning data in the form of CG images; an analysis unit that performs analysis of an attribute distribution, which is a statistical data distribution for attributes of learning data used for learning an AI model, based on an intensity evaluation value that indicates the degree of intensity fluctuation of a weight coefficient of a convolution filter when the learning data is input to the AI model; a presentation processing unit that performs processing to present the analysis results by the analysis unit to a user; an adjustment control unit that controls adjustment of the learning data using the data generation unit based on a target distribution of the attribute distribution set based on user operation related to the analysis results presented by the presentation processing unit; and a learning processing unit that performs learning processing of the AI model using the adjusted learning data. The above-described learning system makes it possible to improve the learning accuracy of the training data used in AI model training while also meeting user needs. Furthermore, by generating training data using CG images, it is no longer necessary to capture images of each attribute in order to obtain training data with a desired attribute distribution, which makes it possible to improve the efficiency of the data adjustment work for the training data.
[0010] 1 is a block diagram showing an example of the configuration of a learning system as an embodiment. FIG. 1 is a diagram showing an example of the hardware configuration of an information processing device as an embodiment. FIG. 2 is a functional block diagram for explaining functions possessed by an information processing device and a data generation device in an embodiment. FIG. 2 is an explanatory diagram of a functional configuration for providing variations in a three-dimensional human body model. FIG. 3 is a diagram for explaining a flow of three-dimensional human body model generation. FIG. 4 is an explanatory diagram of an image correction related to face direction. FIG. 5 is an explanatory diagram of image correction related to a front hair portion. FIG. 6 is an explanatory diagram of image correction related to a closed mouth. FIG. 7 is an explanatory diagram of a problem with conventional eyeball fitting. FIG. 8 is an explanatory diagram of optimization of eyeball fitting. FIG. 9 is an explanatory diagram of optimization of integration of a head model and a body model. FIG. 10 is an explanatory diagram of an adjustment process related to a skin mesh. FIG. 11 is a diagram showing an example of an initial distribution setting screen. FIG. 12 is an example of a sample presentation screen. FIG. 13 is an example of an attribute distribution presentation screen. FIG. 14 is an example of a generated data list display screen. FIG. 15 is an explanatory diagram of a method for estimating attributes of insufficient data and excessive data. FIG. 16 is an example of a surplus / deficient data presentation screen. FIG. 17 is an example of a target distribution setting screen. FIG. 18 is an example of an evaluation result presentation screen. FIG. 19 is a flowchart showing an example of a specific processing procedure for realizing a learning method as an embodiment. 10 is an explanatory diagram of a configuration example in which an information processing device has a data generating unit, and FIG. 11 is an explanatory diagram of a configuration example in which a device other than the information processing device has a learning processing unit.
[0011] Hereinafter, with reference to the accompanying drawings, embodiments of the present technology will be described in the following order: <1. Example of configuration of learning system> <2. Example of hardware configuration of information processing device> <3. Learning method as an embodiment> (3-1. Overview of the method) (3-2. Data generation method) (3-3. Specific learning method including analysis related to attribute distribution) (3-4. Processing procedure) <4. Modification example> <5. Summary of embodiment> <6. The present technology>
[0012] 1. Configuration Example of a Learning System FIG. 1 is a block diagram showing a configuration example of a learning system according to an embodiment, which includes an information processing device 1 according to an embodiment of the present technology. The learning system according to the embodiment is a system for training an AI (Artificial Intelligence) model that performs inference processing for a predetermined task using a predetermined type of data, such as image data, sound data, or text data, as input data. As an example, the learning system according to the embodiment trains an AI model that uses image data as input data to achieve a predetermined image analysis task, such as object detection processing or object recognition processing. Note that the object detection processing here refers to a task of obtaining the position of an object in an image, and the inference result is output as region information, such as a bounding box indicating the area where the object exists. Furthermore, the object recognition processing refers to a task of recognizing the object in an input image, and the inference result is output as a likelihood for each object class, such as a person, dog, or cat. Here, the image analysis task may be a task that combines object detection processing and object recognition processing, specifically, a task of recognizing (classifying) an object while also detecting its area. Furthermore, a task such as semantic segmentation, which classifies objects in predetermined block units such as pixel units, may also be considered. In the following, as an example for the purpose of explanation, the inference processing by the AI model is assumed to be human recognition processing based on facial features.
[0013] In this embodiment, it is assumed that inference processing using a trained AI model is executed within a camera. Specifically, it is assumed that an AI processing unit that performs inference processing using an AI model is provided within an image sensor in the camera. Examples of uses of such cameras include use as surveillance cameras in stores, on roads, parking lots, etc. The AI processing unit performs inference processing targeting, for example, people or automobiles, and the inference results can be used to analyze customer trends in a store, the amount of automobile traffic, the number of parked cars, etc.
[0014] As shown in the figure, the learning system of the embodiment includes at least an information processing device 1, a data generating device 2, and a user terminal 3. In the learning system, the information processing device 1, the data generating device 2, and the user terminal 3 are each configured as a computer device, and are capable of mutual data communication via a network NT, which is a communication network such as the Internet.
[0015] The information processing device 1 performs a learning process for an AI model and creates a trained AI model. The data generation device 2 generates learning data to be used in the learning process for the AI model in the information processing device 1.
[0016] The user terminal 3 is a computer device used by a user who receives at least one of the inference results of the trained AI model and the analysis results based on the inference results. For example, in the case of the above-mentioned surveillance camera application in a store, the user would be an employee of the company that operates the store.
[0017] Here, in the embodiment, the AI model is trained as follows: first, initial training is performed using real-life images associated with annotation information such as those provided by ImageNet, and then CG (Computer Graphics) images are used to train the AI model so that it can achieve performance appropriate for the inference environment.
[0018] In training an AI model, it would be ideal to create a highly versatile AI model that can handle a variety of situations, but realizing such a versatile AI model requires training using a huge amount of training data, which increases the training cost.Furthermore, realizing a highly versatile AI model requires the use of a relatively large-scale neural network, which also leads to an increase in hardware resources.
[0019] In the embodiments, it is assumed that inference processing using a trained AI model is performed on an edge device with relatively limited hardware resources, such as a camera. For this reason, a method is adopted in which, after initial training as described above, training data is generated using CG and then retrained. To improve adaptability to the actual inference environment, it would be ideal to use real-life images captured in the actual camera installation environment, but this would require a lot of effort from the user. Training using CG images eliminates the need to capture real-life images, while enabling the generation of appropriate training data tailored to the inference environment, allowing for efficient retraining of the AI model after initial training. Furthermore, using CG images prevents images of actual people from being included in the training data, thereby avoiding the risk of privacy violations.
[0020] 1, a data generating device 2 is a device that generates training data using such CG images. A method for generating training data according to an embodiment will be described later.
[0021] 2. Example of hardware configuration of information processing device> Fig. 2 is a block diagram showing an example of the hardware configuration of the information processing device 1. Note that the data generation device 2 and the computer device serving as the user terminal 3 shown in Fig. 1 can also have the same hardware configuration as that shown in Fig. 2, and therefore a graphical description of the hardware configuration of the data generation device 2 and the user terminal 3 will be omitted.
[0022] As shown in the figure, the information processing device 1 includes a CPU 11. The CPU 11 executes various processes in accordance with programs stored in a ROM 12 or programs loaded from a storage unit 19 into a RAM 13. The RAM 13 also stores data necessary for the CPU 11 to execute various processes as appropriate.
[0023] The CPU 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output interface (I / F) 15 is also connected to this bus 14.
[0024] An input unit 16 consisting of operators and operation devices is connected to the input / output interface 15. For example, the input unit 16 may be various operators and operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, a remote controller, etc. An input operation is detected by the input unit 16, and a signal corresponding to the input operation is interpreted by the CPU 11.
[0025] A display unit 17, such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) panel, and an audio output unit 18, such as a speaker, are connected integrally or separately to the input / output interface 15. The display unit 17 is used to display various types of information, and may be, for example, a display device provided in the housing of the computer device, or a separate display device connected to the computer device.
[0026] The display unit 17 displays images for various image processing, moving images to be processed, etc. on the display screen based on instructions from the CPU 11. The display unit 17 also displays various operation menus, icons, messages, etc., i.e., a GUI (Graphical User Interface), based on instructions from the CPU 11.
[0027] The input / output interface 15 may be connected to a storage unit 19 configured with a hard disk drive (HDD) or solid-state memory, or a communication unit 20 configured with a modem or the like.
[0028] The communication unit 20 performs communication processing via a transmission path such as the Internet, and communication with various devices via wired / wireless communication, bus communication, and the like.
[0029] A drive 21 is also connected to the input / output interface 15 as required, and a removable recording medium 22 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory is appropriately loaded therein.
[0030] The drive 21 can read data files such as programs used for various processes from the removable recording medium 22. The read data files are stored in the storage unit 19, and images and sounds contained in the data files are output on the display unit 17 and the audio output unit 18. Furthermore, the computer programs and the like read from the removable recording medium 22 are installed in the storage unit 19 as needed.
[0031] In a computer device having the above-described hardware configuration, for example, software for the processing of this embodiment can be installed via network communication by the communication unit 20 or via the removable recording medium 22. Alternatively, the software may be stored in advance in the ROM 12, the storage unit 19, etc. The CPU 11 performs processing operations based on various programs, thereby executing the information processing and communication processing required by the information processing device 1.
[0032] The information processing device 1 (and the data generating device 2) is not limited to being configured as a single computer device as shown in Fig. 3, but may be configured as a system of multiple computer devices. The multiple computer devices may be systemized using a LAN (Local Area Network) or the like, or may be located in a remote location using a VPN (Virtual Private Network) or the like using the Internet or the like. The multiple computer devices may include computer devices as a server group (cloud) available through a cloud computing service.
[0033] 3. Learning Method as an Embodiment> (3-1. Overview of Method) FIG. 3 is a functional block diagram for explaining the functions of the information processing device 1 and the data generating device 2 in the embodiment.
[0034] As shown in the figure, the data generating device 2 has a function as a data generating unit F20. Here, the function of the data generating device 2 is realized by the processing of a CPU 11 included in the data generating device 2.
[0035] The data generation unit F20 performs processing to generate learning data using the above-mentioned CG images, and as shown in the figure, includes a head model generation unit F21, a human body model generation unit F22, and an imaging processing unit F23. The processing of these head model generation unit F21, human body model generation unit F22, and imaging processing unit F23 will be described later in detail as a data generation method.
[0036] The information processing device 1 has functions as a learning processing unit F1, an analysis unit F2, a presentation processing unit F3, an adjustment control unit F4, and a learning control unit F5. Here, in this example, the functions of the learning processing unit F1, the analysis unit F2, the presentation processing unit F3, the adjustment control unit F4, and the learning control unit F5 are realized by processing of a CPU 11 provided in the information processing device 1, but it is not essential to adopt a configuration in which all of these functions are realized by processing of the CPU 11, and it is also possible to adopt a configuration in which at least some of the functions are realized by processing other than that of the CPU 11.
[0037] The learning processing unit F1 has a learning device F11 of the AI model, and inputs learning data to the learning device F11 to perform learning processing of the AI model. The learning processing unit F1 in this embodiment performs learning processing as re-learning for the AI model that has undergone the above-mentioned initial learning. In the learning processing as re-learning, the learning processing unit F1 performs learning processing of the AI model using learning data that has been adjusted under the control of the adjustment control unit F4. It is also possible to perform initial learning using the learning processing unit F1.
[0038] The analysis unit F2 analyzes the attribute distribution, which is a statistical data distribution of the attributes of the training data used to train the AI model, based on a turbulence evaluation value that indicates the degree of turbulence of the weight coefficients of the convolution filter when the training data is input to the AI model. Here, the turbulence evaluation value can be said to be an evaluation value that indicates the reactivity of the convolution filter, and the greater the "degree of turbulence" indicated by the turbulence evaluation value, the more learning remains. By using such a turbulence evaluation value, it is possible to appropriately evaluate the correctness of the operation of the AI model.
[0039] The presentation processing unit F3 performs processing to present various pieces of information related to learning to the user. Specifically, the presentation processing unit F3 performs processing to present the analysis results by the analysis unit F2 to the user.
[0040] The adjustment control unit F4 controls the adjustment of the learning data based on the target distribution of the attribute distribution set based on the user operation related to the analysis results presented by the presentation processing unit F3. Specifically, the adjustment control unit F4 controls the adjustment of the learning data using the data generation unit F20.
[0041] The learning control unit F5 controls the learning process of the AI model performed by the learning processing unit F1. Specifically, it controls the start and end of the learning process, and specifies the learning data to be used in the learning process. In particular, in this embodiment, it controls the learning process so that the learning data adjusted by the adjustment control unit F4 is used.
[0042] In this embodiment, the addition of training data to be used for re-learning is performed in a manner similar to so-called DA (Data Augmentation), in which new data is generated by, for example, transforming or combining the features of the training data (also called "clean data") used in the initial training, thereby expanding the amount of training data and the distribution of its attributes.
[0043] In this embodiment, the learning data is adjusted by using the above-described intensity evaluation value to estimate data attributes that are insufficient for learning and data attributes that are considered to be excessive, and then the insufficient attributes are supplemented and the excessive attributes are eliminated. In this embodiment, such adjustment of the learning data is performed using a data generation unit F20 described below.
[0044] (3-2. Data Generation Method) In FIG. 3 , the head model generation unit F21 in the data generation unit F20 generates a three-dimensional head model by three-dimensionalizing the generated facial image. Here, the term "generated facial image" refers to a facial image based on an image other than a captured image. Examples of "generated facial images" include those generated using a generative AI, such as a generator in GAN (Generative Adversarial Networks), and those generated by processing a captured facial image. Data with various attributes should be prepared as training data for the AI model. In other words, data with a fairly wide attribute distribution should be prepared. Using such generated facial images is advantageous because it allows training data with a wide range of attributes to be obtained without capturing images. For example, if a generative AI is used, facial images with various attributes can be generated by inputting prompts to the AI. Furthermore, if a captured image is processed, facial images with various attributes can be generated by processing a small number of seed facial images.
[0045] The human body model generation unit F22 generates a three-dimensional human body model by integrating the three-dimensional head model generated by the head model generation unit F21 with a three-dimensional body model, which is a three-dimensional model of the body. Furthermore, the imaging processing unit F23 converts the three-dimensional human body model generated by the human body model generation unit F22 into two-dimensional data to generate learning data using CG images.
[0046] In this embodiment, not only are variations in the facial images provided to the 3D human body models varied, but variations are also provided by the types of eyeballs and hair that are combined with the 3D human body models.Furthermore, in this embodiment, variations are provided to the 3D human body models by the types of body models that are integrated with the 3D head models.
[0047] 4 is an explanatory diagram of the functional configuration for providing variations in the three-dimensional human body model. In the diagram, the face image generation unit F30, three-dimensionalization unit F31, eyeball addition unit F32, hair addition unit F33, selection unit F35, and selection / adjustment unit F36 are considered to be part of the functions of the head model generation unit F21. Furthermore, the body integration unit F34 and selection / adjustment unit F37 are considered to be part of the functions of the human body model generation unit F22.
[0048] In this case, the flow of generating a three-dimensional human body model is as follows, as shown by the transition from Figure 5A to Figure 5E, first, the face image is three-dimensionalized to generate a three-dimensional model of the head, then an eyeball model and a hair model are fitted to the three-dimensional model, and finally a body model is integrated into the three-dimensional head model to which the eyeball model and hair model have been fitted.
[0049] In FIG. 4 , the facial image generation unit F30 performs processing related to the generation of facial images. That is, the processing is performed to obtain the above-mentioned "generated facial image." Specifically, the processing involves causing the generation AI to generate a facial image, or generating a facial image by processing the above-mentioned seed image. Here, the facial image is generated using the generation AI. In this case, the facial image generation unit F30 generates a plurality of facial images so that the specified attribute distribution conditions are satisfied by inputting prompts to the generation AI according to specified information on attribute distribution, such as age and gender, to generate facial images. For example, specified information on age distribution and gender ratio is provided as the specified information on attribute distribution (hereinafter referred to as "attribute distribution specified information"), and the facial image generation unit F30 causes the generation AI to generate a plurality of facial images so that the specified conditions on age distribution and gender ratio are satisfied.
[0050] The generation AI may be provided in a device external to the data generation device 2, but a configuration in which the data generation device 2 is provided with the generation AI is also conceivable. When a seed image is used in facial image generation, the seed image may be irreversibly processed. Processing may also change facial features such as age group and facial expression.
[0051] The three-dimensional rendering unit F31 renders the face image obtained by the face image generation unit F30 into a three-dimensional image to obtain a three-dimensional model of the head.
[0052] The eyeball addition unit F32 and the hair addition unit F33 fit an eyeball model selected from the eyeball model group MD1 and a hair model selected from the hair model group MD2, respectively, to the three-dimensional head model obtained by the three-dimensionalization unit F31, and the body integration unit F34 integrates a three-dimensional body model selected from the body model group MD3 with the three-dimensional head model to which the eyeball model and hair model have been fitted to generate a three-dimensional human body model.
[0053] The multiple types of eyeball models as the eyeball model group MD1, the multiple types of hair models as the hair model group MD2, and the multiple types of three-dimensional body models as the body model group MD3 are stored in a predetermined storage device such as the storage unit 19 included in the data generating device 2. In the eyeball model group MD1, the multiple types of eyeball models are eyeball models with at least different pupil colors, and in the hair model group MD2, the multiple types of hair models are hair models with at least different hairstyles. Furthermore, in the body model group MD3, the multiple types of three-dimensional body models are three-dimensional body models with at least different physiques and clothing shapes.
[0054] For the eyeball addition unit F32, the eyeball model to be used for fitting is selected by the selection unit F35 based on attribute distribution designation information. The attribute distribution designation information for the eyeball model is, for example, information designating the ratio of each pupil color, and the selection unit F35 randomly selects an eyeball model from the eyeball model group MD1 so that the designation condition is satisfied.
[0055] For the hair model, the selection and adjustment unit F36 not only selects a hair model from the hair model group MD2 but also adjusts the hair color. The attribute distribution designation information for the hair model is information that designates the ratio of hairstyle (e.g., long hair, short hair, etc.) and hair color, and the selection and adjustment unit F36 randomly selects a hair model from the hair model group MD2 and randomly adjusts the hair color of the selected hair model so that the designated conditions are met.
[0056] Similarly, for three-dimensional body models, the selection and adjustment unit F37 not only selects a body model from the body model group MD3, but also adjusts the skin color and clothing color. The attribute distribution designation information for the three-dimensional body model is, for example, information designating the ratio of physique, skin color, and clothing color, and the selection and adjustment unit F37 randomly selects a three-dimensional human body model from the body model group MD3 and randomly adjusts the skin and clothing color of the selected three-dimensional human body model so that the designated conditions are met.
[0057] The functional configuration shown in FIG. 4 makes it possible to generate a training data group having a specified attribute distribution as a training data group based on CG images.
[0058] In the above example, the head model generation unit F21 generates a plurality of types of three-dimensional head models that have different types of eyeball models and hair models to be combined, but it is also possible to fix either the eyeball model or the hair model. That is, in this embodiment, the head model generation unit F21 only needs to generate a plurality of types of three-dimensional head models that have different types of eyeball models or hair models to be combined.
[0059] Furthermore, in the above, an example was given in which the human body model generation unit F22 generates multiple types of three-dimensional human body models with different types of three-dimensional body models to be integrated into the three-dimensional head model, where the body shape, skin color, and clothing color are different. However, the change elements of the three-dimensional body model to be integrated into the three-dimensional head model are not limited to these body shape, skin color, and clothing color, and can also be other elements, such as body posture or clothing pattern.
[0060] In this embodiment, if a face in a face image to be three-dimensionalized faces in a direction other than the forward direction, the head model generation unit F21 performs image correction so that the face faces forward. Fig. 6 shows an example of such image correction related to face direction. The head model generation unit F21 determines whether the face in the face image to be three-dimensionalized faces is facing forward, and if it is not facing forward, performs image correction so that the face faces forward.
[0061] Generally, algorithms for creating three-dimensional facial images are based on the assumption that the facial image is facing forward, and this embodiment also employs such an algorithm. Therefore, by correcting the facial orientation as described above, the three-dimensional processing of the facial image can be performed appropriately.
[0062] Furthermore, in this embodiment, the head model generation unit F21 performs a process to remove bangs from a face image to be three-dimensionalized if the face has bangs. For example, if a face image with bangs, such as that shown in FIG. 7A, is three-dimensionalized, there is a risk that the bangs will be inherited by the texture of the three-dimensional head model, as shown in FIG. 7B. Therefore, the head model generation unit F21 in this embodiment determines whether bangs are present on the face in the face image to be three-dimensionalized, and if bangs are present, performs a process to remove the bangs from the face image (e.g., a process to mask them with skin color). Note that the determination of whether bangs are present can be made, for example, by determining whether a low-brightness area, such as black, exists in the forehead area identified by its positional relationship with a specific facial feature, such as the position of the eyebrows. FIG. 7C shows an example of an image after the bangs removal process.
[0063] By performing the above-described bangs removal process, it is possible to prevent unnecessary bangs from being carried over to the three-dimensional head model.
[0064] In addition, the head model generation unit F21 in this example performs image correction to close the mouth of a face image to be three-dimensionalized if the mouth is open in the image. For example, if a face image with an open mouth and exposed teeth is three-dimensionalized as shown in FIG. 8A, there is a risk that the image of the teeth will be superimposed on the lip texture in the three-dimensional head model as shown in FIG. 8B. Therefore, the head model generation unit F21 in this example determines whether the mouth of the face to be three-dimensionalized in the image is open, and if the mouth is open, performs image correction to close the mouth of the face image. FIG. 8C shows an example of an image after the mouth closing process.
[0065] By performing the above-described mouth closing process, it is possible to prevent the image of the teeth from being superimposed on the lip portion of the three-dimensional head model.
[0066] In addition, various image correction processes can be considered to ensure that the three-dimensional processing of facial images is performed appropriately, such as correction processes to optimize brightness and contrast, processing to remove glasses, and processing to correct the degree of eye opening.
[0067] Furthermore, in this embodiment, the head model generation unit F21 (eyeball addition unit F32) performs eyeball model fitting using landmark information for parts of the three-dimensional head model other than the eyes (for example, the nose, ears, mouth, etc.). When eyeball fitting is performed using a conventional algorithm, there are problems such as gaps at the corners of the eyes and a tendency for the eyes to cross, as shown in Fig. 9. In contrast, by performing fitting using landmark information for parts other than the eyes, it is possible to achieve natural eyeball fitting without gaps, as shown in Fig. 10.
[0068] Furthermore, in this embodiment, in the integration process of the three-dimensional head model and three-dimensional body model in the human body model generation unit F22 (body integration unit F34), deformation using a lattice modifier is applied, thereby preventing the mesh at the neck of the three-dimensional head model from protruding from the mesh at the neck of the three-dimensional body model, as exemplified by the comparison between Figures 11A and 11B.
[0069] Furthermore, in this embodiment, in response to the fact that some three-dimensional body models do not have skin texture from the neck to the shoulders, skin texture from the neck to the shoulders is added to the three-dimensional body model. If the three-dimensional body model does not have skin texture from the neck to the shoulders, the inside of the mesh may become transparent around the neck, as shown in Fig. 12A. However, by adding skin texture from the neck to the shoulders to the three-dimensional body model as described above, it is possible to prevent such transparency around the neck, as shown in Fig. 12B.
[0070] (3-3. Specific Learning Method Including Analysis of Attribute Distribution) As described above, in this embodiment, it is assumed that adjustment of learning data is performed when re-learning an AI model, and it is assumed that initial learning has already been performed on the AI model using clean data such as the ImageNet learning dataset. Re-learning of the AI model is assumed to involve repeating adjustment of learning data and learning processing using the adjusted learning data until a predetermined learning termination condition is met, such as the performance of the AI model satisfying a certain performance standard.
[0071] 13 shows an example of an initial distribution setting screen for setting a target distribution of learning data in the first data adjustment. This initial distribution setting screen is a screen presented to the user by the presentation processing unit F3, and the presentation processing unit F3 performs processing to display this initial distribution setting screen on the display unit (display unit 17) of the user terminal 3.
[0072] As shown in the figure, the attribute parameters that indicate the target distribution can be roughly divided into parameters related to human models and parameters related to scenes. Parameters related to human models include, for example, parameters such as age distribution, gender ratio, head length range, and head circumference range. Parameters related to scenes include, for example, parameters such as the number of cameras, camera angle, and brightness range.
[0073] In this embodiment, the initial distribution setting screen displays default parameters as various attribute parameters that indicate the target distribution of learning data in the initial data adjustment. For example, default parameter groups are prepared for each destination, such as Japan or the United States, and a corresponding parameter group is selected from these parameter groups and displayed as the default parameter group on the initial distribution setting screen. On the initial distribution setting screen, the parameters displayed as defaults can be changed by user operation, but by displaying the default parameters, the initial data adjustment can be performed even if the user does not have knowledge of data adjustment and is unable to set parameters themselves.
[0074] As shown in the figure, the initial distribution setting screen has a setting button B1, and in response to operating this setting button B1, the adjustment control unit F4 instructs the data generation unit F20 to generate learning data based on information about the target distribution according to the attribute parameter group displayed on the initial distribution setting screen. This realizes the initial data adjustment.
[0075] Here, in this embodiment, regardless of whether it is the first data adjustment or the second or subsequent data adjustment, when a target distribution is set, the adjustment control unit F4 controls the generation of learning data for some attributes within the target distribution before causing the data generation unit F20 to generate learning data for all attributes within the target distribution, and the presentation processing unit F3 performs processing to present the generated learning data for some attributes to the user. In other words, when a target distribution is set, a sample image before actual generation is presented. As such a sample image before actual generation, for example, it is possible to present an image with an average attribute within the target distribution and an image with an attribute corresponding to an outlier.
[0076] 14 shows an example of a sample presentation screen that presents a sample image before actual generation. As shown, a back button B2 and a confirm button B3 are arranged on the sample presentation screen, and in response to operation of the back button B2, the presentation processing unit F3 presents the initial distribution setting screen shown in FIG. 13 again to the user. This makes it possible to adjust the target distribution before actual generation of training data if the sample image is not what the user intended.
[0077] When the confirm button B3 is operated on the sample presentation screen, if it is the first time data is adjusted, the adjustment control unit F4 instructs the data generation unit F20 to generate training data by providing target distribution information based on the attribute parameter group displayed on the initial distribution setting screen. On the other hand, if it is the second or subsequent time data is adjusted, the adjustment control unit F4 instructs the data generation unit F20 to generate training data by providing target distribution information set on a target distribution setting screen (described later).
[0078] It is also conceivable that the process of presenting the sample image before actual generation may be performed only during the first data adjustment or during the second or subsequent data adjustments.
[0079] In this embodiment, the presentation processing unit F3 performs a process of presenting information indicating the results of the training data generation to the user in response to the data generation unit F20 generating training data according to the instructed target distribution, regardless of whether this is the first data adjustment or a second or subsequent data adjustment. Specifically, in this example, the presentation processing unit F3 performs a process of presenting information indicating the attribute distribution of the generated training data (see FIG. 15 ) and the images themselves as training data (see FIG. 16 ) to the user.
[0080] As can be understood from the above description, the data generation unit F20 generates training data with various attributes by randomly (i.e., probabilistically) selecting attributes so as to satisfy specified attribute distribution conditions. Therefore, the set target distribution and the attribute distribution of the actually generated training data do not necessarily match. Therefore, it is meaningful to present the attribute distribution of the actually generated training data to the user. Although not illustrated, the data generation unit F20 associates annotation information indicating each attribute with the generated training data. Based on this annotation information, the presentation processing unit F3 analyzes the attribute distribution of the generated training data from various perspectives, such as age distribution and gender ratio, and presents the analysis results to the user in a predetermined format, such as a graph or a table.
[0081] A "View Image" button B4 and a "Start Learning" button B5 are arranged on the attribute distribution presentation screen shown in Fig. 15. In response to an operation of the "View Image" button B4, the presentation processing unit F3 performs processing to present to the user a screen displaying a list of generated data shown in Fig. 16, i.e., a screen displaying a list of images as generated learning data.
[0082] A back button B6 is arranged on the generated data list screen shown in FIG. 16, and when the back button B6 is operated, the presentation processing unit F3 performs processing to return the screen presented to the user to the attribute distribution presentation screen shown in FIG.
[0083] When the start learning button B5 is operated on the attribute distribution presentation screen, the learning control unit F5 causes the learning processing unit F1 to execute a learning process for the AI model using the generated learning data.
[0084] When a learning process is performed as re-learning using the learning data obtained by data adjustment, the analysis unit F2 analyzes the attribute distribution of the learning data based on the above-mentioned intensity evaluation value. In this embodiment, for the second and subsequent data adjustments during re-learning, the intensity evaluation value is used to analyze data attributes that are insufficient for learning and data attributes that are considered excessive, and the data adjustments are performed in a manner that tends to supplement the insufficient attributes and exclude the excessive attributes, thereby improving the learning accuracy.
[0085] In this example, the analysis unit F2 calculates a value based on saliency as the fluctuation evaluation value. In this specification, "saliency" refers to a value obtained by averaging the weight coefficients in the convolution filter of the AI model. This saliency allows the degree of fluctuation in the weight coefficients of the convolution filter to be appropriately evaluated, thereby improving the accuracy of estimating the attributes of missing data and excess data.
[0086] For saliency, we use the same one as described in Reference 1 below. Reference 1: "Where do Models go Wrong? Parameter-Space Saliency Maps for Explainability" Roman Levin, Manli Shu, Eitan Borgnia, Furong Huang, Micah Goldblum, Tom Goldstein: arXiv:2108.01335
[0087] This saliency can be calculated for each convolution filter in an AI model, allowing for more diverse evaluations than a loss value (error function) that can only be calculated as a single value for the entire AI model. Furthermore, since it is a value that can be calculated for each convolution filter, saliency can be calculated in various units, such as for each pixel of the input image (learning data) or for each layer. When an AI model performs object recognition processing, saliency calculated for each pixel can also be used as a value indicating the magnitude of the influence that pixel has on the score of the predicted class.
[0088] Note that the number of convolution filters in an AI model may reach several thousand, and it may not be appropriate to use the saliency calculated for each filter as an evaluation value as is. Taking this into consideration, the intensity evaluation value may be calculated as a moving average value of the saliency for each filter. For example, when the number of convolution filters is several thousand, the number of sections may be set to 100, and the moving average value of the saliency for each filter may be calculated.
[0089] For example, the analysis unit F2 in this example calculates a saliency-based intensity evaluation value for each piece of training data input to the AI model during relearning, and estimates attributes that are lacking or excessive in the training based on the saliency evaluation value calculated for each piece of training data and attribute information associated with the training data as annotation information.
[0090] Here, attributes that are missing from learning and attributes that are deemed excessive can be estimated based on the magnitude of the violence assessment value. Figures 17A and 17B are explanatory diagrams of an example of estimating attributes of missing data, and Figure 17C is an explanatory diagram of an example of estimating attributes of excessive data. For example, as shown in Figure 17A, a violence assessment value is calculated for each attribute value of age, and the attribute value of an age with a higher violence assessment value than other ages is estimated as the attribute value of missing data. The example in the figure shows a case where the violence assessment value for teenagers is higher than for other age groups, and in this case, it can be estimated that data for teenagers is lacking in learning.
[0091] Also, as shown in FIG. 17B, for attribute values of a target object such as age, attribute values that are far from the statistical average and have a larger violence evaluation value than other attribute values can be estimated as attribute values of missing data.
[0092] In this way, the attribute of the missing data can be identified as an attribute whose violence evaluation value is larger than that of other attributes when comparing the violence evaluation value for each attribute from a certain point of view.
[0093] On the other hand, when comparing the violence evaluation values for each attribute from a certain perspective, the attribute of the excess data can be identified as an attribute whose violence evaluation value is smaller than that of other attributes. For example, Fig. 17C illustrates an example of the violence evaluation value calculated for each attribute value from the perspective of camera angle. In this case, the attribute value of the small angle region, whose violence evaluation value is smaller than that of other attribute values, can be identified as the attribute value of the excess data.
[0094] Furthermore, depending on the magnitude of the violence evaluation value, it is also possible to estimate the degree of deficiency or excess in learning. Specifically, for an attribute estimated to be deficient, the larger the violence evaluation value, the greater the degree of deficiency. For an attribute estimated to be excessive, the smaller the violence evaluation value, the greater the degree of excess.
[0095] The presentation processing unit F3 performs processing to present to the user information indicating the attributes of data that is insufficient and data that is deemed excessive in the learning analyzed by the analysis unit F2.
[0096] FIG. 18 shows an example of an excess / deficient data presentation screen for presenting information indicating the attributes of the missing data and the excess data to the user. As shown, the excess / deficient data presentation screen displays text information indicating the attributes of the missing data analyzed by the analysis unit F2 and text information indicating the attributes of the excess data (omitted in the illustrated example). Here, there are cases where either the missing data or the excess data cannot be estimated, and FIG. 18 shows an example screen in which the excess data cannot be estimated. In addition, the excess / deficient data presentation screen in this example also displays example images of undetectable scenes. As images of undetectable scenes, training data in which the AI model was unable to detect a target object and selected from training data with attributes that the analysis unit F2 estimated to be insufficient in training are displayed.
[0097] In the figure, examples of information indicating the attributes of the missing data include "insufficient learning data for the camera facing forward at a depression angle of 30 degrees," "insufficient learning data for long hair," and "insufficient learning data for men in their 20s." However, these are merely examples for explanatory purposes.
[0098] By presenting the user with information indicating the attributes of the missing or excess data as described above, even if the user does not have knowledge of which attributes are missing or excess in improving learning accuracy, the user can grasp the information indicating such missing or excess attributes, and furthermore, the user can perform operations related to setting a target for attribute distribution while being made aware of such missing or excess attributes. Therefore, it is possible to improve learning accuracy by adjusting the learning data.
[0099] Furthermore, in this embodiment, the analysis unit F2 analyzes the priority of the contribution of attributes of data that are insufficient or excessive in learning to improving learning accuracy. Specifically, the analysis unit F2 in this example analyzes the degree of insufficiency and the degree of excess as this priority. Attributes that are highly insufficient or excessive in learning can be said to have a high degree of contribution to improving learning accuracy. Therefore, analyzing the degree of insufficiency and the degree of excess corresponds to analyzing the priority of the contribution to improving learning accuracy.
[0100] The presentation processing unit F3 performs processing to present information indicating the attributes of data that is insufficient or excessive in learning in a manner according to the above priority. In this example, a screen presenting information indicating the attributes of the insufficient or excessive data in a manner according to priority is presented to the user as a setting screen for a target distribution in learning data adjustment (hereinafter referred to as a "target distribution setting screen").
[0101] Fig. 19 shows an example of a target distribution setting screen. In this example, the target distribution setting screen is presented as a transition screen from the excess / deficiency data presentation screen shown in Fig. 18. Specifically, the excess / deficiency data presentation screen of this example has an OK button B9 arranged thereon as shown in Fig. 18, and in response to operation of this OK button B9, the presentation processing unit F3 performs processing to present the target distribution setting screen shown in Fig. 19 to the user.
[0102] As shown in the figure, the target distribution setting screen displays information indicating the attribute distribution before and after correction, and the priority, for each attribute item of missing data and excess data. Here, an example is shown in which the priority is calculated using three levels of values: "high," "medium," and "low." However, the priority levels are not limited to three levels. For illustrative purposes, the figure shows an example in which the attributes of missing data are estimated to be camera depression angle = -91 to -120, camera depression angle = 91 to 120, brightness = 0 to 50, brightness = 151 to 200, and gender = male. Corresponding to this example, the attribute distributions before and after correction for camera depression angle, brightness, and gender (male / female ratio) are displayed as "-90 to 90," "-120 to 120," "50 to 150," "0 to 200," "5:5," and "7:3." The illustrated example shows a screen in which the priority of the attribute of camera depression angle is calculated as "high," the priority of the attribute of brightness as "medium," and the priority of the attribute of male as "low." Here, as an example of presentation in a manner according to priority, a case is shown in which information of attribute distributions with higher priority is displayed at a higher position. Note that examples of changes in presentation in accordance with priority are not limited to changes in display position, and changes in other elements such as display size and color are also possible.
[0103] By presenting information indicating the attribute distribution of excess and deficiency data in a manner according to priority on the target distribution setting screen in this way, the user can select which attribute distribution items to include in the target distribution using the presented priority as a guide. Specifically, even if the user does not have knowledge regarding the correlation between the attributes of the learning data and their degree of contribution to improving learning accuracy, it is possible to understand the degree of contribution to improving learning accuracy for attributes analyzed as insufficient or excess, and further, it is possible to allow the user to perform operations related to target setting of the attribute distribution after understanding such degree of contribution.
[0104] As shown in the figure, the target distribution setting screen has a check box for each attribute item. By checking the check box of a desired item, the user can select the attribute distribution of that item as the attribute distribution to be included in the target distribution. A confirm button B10 is provided on the target distribution setting screen, and by operating the confirm button B10, the user can specify the attribute distribution of the item selected by the check box as the attribute distribution to be included in the target distribution.
[0105] In addition, the target distribution setting screen may also display information such as an estimated value of the time required to execute the learning process ("Estimated execution time" in the figure) as guideline information regarding the learning process using the learning data after data adjustment.
[0106] Furthermore, the example in Figure 19 shows a case where the attributes of the excess data are not estimated, but if the attributes of the excess data are estimated, the attribute distribution items, information on the attribute distribution before and after correction, and priority can be displayed in a similar manner.
[0107] The adjustment control unit F4 sets, as a target distribution, an attribute distribution including an attribute selected by the user from among the attributes presented in a priority order on the target distribution setting screen. Then, the adjustment control unit F4 controls the data generation unit F20 to generate learning data (learning data group) having an attribute distribution according to the set target distribution. Specifically, the adjustment control unit F4 instructs the data generation unit F20 to generate learning data having an attribute distribution according to the set target distribution by providing attribute distribution designation information (see FIG. 4 ) according to the set target distribution.
[0108] This allows the adjustment of the learning data to be controlled so that the attribute distribution includes attributes selected by the user while taking into consideration the balance between improving learning accuracy and the user's own needs. Therefore, with regard to the adjustment of the learning data, it is possible to achieve both improvement of learning accuracy and response to the user's needs.
[0109] Furthermore, in this embodiment, the presentation processing unit F3 performs a process of presenting information indicating the evaluation result of the learning accuracy to the user each time the learning processing unit F1 performs learning of the AI model. FIG. 20 shows an example of an evaluation result presentation screen for presenting information indicating the evaluation result of the learning accuracy to the user. In the example of FIG. 20, a loss value (value of an error function) and a performance evaluation value of the AI model after learning are presented as information indicating the evaluation result of the learning accuracy. The performance evaluation value of the AI model is not limited to Accuracy (the percentage of correct predictions out of all predictions) shown in the figure, but other performance evaluation values such as Precision (the percentage of correct predictions by the model that were actually correct), Recall (the percentage of correct predictions by the model that were actually correct), and F-measure may also be presented. On the evaluation result presentation screen, the information indicating the evaluation result of the learning accuracy is not limited to numerical information, and information in a format other than numerical, such as graphed information, may also be presented.
[0110] As shown in the figure, the evaluation result presentation screen has a Continue Learning button B7 and an End Learning button B8, allowing the user to refer to the evaluation results and choose whether or not to perform further data adjustment and learning processing. When the Continue Learning button B7 is operated, the presentation processing unit F3 performs the presentation processing of the excess / deficiency data presentation screen illustrated in Figure 18 above, allowing the user to set the target distribution of the learning data to be used in the next learning processing. On the other hand, when the End Learning button B8 is operated, no further learning processing is performed and learning ends.
[0111] (3-4. Processing Procedure) A specific example of a processing procedure for realizing the learning method according to the embodiment described above will be described with reference to the flowchart in Fig. 21. In this example, the processing shown in Fig. 21 is executed by the CPU 11 of the information processing device 1 based on a program stored in a predetermined storage device such as the ROM 12 or the storage unit 19.
[0112] In step S101, the CPU 11 executes a process of presenting an initial distribution setting screen, i.e., a process of presenting to the user an initial distribution setting screen displaying default attribute parameters as described above with reference to FIG.
[0113] In step S102 following step S101, the CPU 11 performs a waiting process for a setting button operation, that is, a process for waiting for an operation of the setting button B1 arranged on the initial distribution setting screen.
[0114] When the setting button B1 is operated, the CPU 11 proceeds to step S103 and executes a sample presentation process before the actual data generation (see FIG. 14 ). The process of step S103 can be performed both during the first data adjustment and during subsequent data adjustments. During the first data adjustment, the process of step S103 causes the data generation unit F20 to generate learning data for some attributes within the target distribution based on the attribute parameter group displayed on the initial distribution setting screen, and presents a sample presentation screen including the generated learning data to the user. During subsequent data adjustments, the process of step S103 causes the data generation unit F20 to generate learning data for some attributes within the target distribution based on the target distribution set on the target distribution setting screen (see FIG. 19 ), and presents a sample presentation screen including the generated learning data to the user.
[0115] In step S104 following step S103, the CPU 11 waits for the operation of the decision button B3 arranged on the sample presentation screen as a waiting process for the decision operation.
[0116] When the enter button B3 is operated, the CPU 11 proceeds to step S105 and executes data generation control processing. That is, if data is being generated for the first time, the CPU 11 instructs the data generation unit F20 to generate training data having an attribute distribution according to the target distribution by specifying attribute distribution designation information (see FIG. 4 ) according to the target distribution based on the attribute parameter group displayed on the initial distribution setting screen. If data is being generated for the second or subsequent time, the CPU 11 instructs the data generation unit F20 to generate training data having an attribute distribution according to the target distribution by specifying attribute distribution designation information according to the target distribution set on the target distribution setting screen.
[0117] In step S106 following step S105, the CPU 11 performs information presentation processing for the generated data. That is, this is processing for presenting information indicating the results of the generation of learning data to the user. Specifically, this is processing for presenting information indicating the attribute distribution of the learning data described above in FIG. 15 and the images themselves (see FIG. 16) serving as learning data to the user. Note that the specific manner in which this information is presented (such as screen transitions in response to the operation of the View Image button B4) has been described above, and therefore a duplicate description will be avoided.
[0118] In step S107 following step S106, the CPU 11 determines whether or not learning has started. Specifically, the CPU 11 waits for operation of the learning start button B5 on the attribute distribution presentation screen shown in FIG.
[0119] If it is determined in step S107 that the learning start button B5 has been operated and learning has started, the CPU 11 proceeds to step S108 and executes learning processing using the generated data. That is, learning processing of the AI model is performed using the learning data generated by the data generation unit F20 in the data generation control processing of step S105. Here, as described above, in this example, the violence evaluation value is calculated for each piece of learning data, so the violence evaluation value is calculated during the learning processing of step S108. During the learning processing, the learner F11 also calculates the loss value described above.
[0120] In step S109 following step S108, the CPU 11 executes a verification process for the trained model. Specifically, the trained AI model executes an inference process using the verification data as input data, and calculates a performance evaluation value of the AI model, such as accuracy.
[0121] In step S110 following step S109, the CPU 11 performs processing for presenting the learning accuracy evaluation result, i.e., presenting to the user an evaluation result presentation screen that displays evaluation information related to the learning accuracy, such as the loss value and accuracy, as exemplified in Fig. 20 .
[0122] In step S111 following step S110, the CPU 11 determines whether the learning has ended. Specifically, it determines whether the learning continue button B7 or the learning end button B8 arranged on the evaluation result presentation screen has been operated.
[0123] If the learning continuation button B7 is operated and it is determined that learning has not ended, the CPU 11 proceeds to step S112 and executes a process for presenting information on missing or excess attributes. Specifically, in this example, a process is performed to present to the user an excess / deficient data presentation screen that displays information on the attributes of missing data and excess data, as illustrated in Fig. 18. To present this excess / deficient data presentation screen, the CPU 11 performs a process for estimating the attributes of missing data and excess data based on the annotation information and the intensity evaluation value of the learning data.
[0124] In step S113 following step S112, the CPU 11 executes a process for presenting a target distribution setting screen. That is, as illustrated in FIG. 19 , a process for presenting to the user a target distribution setting screen that displays the attribute distributions of the missing data and the excess data in a manner according to priority is performed. In this example, the process for presenting the target distribution setting screen in step S113 is executed in response to operation of the OK button B9 arranged on the evaluation result presentation screen. In addition, in order to present the target distribution setting screen, the CPU 11 calculates priorities for the attributes of the missing data and the excess data based on the intensity evaluation values.
[0125] In step S114 following step S113, the CPU 11 waits for the operation of the decision button B10 arranged on the target distribution setting screen as a waiting process for the decision operation.
[0126] If it is determined in step S114 that the decision button B10 has been operated and a decision operation has been performed, the CPU 11 returns to step S103. As a result, a sample presentation process before actual generation is executed for the second and subsequent data adjustments. Subsequently, by the processes of steps S104 to S110, for the second and subsequent data adjustments, an information presentation process for the adjusted learning data (S106), a re-learning process using the adjusted learning data (S108), and a presentation process for the learning accuracy evaluation results (S110) are executed.
[0127] If the CPU 11 determines in step S111 that the learning end button B8 has been operated to end the learning, it ends the series of processes shown in FIG.
[0128] Although the above example shows the calculation of a violence evaluation value for each piece of learning data input to the AI model during relearning, it is also possible to group multiple pieces of learning data with the same or similar attributes and calculate a violence evaluation value for each group. This can improve the efficiency of the process of estimating the attributes of excess or deficiency data.
[0129] 4. Modifications Note that the embodiment is not limited to the specific example described above, and various modified configurations may be adopted. For example, in the above example, the generation of learning data is performed in a device (data generation device 2) separate from the information processing device (information processing device 1) that performs analysis related to the attribute portion based on the violence evaluation value. However, it is also possible to generate learning data in the information processing device. FIG. 22 shows the configuration of an information processing device 1A that has the function of a data generation unit F20.
[0130] Furthermore, in the above example, an information processing device that performs analysis of the attribute portion based on the violence evaluation value performs the learning process of the AI model. However, it is also possible to adopt a configuration in which the learning process of the AI model is performed by a device separate from the information processing device. FIG. 23 illustrates an example configuration of a learning system in this case. As shown in the figure, in this learning system, a learning device 5 having a learning processing unit F1 is added. Furthermore, as the information processing device that performs analysis of the attribute portion based on the violence evaluation value, an information processing device 1B having a configuration in which the learning processing unit F1 is omitted from the information processing device 1 is used. Note that the information processing device 1B can also be configured to include a data generation unit F20, similar to the information processing device 1A shown in FIG. 22.
[0131] Furthermore, although the above provides an example of applying the learning method according to the present technology when an AI model performs inference processing on image data, the present technology can be widely and suitably applied to the learning of AI models that perform inference processing on data other than images, such as sound data.
[0132] 5. Summary of the Embodiments As described above, the information processing device (1, 1A, 1B) according to the embodiment includes an analysis unit (F2) that performs an analysis of an attribute distribution, which is a statistical data distribution of attributes of training data used to train an AI model, based on a turbulence evaluation value indicating the degree of turbulence of the weight coefficients of a convolution filter when the training data is input to the AI model; a presentation processing unit (F3) that performs processing to present the analysis results by the analysis unit to a user; and an adjustment control unit (F4) that controls adjustment of the training data based on a target distribution of the attribute distribution set based on user operation related to the analysis results presented by the presentation processing unit. By using the turbulence evaluation value indicating the degree of turbulence of the weight coefficients of the convolution filter, in other words, the responsiveness of the convolution filter, it is possible to appropriately evaluate the correctness of the operation of the AI model. The presentation processing unit and adjustment control unit then present the analysis results of the attribute distribution based on the turbulence evaluation value, which can appropriately evaluate the correctness of the operation of the AI model, and adjust the training data based on user operation related to the presented analysis results. Therefore, it is possible to realize a method for adjusting the learning data used in AI model learning that can improve learning accuracy while also meeting user requests.
[0133] In the information processing device according to the embodiment, the analysis unit analyzes attributes of data that are insufficient or excessive in learning as analysis related to attribute distribution. This makes it possible to obtain information that is useful for improving the learning accuracy of the AI model as analytical information related to attribute distribution.
[0134] Furthermore, in the information processing device according to the embodiment, the presentation processing unit performs processing to present to the user information indicating the attributes of data that are insufficient or excessive in the learning analyzed by the analysis unit. This allows the user to grasp the information indicating such insufficient or excessive attributes even if the user does not have knowledge of which attributes are insufficient or excessive for improving learning accuracy. Furthermore, it allows the user to perform operations related to setting a target for attribute distribution while being aware of such insufficient or excessive attributes. Therefore, it is possible to improve learning accuracy by adjusting the learning data.
[0135] Furthermore, in an information processing device according to an embodiment, the analysis unit analyzes the priority of the contribution of attributes of data that are insufficient or excessive in learning to improving learning accuracy, and the presentation processing unit presents information indicating the attributes of data that are insufficient or excessive in learning in a manner according to the priority. This allows the user to understand the degree of contribution of attributes analyzed as insufficient or excessive to improving learning accuracy, even if the user does not have knowledge of the correlation between the attributes of the learning data and their contribution to improving learning accuracy. Furthermore, it allows the user to perform operations related to setting a target for the attribute distribution while understanding the degree of contribution. This allows the user to perform operations related to setting a target for the attribute distribution while taking into account the balance between improving learning accuracy and their own needs.
[0136] Furthermore, in the information processing device according to the embodiment, the adjustment control unit sets, as a target distribution, an attribute distribution that includes an attribute selected by the user from among the attributes presented by the presentation processing unit in a manner according to priority. This allows adjustment of the learning data to be controlled so that the attribute distribution includes the attribute selected by the user, taking into consideration a balance between improving learning accuracy and the user's own needs. Therefore, with regard to the adjustment of the learning data, it is possible to achieve both improved learning accuracy and responsiveness to the user's needs.
[0137] Furthermore, in the information processing device according to the embodiment, the adjustment control unit controls the generation of learning data for some attributes within a target distribution for an attribute distribution before generating learning data for all attributes within the target distribution, and the presentation processing unit performs processing to present the generated learning data for some attributes to the user. As described above, by generating learning data for only some attributes and presenting it to the user before generating learning data for all attributes within the target distribution, the user can check whether the target distribution is an unintended distribution before generating learning data for all attributes, thereby preventing the generation of learning data for all attributes from being wasted.
[0138] An information processing method as an embodiment is an information processing method in which an information processing device analyzes an attribute distribution, which is a statistical data distribution of attributes of training data used for training an AI model, based on a fluctuation evaluation value that indicates the degree of fluctuation in the weight coefficients of a convolution filter when the training data is input to the AI model, performs processing to present the analysis results to a user, and controls so that the training data is adjusted based on a target distribution of the attribute distribution that is set based on user operations related to the presented analysis results.With this information processing method, it is possible to obtain the same functions and effects as the information processing device as the above-mentioned embodiment.
[0139] A data generation device (see 2 or information processing device 1A) according to an embodiment includes a head model generation unit (see F21) that generates a three-dimensional head model by three-dimensionalizing a generated facial image; a human body model generation unit (see F22) that generates a three-dimensional human body model by integrating the three-dimensional head model with a three-dimensional body model, which is a three-dimensional model of the body; and an imaging processing unit (see F23) that two-dimensionalizes the three-dimensional human body model to generate training data using CG images. Generating training data using CG images eliminates the need to capture images of each attribute to obtain training data with a desired attribute distribution, thereby improving the efficiency of data adjustment work for the training data. In this case, by using the "generated facial image" as the original data for the head model, the subject of the training data can be a person who is highly realistic but does not actually exist. Therefore, by including subjects with highly realistic faces in the training data, it is possible to improve learning accuracy while avoiding privacy violations.
[0140] In the data generation device according to the embodiment, if the face in the facial image faces in a direction other than the front, the head model generation unit corrects the image so that the face faces forward, thereby enabling the facial image to be three-dimensionally processed appropriately.
[0141] Furthermore, in the data generation device according to the embodiment, the head model generation unit performs processing to remove bangs from the face in the facial image if the bangs are present. This allows the facial image to be properly three-dimensionalized. Specifically, it prevents unnecessary bangs from being carried over to the three-dimensional head model.
[0142] Furthermore, in the data generation device according to the embodiment, the head model generation unit performs image correction to close the mouth when the mouth is open in the facial image. This allows the facial image to be properly three-dimensionalized. Specifically, it prevents the image of teeth from being superimposed on the lip portion of the three-dimensional head model.
[0143] In addition, in the data generation device according to the embodiment, the head model generation unit generates multiple types of 3D head models that are combined with different types of eyeball models or hair models. This makes it possible to generate multiple types of 3D human body models that differ in at least the features of the eyeballs or hair. In other words, it becomes possible to generate multiple types of 3D human body models that differ in at least the attributes related to the eyeballs or hair, thereby increasing the variety of attributes in the training data.
[0144] Furthermore, in the data generation device according to the embodiment, the human body model generation unit generates multiple types of 3D human body models with different types of 3D body models to be integrated into the 3D head model. This makes it possible to generate multiple types of 3D human body models with different body features, thereby increasing the variety of attributes of the training data.
[0145] In a data generation method according to an embodiment, an information processing device generates a three-dimensional head model by three-dimensionally converting a generated facial image, integrates the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body to generate a three-dimensional human body model, and then two-dimensionalizes the three-dimensional human body model to generate training data using CG images. This data generation method can also achieve the same functions and effects as the data generation device according to the embodiment described above.
[0146] The learning system according to an embodiment includes a data generation unit (F20) having a head model generation unit that generates a three-dimensional head model by three-dimensionalizing a generated facial image, a human body model generation unit that generates a three-dimensional human body model by integrating the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body, and an imaging processing unit that two-dimensionalizes the three-dimensional human body model to generate learning data using CG images; an analysis unit that performs an analysis of an attribute distribution, which is a statistical data distribution of attributes of the learning data used to train the AI model, based on an intensity evaluation value that indicates the degree of intensity fluctuation of the weight coefficient of a convolution filter when the learning data is input to the AI model; a presentation processing unit that performs processing to present the analysis results by the analysis unit to a user; an adjustment control unit that controls adjustment of the learning data using the data generation unit based on a target distribution of the attribute distribution set based on user operation related to the analysis results presented by the presentation processing unit; and a learning processing unit (F1) that performs learning processing of the AI model using the adjusted learning data. The above-described learning system enables the adjustment of the learning data used to train the AI model to improve learning accuracy while also meeting user requests. Furthermore, by generating training data using CG images, it is no longer necessary to capture images of each attribute in order to obtain training data with a desired attribute distribution, which makes it possible to improve the efficiency of the data adjustment work for the training data.In this case, by using the "generated face image" to generate the head model, it is possible to make the subject of the training data a person who is highly realistic but does not actually exist.Therefore, it is possible to improve the learning accuracy by including subjects with highly realistic faces in the training data, while also avoiding privacy violations.
[0147] Here, as an embodiment, a program that causes, for example, a CPU, a DSP, or a device including these, to implement the processing described with reference to Figure 21 etc. can be considered. That is, a first program of the embodiment is a program readable by a computer device, and causes the computer device to implement the following functions: analyze an attribute distribution, which is a statistical data distribution of attributes of training data used for training an AI model, based on a fluctuation evaluation value indicating the degree of fluctuation in the weight coefficients of a convolution filter when the training data is input to the AI model; present the analysis results to a user; and control the adjustment of the training data based on a target distribution of the attribute distribution set based on user operation related to the presented analysis results. Such a program allows the functions of the information processing device 1 as the above-mentioned embodiment to be implemented in a computer device.
[0148] As another embodiment, a program that causes, for example, a CPU, a DSP, or a device including these, to implement the processing described with reference to Figure 4, etc., can be considered. That is, a second program of the embodiment is a program readable by a computer device, and causes the computer device to implement the functions of three-dimensionally converting a generated facial image to generate a three-dimensional head model, integrating the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body to generate a three-dimensional human body model, and two-dimensionalizing the three-dimensional human body model to generate training data using CG images. Such a program allows the functions of the data generation device 2 of the above-described embodiment to be implemented in the computer device.
[0149] The above-described programs can be pre-recorded on a hard disk drive (HDD) or solid state drive (SSD) as a recording medium built into a computer or other device, or on a ROM within a microcomputer having a CPU. Alternatively, the programs can be temporarily or permanently stored (recorded) on a removable recording medium such as a flexible disk, a CD-ROM (Compact Disc Read Only Memory), a Magneto Optical (MO) disc, a Digital Versatile Disc (DVD), a Blu-ray Disc (Blu-ray Disc (registered trademark)), a magnetic disk, a semiconductor memory, or a memory card. Such removable recording media can be provided as so-called packaged software. Furthermore, such programs can be installed on a personal computer or the like from a removable recording medium, or can be downloaded from a download site via a network such as a LAN or the Internet.
[0150] Furthermore, such a program is suitable for providing a wide range of data adjustment control methods and data generation methods according to embodiments, and can cause various types of information processing devices to function as devices that realize the data adjustment control method and data generation method of the present disclosure.
[0151] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0152] <6. The Present Technology> The present technology may also have the following configuration. (1) An information processing device comprising: an analysis unit that performs an analysis of an attribute distribution, which is a statistical data distribution of attributes of learning data used in learning an AI model, based on an intensity evaluation value that indicates the degree of intensity fluctuation of a weight coefficient of a convolution filter when the learning data is input to the AI model; a presentation processing unit that performs processing to present the analysis results by the analysis unit to a user; and an adjustment control unit that controls adjustment of the learning data based on a target distribution of the attribute distribution that is set based on a user operation related to the analysis results presented by the presentation processing unit. (2) The information processing device described in (1), in which the analysis unit analyzes attributes of data that are insufficient or excessive in learning as the analysis of the attribute distribution. (3) The information processing device described in (2), in which the presentation processing unit performs processing to present to a user information indicating the attributes of the data that are insufficient or excessive in the learning analyzed by the analysis unit. (4) The information processing device according to any one of (1) to (5), wherein the analysis unit analyzes priorities of attributes of data that are insufficient or excessive in the learning in terms of their contribution to improving learning accuracy, and the presentation processing unit performs processing to present information indicating attributes of data that are insufficient or excessive in the learning in a manner according to the priorities. (5) The information processing device according to (4), wherein the adjustment control unit sets, as the target distribution, an attribute distribution including an attribute selected by a user from the attributes presented by the presentation processing unit in a manner according to the priorities. (6) The information processing device according to any one of (1) to (5), wherein the adjustment control unit controls to generate learning data for some attributes within a target distribution for the attribute distribution before generating learning data for all attributes within the target distribution, and the presentation processing unit performs processing to present the generated learning data for the some attributes to a user.(7) An information processing method in which an information processing device performs an analysis of an attribute distribution, which is a statistical data distribution for attributes of learning data used for learning an AI model, based on a fluctuation evaluation value indicating the degree of fluctuation in weighting coefficients of a convolution filter when the learning data is input to the AI model, performs processing to present the analysis results to a user, and controls so that the learning data is adjusted based on a target distribution of the attribute distribution set based on a user operation related to the presented analysis results. (8) A program readable by a computer device, which causes the computer device to realize functions of performing an analysis of an attribute distribution, which is a statistical data distribution for attributes of learning data used for learning an AI model, based on a fluctuation evaluation value indicating the degree of fluctuation in weighting coefficients of a convolution filter when the learning data is input to the AI model, performs processing to present the analysis results to a user, and controls so that the learning data is adjusted based on a target distribution of the attribute distribution set based on a user operation related to the presented analysis results. (9) A data generation device comprising: a head model generation unit that generates a three-dimensional head model by three-dimensionalizing a generated face image; a human body model generation unit that generates a three-dimensional human body model by integrating the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body; and an imaging processing unit that two-dimensionalizes the three-dimensional human body model and generates learning data using CG images. (10) The data generation device described in (9), wherein the head model generation unit performs image correction so that the face in the face image faces in a direction other than forward, so that the face faces forward. (11) The data generation device described in (9) or (10), wherein the head model generation unit performs processing to remove bangs from the face in the face image if the face has bangs. (12) The data generation device described in any of (9) to (11), wherein the head model generation unit performs image correction so that the mouth is closed if the mouth is open in the face image. (13) The data generation device according to any one of (9) to (12), wherein the head model generation unit generates a plurality of types of the three-dimensional head models that are combined with different types of eyeball models or hair models.(14) The data generation device according to any of (9) to (13), wherein the human body model generation unit generates a plurality of types of three-dimensional human body models to be integrated into the three-dimensional head model, each of which has a different type of three-dimensional body model. (15) A data generation method in which an information processing device three-dimensionalizes a generated face image to generate a three-dimensional head model, integrates the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body to generate a three-dimensional human body model, and two-dimensionalizes the three-dimensional human body model to generate training data in the form of CG images. (16) A program readable by a computer device, causing the computer device to realize functions of three-dimensionalizing a generated face image to generate a three-dimensional head model, integrating the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body to generate a three-dimensional human body model, and two-dimensionalizing the three-dimensional human body model to generate training data in the form of CG images. (17) A learning system comprising: a data generation unit having a head model generation unit that three-dimensionalizes a generated face image to generate a three-dimensional head model; a human body model generation unit that generates a three-dimensional human body model by integrating the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body; and an imaging processing unit that two-dimensionalizes the three-dimensional human body model to generate learning data using CG images; an analysis unit that performs analysis of an attribute distribution, which is a statistical data distribution for attributes of learning data used for learning an AI model, based on a fluctuation evaluation value that indicates the degree of fluctuation in the weight coefficient of a convolution filter when the learning data is input to the AI model; a presentation processing unit that performs processing to present the analysis results by the analysis unit to a user; an adjustment control unit that controls adjustment of the learning data using the data generation unit based on a target distribution of the attribute distribution set based on user operation related to the analysis results presented by the presentation processing unit; and a learning processing unit that performs learning processing of the AI model using the adjusted learning data.
[0153] 1, 1A, 1B Information processing device 2 Data generation device 3 User terminal NT Network 11 CPU F1 Learning processing unit F11 Learning device F2 Analysis unit F3 Presentation processing unit F4 Adjustment control unit F5 Learning control unit F20 Data generation unit F21 Head model generation unit F22 Human body model generation unit F23 Imaging processing unit F30 Facial image generation unit F31 Three-dimensionalization unit F32 Eyeball addition unit F33 Hair addition unit F34 Body integration unit F35 Selection unit F36, F37 Selection and adjustment unit MD1 Eyeball model group MD2 Hair model group MD3 Body model group 5 Learning device
Claims
1. An information processing device comprising: an analysis unit that performs analysis of an attribute distribution, which is a statistical data distribution for the attributes of learning data used to train an AI model, based on a fluctuation evaluation value that indicates the degree of fluctuation in the weight coefficients of a convolution filter when the learning data is input to the AI model; a presentation processing unit that processes to present the analysis results by the analysis unit to a user; and an adjustment control unit that controls adjustment of the learning data based on a target distribution of the attribute distribution that is set based on user operation related to the analysis results presented by the presentation processing unit.
2. The information processing device according to claim 1, wherein the analysis unit analyzes attributes of data that are insufficient or excessive in learning as the analysis related to the attribute distribution.
3. The information processing device according to claim 2, wherein the presentation processing unit performs processing to present to a user information indicating attributes of data that is insufficient or excessive in the learning analyzed by the analysis unit.
4. The information processing device described in claim 3, wherein the analysis unit analyzes the priority of the attributes of data that are insufficient or excessive in the learning in terms of their contribution to improving learning accuracy, and the presentation processing unit performs processing to present information indicating the attributes of data that are insufficient or excessive in the learning in a manner according to the priority.
5. The information processing device according to claim 4, wherein the adjustment control unit sets, as the target distribution, an attribute distribution including an attribute selected by the user from among the attributes presented by the presentation processing unit in a manner according to the priority.
6. The information processing device according to claim 1, wherein the adjustment control unit controls the generation of learning data for some attributes within a target distribution before generating learning data for all attributes within the target distribution for the attribute distribution, and the presentation processing unit performs processing to present the generated learning data for some attributes to a user.
7. An information processing method in which an information processing device performs an analysis of an attribute distribution, which is a statistical data distribution of attributes of learning data used to train an AI model, based on a fluctuation evaluation value that indicates the degree of fluctuation in the weight coefficient of a convolution filter when the learning data is input to the AI model, performs a process to present the analysis results to a user, and controls so that the learning data is adjusted based on a target distribution of the attribute distribution that is set based on user operation related to the presented analysis results.
8. A computer-readable program that causes the computer to perform the following functions: analyze attribute distribution, which is a statistical data distribution for the attributes of training data used in training an AI model, based on a fluctuation evaluation value that indicates the degree of fluctuation in the weight coefficient of a convolution filter when the training data is input to the AI model; present the analysis results to a user; and control the adjustment of the training data based on a target distribution of the attribute distribution that is set based on user operations related to the presented analysis results.
9. A data generation device comprising: a head model generation unit that three-dimensionalizes a generated facial image to generate a three-dimensional head model; a human body model generation unit that integrates the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body to generate a three-dimensional human body model; and an imaging processing unit that two-dimensionalizes the three-dimensional human body model to generate learning data in the form of CG images.
10. The data generation device according to claim 9, wherein, if the face in the facial image faces in a direction other than the front, the head model generation unit corrects the image so that the face faces forward.
11. The data generation device according to claim 9, wherein the head model generation unit performs processing to remove bangs from the face in the facial image if such bangs are present.
12. The data generation device according to claim 9, wherein the head model generation unit performs image correction to close the mouth when the mouth of the face in the face image is open.
13. The data generation device according to claim 9, wherein the head model generation unit generates a plurality of types of the three-dimensional head models that are combined with different types of eyeball models or hair models.
14. The data generation device according to claim 9, wherein the human body model generation unit generates a plurality of types of three-dimensional human body models, each of which has a different type of three-dimensional body model to be integrated into the three-dimensional head model.
15. A data generation method in which an information processing device three-dimensionalizes a generated facial image to generate a three-dimensional head model, integrates the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body to generate a three-dimensional human body model, and two-dimensionalizes the three-dimensional human body model to generate learning data in the form of CG images.
16. A computer-readable program that causes a computer to realize the following functions: converting a generated facial image into a three-dimensional image to generate a three-dimensional head model; integrating the three-dimensional head model with a three-dimensional body model to generate a three-dimensional human body model; and converting the three-dimensional human body model into a two-dimensional image to generate learning data using CG images.
17. A learning system comprising: a data generation unit having a head model generation unit that three-dimensionalizes a generated facial image to generate a three-dimensional head model; a human body model generation unit that generates a three-dimensional human body model by integrating the three-dimensional head model with a three-dimensional body model that is a three-dimensional model of the body; and an imaging processing unit that two-dimensionalizes the three-dimensional human body model to generate learning data in the form of CG images; an analysis unit that performs analysis of an attribute distribution, which is a statistical data distribution of attributes of learning data used for learning an AI model, based on a fluctuation evaluation value that indicates the degree of fluctuation in the weight coefficient of a convolution filter when the learning data is input to the AI model; a presentation processing unit that performs processing to present the analysis results by the analysis unit to a user; an adjustment control unit that controls adjustment of the learning data using the data generation unit based on a target distribution of the attribute distribution set based on user operation related to the analysis results presented by the presentation processing unit; and a learning processing unit that performs learning processing for the AI model using the adjusted learning data.
Citation Information
Patent Citations
Device and method for processing image, and recording medium
JP2001319245A
Device and method for controlling display and computer readable recording medium with display control program recorded thereon
JP2002032785A
Automatic face-tracking system and automatic face- tracking method
JP2002269546A
Information processing method, information processing system, and information processing program
JP2024030579A