Information processing device, information processing method, and program
The information processing apparatus addresses the inefficiencies in existing methods for improving AI model biases by using an integrated system to extract attributes, label data, detect biases, and adjust datasets within the AI processing apparatus, thereby enhancing prediction accuracy and efficiency.
Patent Information
- Application Number
- PCT/JP2024/039003
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-11-01
- Publication Date
- 2025-06-19
AI Technical Summary
Existing AI models learn biases from datasets, leading to misidentification of races and negative evaluations of genders, and current methods for improving bias are inefficient or unable to determine the source of bias.
An information processing apparatus and method that includes an attribute extraction unit, a labeling unit, a bias detection unit, and a dataset adjustment unit to extract attributes, label data, detect biases, and adjust the dataset to improve biases detected in AI models.
The solution efficiently improves biases in AI models by directly addressing the dataset biases, enhancing the accuracy of AI model predictions, and allowing for automatic detection and improvement of biased attributes without repeated dataset generation.
Smart Images

Figure JP2024039003_19062025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] The present technology relates to an information processing device, an information processing method, and a program.
[0002] Advances in the field of artificial intelligence (AI) have led to the increasing use of AI in a variety of fields. However, it is known that existing AI models (learning models) learn various biases. One cause of bias is the inclusion of bias in the training dataset (hereinafter referred to as the dataset as appropriate) from which the AI model learns. If a dataset contains bias, an AI model obtained using the dataset may mistakenly recognize a specific race as an animal or give a negative evaluation to a specific gender. Therefore, methods for improving such bias have been proposed (see, for example, Patent Documents 1 and 2 listed below).
[0003] JP 2020-2255 A JP 2021-189553 A
[0004] The technology described in Patent Literature 1 has a problem in that it is inefficient because it is necessary to repeatedly generate a dataset to improve the bias. Also, the technology described in Patent Literature 2 is a method for reducing bias during learning of an AI model, but has a problem in that it is not possible to determine whether the bias is caused by the model or the dataset.
[0005] An object of the present technology is to provide an information processing device, an information processing method, and a program that can improve bias corresponding to at least one attribute included in a dataset.
[0006] The present technology is an information processing device having, for example, an attribute extraction unit that extracts attributes from an input dataset; a labeling unit that labels attributes for each piece of data that constitutes the dataset; a bias detection unit that detects bias for each attribute based on the labeling results; and a dataset adjustment unit that adjusts the dataset to improve at least one bias detected by the bias detection unit.
[0007] This technology is an information processing method in which, for example, an attribute extraction unit extracts attributes from an input dataset; a labeling unit labels attributes for each piece of data that makes up the dataset; a bias detection unit detects bias for each attribute based on the labeling results; and a dataset adjustment unit adjusts the dataset to improve at least one bias detected by the bias detection unit.
[0008] This technology is a program that causes a computer to execute an information processing method, for example, in which an attribute extraction unit extracts attributes from an input dataset; a labeling unit labels attributes for each piece of data that makes up the dataset; a bias detection unit detects bias for each attribute based on the labeling results; and a dataset adjustment unit adjusts the dataset to improve at least one bias detected by the bias detection unit.
[0009] 1 is a block diagram illustrating an example of the configuration of an information processing apparatus according to a first embodiment; FIG. 2 is a diagram illustrating an overview of processing performed in the information processing apparatus according to the first embodiment; FIG. 3 is a flowchart illustrating the flow of processing performed in the information processing apparatus according to the first embodiment; FIG. 4 is a diagram illustrating a first display example of a labeling result by the labeling unit; FIG. 5 is a diagram illustrating a second display example of a labeling result by the labeling unit; FIGS. 6A to 6C are diagrams illustrating a display example of a bias detection result by the bias detection unit; FIGS. 7A to 7C are diagrams referred to when describing processing by the dataset adjustment unit; FIGS. 8A to 8C are diagrams referred to when describing processing by the dataset adjustment unit; FIGS. 9A to 9C are diagrams referred to when describing processing by the dataset adjustment unit; FIGS. 10A to 10C are diagrams referred to when describing processing by the dataset adjustment unit; FIGS. 11A to 11C are diagrams referred to when describing processing by the dataset adjustment unit; FIGS. 12A to 12C are diagrams referred to when describing processing by the dataset adjustment unit; 19 is a block diagram for explaining a configuration example of an information processing device according to a fourth embodiment. It is a schematic block diagram showing the overall configuration of an information processing system according to an example of the present disclosure. It is a block diagram showing the configuration of each device that registers or downloads an artificial intelligence (AI) model or an AI application via a marketplace function provided in an information processing device on the cloud side in the information processing system shown in FIG. 18. It is a flowchart showing the flow of a part of the process executed by each device when registering or downloading an AI model or an AI application via the marketplace function. It is a flowchart showing the flow of another part of the process executed by each device when registering or downloading an AI model or an AI application via the marketplace function.1 is a block diagram illustrating a connection mode between a cloud-side information processing device and an edge-side information processing device. FIG. 2 is a block diagram illustrating an example configuration of a cloud-side information processing device. FIG. 3 is a block diagram illustrating an example internal configuration of a camera as an imaging device. FIG. 4 is a schematic configuration diagram illustrating an example configuration of an image sensor as an imaging device. FIG. 4 is a block diagram illustrating an example software configuration of an imaging device. FIG. 5 is a block diagram illustrating an operating environment of a container when container technology is used. FIG. 6 is a block diagram illustrating an example hardware configuration of an information processing device. FIG. 7 is a diagram illustrating an imaging device and an imaging device system to which the present technology can be applied. FIG. 8 is a diagram illustrating an imaging device and an imaging device system to which the present technology can be applied. FIG. 9 is a diagram illustrating an imaging device and an imaging device system to which the present technology can be applied. FIG. 10 is a diagram illustrating a first application example of a cloud-side information processing device. FIG. 11 is a diagram illustrating a second application example of a cloud-side information processing device. FIGS. 1A and 1B are diagrams illustrating a third application example of a cloud-side information processing device. FIG. 11 is a diagram illustrating a fourth application example of a cloud-side information processing device. FIG. 12 is a diagram illustrating a modified example of the fourth application example of a cloud-side information processing device. FIG. 13 is a diagram illustrating a fifth application example of a cloud-side information processing device.
[0010] Hereinafter, embodiments of the present technology will be described with reference to the drawings. The description will be given in the following order. In this specification and the drawings, components having substantially the same functions or configurations will be denoted by the same reference numerals, and duplicated descriptions will be omitted as appropriate. <First embodiment> <Second embodiment> <Third embodiment> <Fourth embodiment> <Fifth embodiment> <Example> <Modification>
[0011] First Embodiment [Configuration Example of Information Processing Apparatus] Fig. 1 is a block diagram showing a configuration example of an information processing apparatus (information processing apparatus 1) according to a first embodiment. One example of the information processing apparatus 1 is a cloud server, but the information processing apparatus 1 is not limited to this. The information processing apparatus 1 may also be an edge device, a fog server located between the edge device and the cloud server, or the like. Specific examples of the cloud server, edge device, and fog server will be described later.
[0012] A dataset is input to the information processing device 1. The dataset is made up of multiple pieces of data (a large amount of data). The dataset may be stored in a storage device possessed by the information processing device 1, or may be supplied from outside the information processing device 1 via a network such as the Internet. Note that the type of data constituting the dataset is not particularly limited, but the following description will be given using image data as an example of data. Furthermore, the description will be given assuming that each piece of image data is annotated with the attribute "male" or "female."
[0013] The information processing device 1 includes, for example, an attribute extraction unit 2, a labeling unit 3, a bias detection unit 4, a data set adjustment unit 5, and an output unit 6.
[0014] The attribute extraction unit 2 extracts attributes from the input dataset. While the attribute extraction method is not limited to a specific method, in this embodiment, CLIP2StyleGAN (Generative Adversarial Network) is used as an example. CLIP2StyleGAN is a technique that, broadly speaking, extracts features from image data constituting a dataset using a CLIP encoder, then applies PCA (Principal Component Analysis) to obtain the i-th principal component direction, and automatically extracts (acquires) linguistic expressions corresponding to changes in edit-directions corresponding to the i-th principal component direction, i.e., attributes. Automatic extraction of attributes is made possible by searching the vocabulary of CLIP for vocabulary similar to the semantic concept of the i-th principal component. Automatically extracted attributes are those that are human-interpretable and capture the characteristics of the dataset. Note that semantic concepts are acquired, for example, according to the following process. Each principal component of the PCA described above includes several attributes (e.g., two attributes: "age" and "smile"). Next, data that responds to each principal component is extracted from the dataset (a group of samples including age and smiley faces is extracted from face images). To acquire linguistic expressions, attributes are searched for using the CLIP image encoder, CLIP text encoder, and CLIP vocabulary of the extracted data. Specifically, a search is performed within CLIP's vocabulary to find words similar to the expressions of the extracted sample group. Here, smiley faces and ages are acquired as language.
[0015] The labeling unit 3 labels each piece of data (e.g., image data) constituting the dataset with the attribute extracted by the attribute extraction unit 2. The labeling method is not limited to a specific method, but in this embodiment, CLIP is used as the labeling method. CLIP is a method in which the attribute extracted by the attribute extraction unit 2 is used as a label to determine the similarity with each piece of image data, and a label is assigned to image data that has a similarity equal to or greater than a certain level. This allows each piece of image data to be classified as to whether or not it has the attribute extracted by the attribute extraction unit 2. This makes it possible to classify each piece of image data as to whether or not it has an attribute, without the information processing device 1 having to perform learning.
[0016] The bias detection unit 4 detects bias for each attribute based on the labeling results by the labeling unit 3. The attributes extracted by the attribute extraction unit 2 may include attributes that are not biased within the dataset. Therefore, it is necessary to identify attributes that have bias. The bias detection unit 4, for example, calculates the degree of imbalance for each attribute, and detects attributes with a degree of imbalance equal to or greater than a certain level as attributes that have bias. Examples of distances and amounts of information that can be used when calculating the degree of imbalance include "Euclidean distance," "Chebyshev distance," "Kullback Leibler divergence," "Hellinger distance," "Total variation distance," and "Chi-square divergence."
[0017] The data set adjustment unit 5 adjusts the data set so as to improve at least one bias detected by the bias detection unit 4. Note that a specific example of adjusting the data set will be described later.
[0018] The output unit 6 is a general term for devices that output various types of information. The output unit 6 includes, for example, a display unit 6A, an audio output unit 6B, and a communication unit 6C. The display unit 6A is an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display. The audio output unit 6B is a device that reproduces sound, such as a speaker. The communication unit 6C is a communication device that transmits and receives data and commands to and from devices that can communicate with the information processing device 1 via a network such as the Internet or a wireless LAN (Local Area Network).
[0019] The output unit 6 outputs, for example, display information that visualizes the results of labeling by the labeling unit 3. The output unit 6 also outputs, for example, display information that visualizes information based on the bias for each attribute detected by the bias detection unit 4. The output unit 6 also outputs, for example, display information that visualizes the results of bias improvement by the dataset adjustment unit 5. Specific output examples will be described later. The output display information may be displayed on the display unit 6A of the information processing device 1, or on a display unit of another device (for example, a user terminal 200 described later).
[0020] [Processing Flow] (Overview) An overview of the processing performed by the information processing device 1 will be described with reference to FIG. 2 . Attributes are automatically extracted from an input dataset by the attribute extraction unit 2. Then, the labeling unit 3 labels each image data with an attribute. In the illustrated example, attribute A is gender, attribute B is race, and attribute C is age. Each attribute includes a set of image data to which the attribute is assigned. Image data may be labeled with multiple attributes. The results of labeling by the labeling unit 3 are displayed in a table format (e.g., a bar graph format) as shown on the left side of FIG. 2 . For example, the number of image data annotated as group 1 to which the attribute "gender" is assigned is displayed side by side with the number of image data annotated as group 2 to which the attribute "gender" is assigned. Furthermore, the number of pieces of image data to which "race" has been assigned as an attribute among the image data annotated as group 1 and the number of pieces of image data to which "race" has been assigned as an attribute among the image data annotated as group 2 are displayed side by side. Furthermore, the number of pieces of image data to which "age" has been assigned as an attribute among the image data annotated as group 1 and the number of pieces of image data to which "age" has been assigned as an attribute among the image data annotated as group 2 are displayed side by side.
[0021] Based on the labeling results by the labeling unit 3, the bias detection unit 4 detects bias for each attribute. For example, assume that bias is detected for all attributes (gender, race, and age in this example). The dataset adjustment unit 5 adjusts the contents of the dataset so as to improve the bias for each attribute for which bias has been detected.
[0022] For example, the data set adjustment unit 5 performs processing to equalize the number of image data annotated as group 1 that has been assigned the attribute "gender" and the number of image data annotated as group 2 that has been assigned the attribute "gender."
[0023] (Processing Flow) Fig. 3 is a flowchart showing the processing flow performed by the information processing device 1. Note that the processing flow will be described below assuming an AI model for a gender classification task.
[0024] When the process begins, in step S101, the attribute extraction unit 2 extracts interpretable attributes from the dataset. Specifically, a latent space representation is obtained for the images of the dataset using CLIP's image encoder. The average of this latent space representation is the common theme of the dataset. For example, in the case of a face image dataset, the theme represents "face image." The average is subtracted from each latent space representation. Next, PCA is performed on the obtained latent space representation. It is assumed that the first principal component obtained here will significantly change the linguistic representation learned by CLIP. However, at this stage, multiple attributes may be changed. It is assumed that each attribute is an interpretable attribute. Next, linguistic representations are automatically assigned to the extracted attributes. Specifically, linguistic representations are assigned to the attributes using CLIP's text encoder. The assignment of linguistic representations may be performed automatically, or manual filtering may be performed to address the inclusion of offensive words in CLIP text.
[0025] In this example, each image constituting the dataset is pre-annotated with a label of "male" or "female." The description will be made assuming that the attribute extraction unit 2 extracts "whether or not a smile is present," "race," and "age" as examples of attributes. The attribute extraction unit 2 extracts not only biased attributes but also unbiased attributes. The attributes extracted by the attribute extraction unit 2 correspond to the perturbation ΔW assigned to the latent variable W in CLIP2StyleGAN. Then, the process proceeds to step S102.
[0026] In step S102, the labeling unit 3 labels each image data with the attribute extracted by the attribute extraction unit 2. Specifically, for the attribute extracted by the attribute extraction unit 2, each image data constituting the dataset is classified into two categories, namely, the presence or absence of the attribute. This makes it possible to calculate the frequency of each attribute in the dataset. Note that the frequency calculated can be presented to the user in a graph or the like, thereby visualizing the quality of the dataset before bias improvement. Then, the process proceeds to step S103.
[0027] In step S103, the bias detection unit 4 detects the presence or absence of attribute bias, for example, by calculating the degree of imbalance for each attribute. In step S101, interpretable attributes are automatically extracted. However, it is possible that not all of the extracted attributes are biased in the dataset. Therefore, it is necessary to extract attributes that are likely to be biased from among these. Dataset bias occurs when attributes are imbalanced within the dataset or when unrelated attributes are correlated due to dataset sampling. The following is an example of a correlation between unrelated attributes: Consider a dataset consisting of "dogs" and "cats." Suppose the dataset contains many images of "cats" taken in bedrooms. In this case, an inappropriate correlation occurs between "beds" and "cats" in the bedroom, leading to the determination that an image containing a "bed" and a "dog" is a "cat." Therefore, attributes that are likely to be biased are attributes that account for the majority of the explanatory variables in the dataset despite being unrelated to them, or when there is a bias in the distribution between the attributes and explanatory variables in the dataset. Ultimately, the bias attribute can be identified by calculating the degree of imbalance for each attribute from the graph and attribute frequency described in step S102. The above-mentioned "Euclidean distance" or the like is used to calculate the degree of imbalance. Then, the process proceeds to step S104.
[0028] In step S104, the dataset adjustment unit 5 improves the bias by, for example, expanding the image data that constitutes the dataset. Note that improving the bias means, for example, making the degree of imbalance zero or making the degree of imbalance equal to or less than a certain level. Specific examples of processing by the dataset adjustment unit 5 will be described later.
[0029] The bias-improved dataset may be stored in a storage device (not shown) included in the information processing device 1. The stored bias-improved dataset may be accessible and downloadable by other devices. Furthermore, the bias-improved dataset may be transmitted from the information processing device 1 to other devices via a communication unit 6C included in the output unit 6.
[0030] [Display Examples of Labeling Results] (First Display Example) Next, a description will be given of a display example of display information that visualizes the labeling results by the labeling unit 3. The display example shown below may be displayed on the display unit 6A, or may be displayed on a display unit included in a device that can communicate with the information processing device 1.
[0031] 4 is a diagram illustrating a first display example of the labeling results by the labeling unit 3. The first display example is an example in which the labeling results by the labeling unit 3 (in other words, the contents of the dataset before bias improvement) are visualized in a graph format, more specifically, in a bar graph format.
[0032] As shown in FIG. 4, the attributes extracted by the attribute extraction unit 2 (in this example, "whether or not a person is smiling," "race," and "age") are displayed along the horizontal axis. The vertical axis (height of the bar) indicates the number of images. Note that in the example of FIG. 4, 5,000 images are shown as the number of images, but this is just an example. To allow the user to recognize the number of images in more detail, lines may be displayed every 1,000 images, for example.
[0033] The number of image data annotated with "male" that constitutes the data set and assigned the attribute "whether or not the subject is smiling" is displayed side by side, along with the number of image data annotated with "female" that constitutes the data set and assigned the attribute "whether or not the subject is smiling." In this example, two bar graphs showing the respective numbers are displayed on the left side.
[0034] The number of image data annotated with "male" that has the attribute "race" added to it, and the number of image data annotated with "female" that has the attribute "race" added to it, are displayed side by side. In this example, two bar graphs showing the respective numbers are displayed near the center.
[0035] The number of image data annotated with "male" that has "age" assigned as an attribute, and the number of image data annotated with "age" that has "race" assigned as an attribute are displayed side by side. In this example, two bar graphs showing the respective numbers are displayed near the right side.
[0036] By displaying the data in bar graph format, the user can intuitively recognize whether there is a difference in the number of images labeled with a certain attribute between men and women, i.e., whether there is a bias in the attribute. Furthermore, since the user can intuitively recognize the difference in the number of images, the user can also intuitively recognize the degree of bias. In the example shown in Figure 4, there is a large difference in the number of images labeled with the attribute "smiling or not" between "male" and "female," so the user can recognize that the degree of bias in the attribute "smiling or not" is large.
[0037] (Second Display Example) Fig. 5 is a diagram for explaining a second display example. As shown in Fig. 5, a bar graph showing the number of images for each attribute is also displayed in this example. In this example, in addition to the bar graph, representative sample images showing the attributes are displayed.
[0038] As shown in FIG. 5 , a sample image IM corresponding to the attribute "presence or absence of a smile" is displayed, for example, on the left side of the bar graph. Specifically, sample images IM including a sample image IMA corresponding to a male smile and a sample image IMB corresponding to a female smile are displayed on the left side of the bar graph. Although not shown, sample images corresponding to the attribute "race" (e.g., images of men and women of various races) or sample images corresponding to the attribute "age" (e.g., images of men and women of various ages) may also be displayed. Furthermore, an attribute may be designated by a user operation, and sample images corresponding to the designated attribute may be displayed.
[0039] In the case of data that requires specialized knowledge, such as medical image classification, simply displaying attributes as text can make it difficult for users to recognize their attributes. Visualizing attributes using sample images, as in this example, makes it easier for users to recognize their attributes.
[0040] [Display Example of Bias Detection Results] Next, a description will be given of a display example of display information that visualizes the bias detection results by the bias detection unit 4. The bias detection unit 4 quantifies the degree of imbalance for each attribute and displays it as text information. A larger degree of imbalance indicates a larger degree of bias (for example, a difference in the number of images). Note that the display example shown below may be displayed on the display unit 6A, or may be displayed on a display unit included in a device that can communicate with the information processing device 1.
[0041] For example, as shown in FIG. 6A , the imbalance degree corresponding to the attribute “smiling” is displayed as “40,” the imbalance degree corresponding to the attribute “race” is displayed as “20,” and the imbalance degree corresponding to the attribute “age” is displayed as “10.” As shown in FIG. 6B , the imbalance degree of the entire data set may be displayed. The imbalance degree of each attribute may be averaged to represent the imbalance degree of the entire data set. Note that the average may take into account only the imbalance degree of a user-defined attribute. For example, if the task of the AI model is to automatically identify attendance at university lectures, it is reasonable for age to be biased, and therefore the imbalance degree of the attribute “age” may be excluded from the calculation of the average imbalance degree. As shown in FIG. 6C , both the imbalance degree of each attribute and the imbalance degree of the entire data set may be displayed.
[0042] [Processing of Dataset Adjustment Unit] Next, a specific example of the processing of the dataset adjustment unit will be described. As shown in FIGS. 4 and 5 , if learning is performed using a dataset in which bias exists in an attribute and an AI model generated by learning is used, problems will occur in the judgment results of the AI model. For example, assume that, among the three attributes "whether or not a person smiles," "race," and "age," the degree of imbalance in the attribute "whether or not a person smiles" is greater than a certain level, and a bias exists in the attribute "whether or not a person smiles." In this case, an AI model for a gender classification task obtained using a dataset in which the bias is not corrected may classify smiling women as men. Therefore, the dataset adjustment unit 5 corrects the bias in the dataset.
[0043] For example, the dataset adjustment unit 5 improves the bias by expanding the dataset. As shown in FIG. 7 , for example, the dataset adjustment unit 5 expands the dataset by adding new image data to the dataset so that the number of images of men and women corresponding to each attribute is equal. Note that the numbers do not have to be strictly equal (the same number). Adjusting the number of image data so that the difference in the number of images is small enough to ignore the bias is also considered to be equal.
[0044] New image data is collected manually. For example, a person searches an image database, obtains image data of a smiling woman, and adds the obtained image data to a data set. Although this is a time-consuming process, it is easy to collect image data because the direction (characteristics) of the target image data is fixed, such as "a smiling woman." Of course, image data may also be collected automatically by a computer instead of manually.
[0045] The new image data may be generated image data. For example, it may be image data generated by a trained image generation model (image generator). Furthermore, the new image data may be image data based on a virtual viewpoint generated based on image data from a real viewpoint (real image data). Furthermore, the new image data may be a combination of the above-mentioned image data.
[0046] The dataset adjustment unit 5 may improve the bias by deleting some of the image data constituting the dataset to equalize the number of images. For example, the dataset adjustment unit 5 may delete some of the image data of "males" labeled with the attribute "whether or not smiling" so that the number of image data of "males" labeled with the attribute "whether or not smiling" becomes equal to the number of image data of "females" labeled with the attribute "whether or not smiling".
[0047] Display information visualizing the contents of the dataset after bias improvement may be displayed on the display unit 6A or a display unit of a device capable of communicating with the information processing device 1. For example, as shown in Fig. 8, information indicating that the number of image data of men and women labeled with each attribute has become equal, i.e., that the bias has been improved, may be displayed in the form of a bar graph. Of course, instead of the bar graph format, text information indicating that the degree of imbalance has become 0 may also be displayed.
[0048] [Effects Obtained by the Present Embodiment] The information processing device according to the present embodiment can obtain, for example, the following effects. The information processing device according to the present embodiment makes it possible to improve bias contained in a dataset. The bias of a dataset can be improved regardless of the type of image data constituting the dataset. There is no need to perform bias improvement processing multiple times, and the bias of a dataset can be improved efficiently. Bias can be explained in natural language, making it easy for users to interpret. Unlike conventional methods, bias can be improved by improving an AI model, but the bias of the dataset itself can be improved. Improving the dataset itself is expected to improve prediction accuracy more than improving an AI model. Bias of a dataset can be improved without being bound by subsequent tasks (classification, object recognition, segmentation, etc.). Attributes that contain bias can be automatically detected, and the results can be presented to the user. There is no need for the user to predefine attributes that are thought to contain bias. Furthermore, attributes can be extracted from a dataset even if the attributes of the dataset are not known.
[0049] Second Embodiment Next, a second embodiment will be described. In the description of the second embodiment, the same or similar components as those in the above description will be denoted by the same reference numerals, and duplicated descriptions will be omitted as appropriate. Furthermore, unless otherwise specified, the matters described in the first embodiment can be applied to the second embodiment.
[0050] Fig. 9 is a flowchart showing the flow of processing executed in the second embodiment. The processing shown in Fig. 9 is basically the same as the flow of processing executed by the information processing device 1 described in the first embodiment (see Fig. 3).
[0051] In the first embodiment, in the process of step S101, attributes are automatically extracted from a dataset using CLIP2StyleGAN. This embodiment differs from the first embodiment in that the user can specify attributes. This makes it possible to handle attributes that cannot be automatically extracted by CLIP2StyleGAN. For example, it becomes possible to label attributes that cannot be automatically extracted by CLIP2StyleGAN and to detect bias in those attributes.
[0052] The processing from step S102 onwards is the same as that in the first embodiment. For example, the attributes labeled by the labeling unit 3 include the attributes specified by the user in step S101. Furthermore, the bias detection unit 4 detects bias for at least the attributes specified by the user in step S101 based on the labeling results of the labeling unit 3. The dataset adjustment unit 5 adjusts the dataset so as to improve the bias for the attributes specified by the user that has been detected by the bias detection unit 4.
[0053] <Third Embodiment> Next, a third embodiment will be described. In the description of the third embodiment, the same or similar components as those in the above description will be denoted by the same reference numerals, and duplicated descriptions will be omitted as appropriate. Furthermore, unless otherwise specified, the matters described in the first and second embodiments can be applied to the third embodiment.
[0054] 10 is a flowchart showing the flow of processing executed in the third embodiment. In this embodiment, an AI model trained using a dataset is used to consistently perform attribute extraction by the attribute extraction unit 2, attribute labeling by the labeling unit 3, bias detection for each attribute by the bias detection unit 4, and dataset adjustment by the dataset adjustment unit 5.
[0055] In step S201, learning is performed using a dataset to generate a StyleGAN2 model, which is an example of an AI model. Then, the process proceeds to step S202.
[0056] In step S202, the dataset is transformed into the latent space of StyleGAN2. After transformation, the matrices shown in FIG. 11 are obtained for both males and females. In the matrix shown in FIG. 11, the rows correspond to the number of training data (d), and the columns correspond to the number of attributes (k) in the latent space. Then, the process proceeds to step S203.
[0057] In step S203, the distribution of tasks for each attribute is calculated to identify significant attributes. FIG. 12 shows an example of the task distribution results. The horizontal axis of the distribution (graph) in FIG. 12 represents the attribute value (k=4989 in this example), and the vertical axis represents the distribution of the number of data points. Furthermore, dots represent the distribution corresponding to "male," and diagonal lines represent the distribution corresponding to "female." Then, processing proceeds to step S204.
[0058] In step S204, the distance between the task distributions is calculated using Wasserstein-1 distance or similar to identify attributes that are important to the task. Important means that the distribution distance is large. Here, meaningful dimensions are identified from the StyleGAN2 latent space by calculating the distance between distributions. As shown in FIG. 13A, if there is little overlap between the distributions of males and females for a certain attribute, attributes such as age and facial expression can be manipulated by changing the latent space. On the other hand, as shown in FIG. 13B, if there is a large overlap in the distributions, even if the latent space is manipulated, it is expected that the image will be close to the original (the dimension does not retain a significant attribute). Processing then proceeds to step S205.
[0059] In step S205, the numerical value of the target attribute is changed and an image is generated using the StyleGAN2 generator. As shown in Fig. 14, an image (center image) is generated by changing the numerical value of the attribute (k = 4989) of the source image (left image), and an image (right image) is generated by changing the numerical value of the attribute (k = 5043) of the source image (left image). Then, the process proceeds to step S206.
[0060] In step S206, the semantic concept of the attribute is identified by CLIP using a pair of the source image (original image) and the image generated in step S205. For example, as shown in FIG. 15, the semantic concept of the attribute (k=4989) is identified as "aging of the mouth" using a pair of the source image and an image in which the value of the attribute (k=4989) has been changed. The attribute with the identified semantic concept is labeled for each image data. Note that, for example, two methods can be considered for identifying the semantic concept of the attribute. The first method is a method of identifying the semantic concept of the attribute by visual judgment. Most of the latent space of StyleGAN is considered to be meaningless dimensions, and the number of dimensions that can be manipulated by attributes is thought to be small, and overlaps are also observed. Therefore, it is thought that labeling by visual judgment does not require a large amount of labor. The second method is a method of identifying the semantic concept of the attribute using CLIP2StyleGAN. This method makes it possible to automatically identify the semantic concept of the attribute. Then, the process proceeds to step S207.
[0061] In step S207, for undesirable attributes that have a large impact on the task (for example, a gender classification task), the number of image data is adjusted by generating images using the StyleGAN2 generator, thereby improving bias.
[0062] For example, if the task is a gender classification task, it is not applicable to classify men and women based on the presence or absence of the attribute "glasses." In this case, image data is generated and adjusted so that the number of images labeled with the attribute "glasses" is equal for men and women. Figure 16 shows an example of the distribution after image data adjustment.
[0063] In the above-described processing, for example, the processing performed in step S202 corresponds to processing by the attribute extraction unit 2, the processing performed in steps S205 and S206 corresponds to processing by the labeling unit 3 and bias detection unit 4, and the processing performed in step S207 corresponds to processing by the dataset adjustment unit 5.
[0064] <Fourth Embodiment> Next, a fourth embodiment will be described. In the description of the fourth embodiment, the same or similar components as those in the above description will be denoted by the same reference numerals, and duplicate descriptions will be omitted as appropriate. Furthermore, unless otherwise specified, the matters described in the first to third embodiments can be applied to the fourth embodiment.
[0065] 17 is a block diagram showing an example of the configuration of an information processing device (information processing device 1A) according to the third embodiment. The information processing device 1A differs in configuration from the information processing device 1 in that it includes a learning unit 7.
[0066] The learning unit 7 performs learning using the dataset whose bias has been improved by the dataset adjustment unit 5. An AI model corresponding to a predetermined task is generated through this learning. Because the bias of the dataset used for learning has been improved, it is possible to prevent the generated AI model from mistakenly recognizing a specific race as an animal or giving a negative evaluation to a specific gender.
[0067] The learning unit 7 does not necessarily have to be included in the information processing device 1A, but may be included in a user terminal that can communicate with the information processing device 1A.
[0068] Fifth Embodiment Next, a fifth embodiment will be described. In the description of the fifth embodiment, the same or similar components as those in the above description will be denoted by the same reference numerals, and duplicate descriptions will be omitted as appropriate. Furthermore, unless otherwise specified, the matters described in the first to fourth embodiments can be applied to the fifth embodiment.
[0069] In the first embodiment described above, the data set adjustment unit 5 reduces the bias for all attributes in which bias exists (see FIG. 8 ). However, it is also possible to reduce the bias for only some attributes in which bias exists, rather than for all attributes in which bias exists. An example of reducing bias according to this embodiment will be described below.
[0070] (First Improvement Example) The first improvement example is an example in which, when there is a bias of a certain level or more for a certain attribute, the bias is improved. For example, when the degree of imbalance detected by the bias detection unit 4 is a certain level or more, the dataset adjustment unit 5 determines that there is a bias of a certain level or more, and performs processing to improve the bias. When the degree of imbalance is equal to or less than a predetermined threshold, it is possible not to improve the bias, as it is determined that the bias does not have a significant impact on the generation of the AI model. This improves processing efficiency.
[0071] (Second Improvement Example) The second improvement example is an example in which, when a bias exists in an attribute specified by a user, the data set adjustment unit 5 performs processing to improve at least the bias of the attribute, as in the second embodiment.
[0072] (Third Improvement Example) In the third improvement example, the user specifies an attribute for which bias is to be reduced, and the dataset adjustment unit 5 performs processing to reduce the bias of the specified attribute. As described in the first embodiment, the labeling results by the labeling unit 3 and the detection results by the bias detection unit 4 are presented to the user by display or the like. The user specifies the attribute for which bias is to be reduced based on the presented content.
[0073] (Fourth Improvement Example) This example is an example in which the attribute for which bias should be improved is changed depending on the task content of the AI model. As described above, in order to suppress erroneous recognition by the AI model, it is basically preferable to improve the bias for all attributes. However, depending on the task content of the AI model, it may be preferable not to improve the bias for some attributes.
[0074] For example, consider a case where the task of an AI model is "person recognition in a dark place," which can be applied to security cameras or cameras that capture images of escalators inside buildings, and the attribute "bright background" is extracted or user-defined. Assume that the attribute "bright background" has a bias.
[0075] In the above example, when processing similar to that described in the first embodiment is performed, the number of image data to which the attribute "bright background" is assigned and the number of image data to which the attribute "bright background" is not assigned are equalized. However, when the task content of the AI model is "person recognition in a dark place," learning with a dataset containing a large number of image data to which the attribute "bright background" is not assigned, in other words, learning with a dataset with a bias, is more likely to improve the accuracy of the generated AI model. Therefore, for example, if the number of image data to which the attribute "bright background" is assigned is smaller than the number of image data to which the attribute is not assigned, i.e., image data with a dark background, the dataset adjustment unit 5 does not perform processing to improve the bias. If the number of image data to which the attribute "bright background" is assigned is greater than the number of image data to which the attribute is not assigned, i.e., image data with a dark background, the dataset adjustment unit 5 may create a bias by, for example, expanding the image data so that the number of image data to which the attribute "bright background" is not assigned increases. For example, if a human recognition dataset exists in advance and contains a large number of bright backgrounds and few dark backgrounds, this may be useful because it eliminates the need to build a dataset from scratch for low-light recognition.
[0076] Furthermore, if the task content of the AI model is, for example, a person recognition model for an entrance where external light always enters or where illumination light is irradiated, the accuracy of the AI model will improve if it is trained with a dataset containing a large number of image data to which the attribute "bright background" is assigned. Therefore, if the number of image data to which the attribute "bright background" is assigned is greater than the number of image data to which the attribute is not assigned, i.e., image data with a dark background, the dataset adjustment unit 5 will not perform processing to improve the bias. If the number of image data to which the attribute "bright background" is assigned is less than the number of image data to which the attribute is not assigned, i.e., image data with a dark background, the dataset adjustment unit 5 may create a bias by, for example, expanding the image data so that the number of image data to which the attribute "bright background" is assigned is greater.
[0077] The user may define how to adjust the bias of each attribute depending on the task content of the AI model, or a predetermined discrimination unit may automatically determine the bias depending on the task content.
[0078] In any of the above improvement examples, it is preferable to improve bias in ethical attributes. Ethical attributes refer to attributes from which bias should be eliminated from the perspective of fairness, and are also called sensitive attributes. Ethical attributes include gender, race, age, disability, skin color, nationality, religion or belief, medical history, etc. However, even these attributes may not be ethical attributes depending on the task content of the AI model.
[0079] Next, examples of the present technology will be described, but the present technology is not limited to the contents of the following examples.
[0080] 18 shows an example of a schematic configuration of an information processing system 1000 that constitutes an imaging device system according to an embodiment. That is, the information processing system 1000 can be an example of a system to which the present technology is applied.
[0081] 18, an information processing system 1000 according to the embodiment includes at least a cloud server 100, a user terminal 200, cameras 300 as a plurality of imaging devices, a fog server 400, and a management server 500. Here, at least the cloud server 100, the user terminal 200, the fog server 400, and the management server 500 are configured to be able to communicate with each other via a network 600 such as the Internet.
[0082] The cloud server 100, the user terminal 200, the fog server 400, and the management server 500 are all configured as information processing devices equipped with a microcomputer having a CPU, a ROM (Read Only Memory), and a RAM (Random Access Memory).
[0083] The camera 300 as an imaging device includes an image sensor, such as a charge-coupled device (CCD) image sensor or a complementary metal oxide semiconductor (CMOS) image sensor. These image sensors constitute an imaging unit (see reference numeral 410 in FIG. 24 ). The camera 300 captures an image of a subject and obtains image information (captured image information) as digital data. The camera 300 also has a function for performing AI-based processing on the captured image. Examples of this processing include image recognition processing and image detection processing. In the following description, various types of image processing, such as image recognition processing and image detection processing, will be simply referred to as "image processing." For example, various types of image processing using AI or an AI model will be referred to as "AI image processing."
[0084] The multiple cameras 300 are configured to be able to communicate data with the fog server 400. For example, various data such as processing result information indicating the results of image processing using AI is transmitted from the cameras 300 to the fog server 400. In addition, the cameras 300 receive various data from the fog server 400.
[0085] Here, the information processing system 1000 is expected to be used, for example, in the following manner: First, the fog server 400 or the cloud server 100 generates analytical information of the subject based on the processing result information obtained by image processing of the multiple cameras 300. This generated analytical information can be viewed by the user via the user terminal 200.
[0086] In this case, the multiple cameras 300 are used as surveillance cameras. For example, they can be used as surveillance cameras for monitoring indoor spaces such as stores, offices, and homes, or as surveillance cameras for monitoring outdoor spaces such as parking lots and city streets. Examples of surveillance cameras for monitoring outdoor spaces include traffic surveillance cameras for monitoring traffic conditions. They can also be used as surveillance cameras for monitoring production lines for FA (Factory Automation), IA (Industrial Automation), and the like. They can also be used as surveillance cameras for monitoring the interior or exterior of automobiles, trains, and the like.
[0087] Furthermore, when used as a surveillance camera in a store, multiple cameras 300 can be placed at predetermined locations within the store. Using multiple cameras 300 allows the user to check the demographics of customers (gender, age, etc.) and their behavior (traffic flow) within the store. In this case, information on the demographics of customers, information on the traffic flow within the store, and information on the congestion status at the cash registers (for example, waiting times at the cash registers) can be generated as analysis information.
[0088] Furthermore, for traffic monitoring camera applications, multiple cameras 300 can be placed at various locations near roads. Using multiple cameras 300 allows the user to recognize information such as the license plate number (vehicle number), vehicle color, and vehicle model of passing vehicles. In this case, information such as the license plate number, vehicle color, and vehicle model can be generated.
[0089] Furthermore, when used as a surveillance camera in a parking lot, the camera 300 can be placed in a position where it can monitor parked vehicles. The camera 300 can be used to monitor, for example, whether there is a suspicious person behaving suspiciously around the vehicle. Furthermore, if a suspicious person is found, a notification device can be provided that notifies the driver of the presence of the suspicious person and their attributes (gender or age group), etc.
[0090] Furthermore, if the camera is used as a surveillance camera to monitor available spaces in towns or parking lots, it can notify users of the locations of spaces where they can park their cars.
[0091] For example, in the above-mentioned store monitoring application, the fog server 400 is placed in the store to be monitored together with the multiple cameras 300. In other words, a fog server 400 is placed for each monitored object. When a fog server 400 is placed for each monitored object such as a store, the cloud server 100 does not need to directly receive data transmitted from the multiple cameras 300 at the monitored object. As a result, the processing load on the cloud server 100 can be reduced.
[0092] In addition, when there are multiple stores to be monitored and all of the stores belong to the same chain, it is preferable to place a fog server 400 for each of the stores, rather than for each individual store. In other words, it is not limited to placing one fog server 400 for each monitored object, but it is possible to place one fog server 400 for each of multiple monitored objects.
[0093] Furthermore, if the cloud server 100 or the plurality of cameras 300 has processing capabilities, the cloud server 100 or the plurality of cameras 300 can have the functions of the fog server 400. As a result, in the information processing system 1000, the fog server 400 can be omitted, and the plurality of cameras 300 can be directly connected to the network 600, allowing the cloud server 100 to directly receive data transmitted from the plurality of cameras 300.
[0094] The various devices described above are broadly divided into cloud-side information processing devices and edge-side information processing devices. Cloud-side information processing devices include the cloud server 100 and the management server 500. Cloud-side information processing devices are a group of devices that provide services that are expected to be used by multiple users. Edge-side information processing devices include the camera 300 and the fog server 400. Edge-side information processing devices are a group of devices that are prepared by users who use cloud services and placed in the environment.
[0095] However, both the cloud-side information processing device and the edge-side information processing device may be arranged in an environment prepared by the same user. The fog server 400 may be an on-premise server.
[0096] (2) Registration of AI Model and AI Application As described above, in the information processing system 1000, AI image processing is performed in the camera 300, which is an edge-side information processing device. Then, in the cloud server 100, which is a cloud-side information processing device, advanced application functions are realized using result information of the AI image processing on the edge side. The result information of the AI image processing is, for example, result information of image recognition processing using AI.
[0097] Here, various methods for registering application functions in the cloud server 100, which is a cloud-side information processing device, or the cloud server 100 including the fog server 400, are as follows. FIG. 19 shows an example configuration of each device in the information processing system 1000 that registers or downloads AI models and AI applications via a marketplace function provided in the cloud-side information processing device. Note that while the fog server 400 is not shown in FIG. 19, the fog server 400 may be provided. In this case, the fog server 400 may take on part of the edge-side functions.
[0098] The cloud server 100 and management server 500 described above are information processing devices that constitute the cloud-side environment. The camera 300 is an information processing device that constitutes the edge-side environment. The camera 300 can be configured as a device equipped with a control unit that performs overall control of the camera 300. The camera 300 can also be configured as a device equipped with an image sensor IS (see FIG. 22 ) that has an arithmetic processing unit that performs various processes, including AI image processing, on captured images. That is, the camera 300, which is an edge-side information processing device, may be equipped with an image sensor IS that is another edge-side information processing device.
[0099] Furthermore, the user terminals 200 used by users who use various services provided by the cloud-side information processing device include an application developer terminal 200A, an application user terminal 200B, an AI model developer terminal 200C, etc. The application developer terminal 200A is used by a user who develops an application used in AI image processing. The application user terminal 200B is used by a user who uses an application. The AI model developer terminal 200C is used by a user who develops an AI model used in AI image processing. Note that the application developer terminal 200A may also be used by a user who develops an application that does not use AI image processing.
[0100] The cloud-side information processing device has prepared therein a training dataset for AI learning. A user developing an AI model communicates with the cloud-side information processing device using the AI model developer terminal 200C and downloads these training datasets (the datasets described in the above-described embodiments). At this time, the training datasets may be provided for a fee. For example, the AI model developer may purchase the training dataset in a state where he or she can purchase various functions and materials registered in a marketplace (electronic marketplace) provided as a function on the cloud side by registering personal information in the marketplace.
[0101] After developing an AI model using the learning dataset, the AI model developer registers the developed AI model in the marketplace using the AI model developer terminal 200C. As a result, an incentive may be paid to the AI model developer when the AI model is downloaded.
[0102] Furthermore, a user who develops an application downloads an AI model from the marketplace using the application developer terminal 200A and develops an application that uses this AI model (hereinafter simply referred to as an "AI application"). At this time, as described above, an incentive may be paid to the AI model developer.
[0103] A user who develops an application registers the developed AI application in the marketplace using the application developer terminal 200A. In this way, an incentive may be paid to the user who developed the AI application when the AI application is downloaded.
[0104] A user of an AI application uses the application user terminal 200B to deploy the AI application and AI model from the marketplace to the camera 300, which serves as an edge-side information processing device that the user manages. At this time, an incentive may be paid to the AI model developer. This enables the camera 300 to perform AI image processing using the AI application and AI model. Specifically, in addition to capturing images, the camera 300 can also detect customers and vehicles through AI image processing.
[0105] Here, deployment of an AI application and an AI model refers to installing the AI application or AI model in a target (device) as an execution subject so that the target (device) as an execution subject can use the AI application or AI model. Furthermore, deployment includes installing the AI application or AI model in a target as an execution subject so that at least a part of the program as the AI application can be executed.
[0106] Furthermore, the camera 300 may be configured to extract attribute information of customers by AI image processing from images captured by the camera 300. The attribute information is transmitted from the camera 300 to an information processing device on the cloud side via the network 600.
[0107] Cloud applications are deployed on the cloud-side information processing device. Each user can use the cloud applications via network 600. The cloud applications include an application that analyzes the movement of customers using their attribute information and captured images. Such cloud applications are uploaded by application developers and users.
[0108] A user of the application uses the cloud application for flow analysis using the application user terminal 200B. This allows the user to analyze the flow of customers visiting their own store and view the analysis results. Viewing the analysis results refers to viewing the flow of customers graphically displayed on a store map, for example. The results of the flow analysis may also be displayed in the form of a heat map, showing the density of customers, etc., for viewing the analysis results. The information may also be displayed sorted by attribute information of the customers.
[0109] In the cloud-side marketplace, AI models optimized for each user may be registered. For example, images captured by a camera 300 installed in a store managed by a user are uploaded and stored in an information processing device on the cloud side as appropriate.
[0110] In the information processing device on the cloud side, a re-learning process of the AI model is performed each time a certain number of uploaded captured images are accumulated, and a process of updating the AI model and re-registering it in the marketplace is executed. Note that the re-learning process of the AI model may be made selectable as an option by the user on the marketplace, for example.
[0111] For example, an AI model retrained using dark images from a camera 300 installed inside a store is deployed to the camera 300. This can improve the recognition rate, etc. of image processing for images captured in dark places. Also, an AI model retrained using bright images from a camera 300 installed outside the store is deployed to the camera 300. This can improve the recognition rate, etc. of image processing for images captured in bright places. In other words, a user using an application can always obtain optimized processing result information by redeploying an updated AI model to the camera 300 again. The retraining process of the AI model will be explained later.
[0112] Furthermore, if personal information is included in information (e.g., information about captured images) uploaded from the camera 300 to the cloud-side information processing device, the data may be uploaded with the privacy information deleted from the viewpoint of privacy protection. The data with the privacy information deleted may be made available to users developing AI models and applications.
[0113] 20 and 21 are flowcharts showing an example of the flow of the above-mentioned processing. Note that the information processing device on the cloud side corresponds to the cloud server 100, management server 500, etc. shown in FIG.
[0114] The AI model developer browses a list of data sets registered in the marketplace using the AI model developer terminal 200C, which has a display unit consisting of an LCD, an organic EL panel, etc. When the AI model developer selects a desired data set, in response to this selection, the AI model developer terminal 200C transmits a download request for the selected data set to the cloud-side information processing device (step S21).
[0115] The cloud-side information processing device accepts the request (step S1), and then performs processing to transmit the requested data set to the AI model developer terminal 200C (step S2).
[0116] The AI model developer terminal 200C performs a process to receive the data set (step S22), which enables the AI model developer to develop an AI model using the data set.
[0117] After the AI model developer finishes developing the AI model, the AI model developer performs an operation to register the developed AI model in the marketplace. This operation is, for example, an operation to specify the name of the AI model, the address where the AI model is located, etc. As a result, the AI model developer terminal 200C transmits a request to register the AI model in the marketplace to the cloud-side information processing device (step S23).
[0118] The cloud-side information processing device receives the registration request (step S3). The cloud-side information processing device performs registration processing for the AI model (step S4). The cloud-side information processing device can, for example, display the AI model on a marketplace. This allows users other than the AI model developer to download the AI model from the marketplace.
[0119] For example, an application developer who wishes to develop an AI application uses the application developer terminal 200A to browse a list of AI models registered in the marketplace. In response to an operation by the application developer, the application developer terminal 200A transmits a download request for the selected AI model to the cloud-side information processing device (step S31). The operation here is, for example, an operation to select one of the AI models on the marketplace.
[0120] The cloud-side information processing device accepts the request (step S5) and transmits the AI model to the application developer terminal 200A (step S6).
[0121] The application developer terminal 200A receives the AI model (step S32), which enables the application developer to develop an AI application that uses an AI model developed by another person.
[0122] After completing development of an AI application, the application developer performs an operation to register the AI application in the marketplace. This operation involves, for example, specifying the name of the AI application and the address where the AI model is located. As a result, the application developer terminal 200A transmits a registration request for the AI application to the cloud-side information processing device (step S33).
[0123] The cloud-side information processing device receives the registration request (step S7). The cloud-side information processing device registers the AI application (step S8). The cloud-side information processing device can, for example, display the AI application on a marketplace. This allows users other than the application developer to select and download the AI application on the marketplace.
[0124] 21, a user who intends to use an AI application selects a purpose on the application user terminal 200B (step S41). In the purpose selection, the selected purpose is transmitted to the information processing device on the cloud side.
[0125] In response to this, the cloud-side information processing device selects an AI application according to the purpose (step S9), and then performs preparation processing (deployment preparation processing) for deploying the AI application and AI model to each device (step S10).
[0126] In the deployment preparation process, an AI model is determined based on information about the device targeted for deployment of the AI model or AI application, such as information about the camera 300 or fog server 400, the performance required by the user, etc. In addition, in the deployment preparation process, it is determined on which device each SW (Software) component constituting the AI application for realizing the function desired by the user is to be executed, based on the performance information of each device and the user's request information.
[0127] Each SW component may be a container (described later) or a microservice. Note that the SW components can also be implemented using Web Assembly technology.
[0128] An AI application that counts the number of customers visiting a store by attribute such as gender or age would include a SW component that uses an AI model to detect people's faces from captured images, and would also include SW components that extract people's attribute information from the face detection results, SW components that aggregate the results, SW components that visualize the aggregated results, etc.
[0129] The deployment preparation process will be described again with some examples.
[0130] In the cloud-side information processing device, a process of deploying each SW component to each device is performed (step S11). In this process, the AI application and the AI model are transmitted to each device such as the camera 300.
[0131] In response to this, the camera 300 performs a deployment process of the AI application and the AI model (step S51). This enables AI image processing to be performed on the captured image captured by the camera 300. Although not shown in Fig. 21 , the fog server 400 also performs a deployment process of the AI application and the AI model as needed.
[0132] However, if all the processing is performed in the camera 300, the deployment processing to the fog server 400 is not performed.
[0133] The camera 300 performs an imaging operation to acquire an image (step S52). Then, the camera 300 performs AI image processing on the acquired image to obtain, for example, an image recognition result (step S53).
[0134] Furthermore, the camera 300 performs a process of transmitting the captured image and the result information of the AI image processing (step S54). The information transmitted in step S54 may include both the captured image and the result information of the AI image processing. Alternatively, only one of the information may be transmitted.
[0135] The cloud-side information processing device that receives this information performs analysis processing (step S12), such as analyzing the flow of customers visiting the store and analyzing vehicles for traffic monitoring.
[0136] The cloud-side information processing device performs processing to present the analysis results (step S13). This processing is realized, for example, by the user using the above-mentioned cloud application.
[0137] The application user terminal 200B receives the analysis result presentation process and performs a process of displaying the analysis result on a monitor or the like (step S42).
[0138] Once the processing up to this point is complete, the user of the AI application can obtain analysis results according to the purpose selected in step S41.
[0139] The information processing device on the cloud side may update the AI model after step S13. By updating and expanding the AI model, it is possible to obtain analysis results suited to the user's usage environment.
[0140] (3) Overview of System Functions In the embodiment, a service using the information processing system 1000 is assumed in which a user as a customer can select a function type for AI image processing of multiple cameras 300. For example, an image recognition function, an image detection function, or the like may be selected as the function type, or a more detailed type may be selected so as to perform an image recognition function, an image detection function, or the like for a specific subject.
[0141] For example, as a business model, a service provider sells cameras 300 and fog servers 400 with AI image recognition functions to users, and has the users install the cameras 300 and fog servers 400 in locations to be monitored.The service provider then develops a service that provides the above-mentioned analytical information to users.
[0142] In this case, the purpose for which each customer desires the system varies, such as store monitoring, traffic monitoring, etc. Therefore, the AI image processing function of the camera 300 can be selectively set so as to obtain analytical information corresponding to the purpose desired by the customer. In the embodiment, the management server 500 has a function for selectively setting the AI image processing function of the camera 300. Note that the cloud server 100 or the fog server 400 may also have the function of the management server 500.
[0143] FIG. 22 shows an example of a connection between a cloud server 100 and a management server 500, which are cloud-side information processing devices, and a camera 300, which is an edge-side information processing device.
[0144] As shown in FIG. 22, the cloud-side information processing device is equipped with a re-learning function, a device management function, and a marketplace function, which are functions available via a hub.
[0145] The Hub performs highly reliable communication with the edge-side information processing device while being protected by security, thereby providing various functions to the edge-side information processing device.
[0146] The re-learning function is a function that performs re-learning and provides a newly optimized AI model, thereby providing an appropriate AI model based on new learning materials.
[0147] The device management function is a function for managing the camera 300 and the like as an edge-side information processing device. For example, the device management function manages and monitors the AI model deployed in the camera 300, and provides functions such as problem detection and troubleshooting.
[0148] The device management function also manages information about the camera 300 and the fog server 400. The information about the camera 300 and the fog server 400 includes information about the chips used as the processing units, memory capacity, storage capacity, CPU and memory usage, etc. Furthermore, the information also includes information about software such as the operating system (OS) installed in each device. Furthermore, the device management function protects secure access by authenticated users.
[0149] The marketplace function provides functions for registering AI models developed by the above-mentioned AI model developers and AI applications developed by application developers, and for deploying these developments to authorized edge-side information processing devices, etc. The marketplace function also provides functions related to the payment of incentives according to the deployment of the developments.
[0150] The camera 300 as an edge-side information processing device includes an edge runtime, an AI application and AI model, and an image sensor IS.
[0151] The edge runtime functions as embedded software for managing applications deployed on the camera 300 and for communicating with an information processing device on the cloud side.
[0152] As described above, the AI model is an AI model registered in the marketplace of the cloud-side information processing device, which allows the camera 300 to use the captured image to obtain information on the results of AI image processing according to the purpose.
[0153] Fig. 23 shows an example of functions provided in a cloud-side information processing device. The cloud-side information processing device is a collective term for devices such as the cloud server 100 and the management server 500. As shown in Fig. 23, the cloud-side information processing device has a license authorization function F1, an account service function F2, a device monitoring function F3, a marketplace function F4, and a camera service function F5.
[0154] The license authorization function F1 is a function that performs various authentication-related processes. Specifically, the license authorization function F1 performs processes related to device authentication of multiple cameras 300 and processes related to authentication of each of the AI models, software, and firmware used in the cameras 300.
[0155] Here, the above software is software required for the camera 300 to properly perform AI image processing. In order to properly perform AI image processing based on captured images and to transmit the results of the AI image processing to the fog server 400 or the cloud server 100 in an appropriate format, it is necessary to control the data input to the AI model and properly process the output data of the AI model. The above software is software that includes peripheral processing necessary to properly perform AI image processing. Such software is software for realizing desired functions using the AI model and corresponds to the above-mentioned AI application.
[0156] Note that an AI application is not limited to one that uses only one AI model, and may use two or more AI models. For example, an AI application may be formed having a processing flow in which information on the recognition result obtained by an AI model that performs AI image processing using a captured image as input data is input to another AI model to perform second AI image processing. Here, the information on the recognition result is image data or the like, and hereinafter referred to as "recognition result information."
[0157] In the license authorization function F1, when the camera 300 is connected to the network 600, a process of issuing a device ID (Identification) for each camera 300 is performed to authenticate the camera 300. Furthermore, when the AI model or software is authenticated, a process of issuing a unique ID for each AI model or AI application for which registration has been applied from the AI model developer terminal 200C or the software developer terminal 700 is performed. The unique ID is an AI model ID, a software ID, or the like.
[0158] The license authorization function F1 also issues various keys, certificates, and the like to the manufacturer of the camera 300 (particularly the manufacturer of the image sensor IS described below), the AI model developer, and the software developer to ensure secure communication between the camera 300, the AI model developer terminal 200C, the software developer terminal 700, and the cloud server 100. Additionally, the license authorization function F1 also performs processing for updating and suspending the validity of the certificate. Furthermore, when user registration is performed by the account service function F2 described below, the license authorization function F1 also performs processing for linking the camera 300 (the above-mentioned device ID) purchased by the user with the user ID. Here, user registration is the registration of account information accompanied by the issuance of a user ID.
[0159] The account service function F2 is a function that generates and manages user account information. The account service function F2 accepts input of user information and generates account information based on the input user information. Here, the account information generated includes at least a user ID and password information. The account service function F2 also performs registration processing (account information registration) for AI model developers and AI application developers (hereinafter sometimes abbreviated as "software developers").
[0160] The device monitoring function F3 is a function that performs processing to monitor the usage status of the camera 300. For example, it monitors information such as the usage rate of the CPU and memory described above as various elements related to the usage status of the camera 300, such as the location where the camera 300 is used, the output frequency of output data from AI image processing, and the free space of the CPU and memory used for AI image processing.
[0161] The marketplace function F4 is a function for selling AI models and AI applications. For example, users can purchase AI applications and AI models used by AI applications via a sales website (sales site) provided by the marketplace function F4. Software developers can also purchase AI models for creating AI applications via the sales site.
[0162] The camera service function F5 is a function for providing the user with services related to the use of the camera 300. One example of this camera service function F5 is the function related to the generation of the analysis information described above. That is, one function of the camera service function F5 is a function for generating analysis information of a subject based on the processing result information of the image processing in the camera 300, and for performing processing to allow the user to view the generated analysis information via the user terminal 200.
[0163] The camera service function F5 also includes an imaging setting search function. Specifically, this imaging setting search function acquires recognition result information of AI image processing from the camera 300 and searches for imaging setting information of the camera 300 using AI based on the acquired recognition result information. Here, the imaging setting information is setting information related to the imaging operation for obtaining a captured image. Specific setting information includes at least optical setting information such as focus and aperture, setting information related to the readout operation of the captured image signal such as frame rate, exposure time, gain, and further setting information related to image signal processing of the readout captured image signal such as gamma correction processing, noise reduction processing, and super-resolution processing.
[0164] The camera service function F5 also includes an AI model search function. This AI model search function acquires recognition result information of AI image processing from the camera 300 and, based on the acquired recognition result information, uses AI to search for an optimal AI model to be used for the AI image processing in the camera 300. The AI model search here refers to, for example, a process of optimizing various processing parameters such as weighting coefficients and setting information related to the neural network structure (including, for example, kernel size information) when the AI image processing is realized by a convolutional neural network (CNN) that includes a convolution operation.
[0165] The camera service function F5 also has a processing allocation determination function. In the processing allocation determination function, when an AI application is deployed to an edge-side information processing device, a process for determining a deployment destination device for each SW component is performed as the above-mentioned deployment preparation process. Note that some SW components may be determined to be executed on a cloud-side device. In this case, the deployment process may not be performed because the SW components have already been deployed to the cloud-side device.
[0166] For example, in the case of an AI application that includes a SW component for detecting a person's face, a SW component for extracting person's attribute information, a SW component for aggregating the extraction results, and a SW component for visualizing the aggregation results, as in the example described above, the camera service function F5 makes the following decisions: The SW component for detecting a person's face determines the image sensor IS of the camera 300 as the deployment destination device; The SW component for extracting person's attribute information determines the camera 300 as the deployment destination device; The SW component for aggregating the extraction results determines the fog server 400 as the deployment destination device; and the SW component for visualizing the aggregation results determines to run on the cloud server 100 without newly deploying it to a device.
[0167] In this way, the camera service function F5 determines the deployment destination of each SW component, thereby determining the processing load for each device. Note that the processing load is determined taking into consideration the specifications and performance of each device, as well as the user's requests.
[0168] By providing the above-described imaging setting search function and AI model search function, imaging settings can be set to improve the results of AI image processing, and AI image processing can be performed using an appropriate AI model according to the actual usage environment. In addition, by providing a processing allocation determination function, AI image processing and its analysis processing can be performed in a device that is appropriate for the AI image processing and its analysis processing.
[0169] The camera service function F5 has an application setting function that sets an appropriate AI application according to the user's purpose before deploying each SW component.
[0170] For example, an appropriate AI application is selected in response to a user's selection of an application such as store monitoring or traffic monitoring. This automatically determines the SW components that make up the AI application. As will be described later, there may be multiple combinations of SW components for achieving the user's objectives using the AI application. In this case, one combination of SW components is selected in response to information from the edge-side information processing device and a user request.
[0171] For example, when a user intends to monitor a store, the combination of SW components may differ depending on whether the user's requirement is privacy-oriented or speed-oriented.
[0172] The application setting function involves processes such as accepting the user's operation to select a purpose (application) on the user terminal 200 (here, this corresponds to the application user terminal 200B shown in Figure 19), and selecting an appropriate AI application based on the selected application.
[0173] Here, in the above description, a configuration is exemplified in which the cloud server 100 alone realizes the license authorization function F1, account service function F2, device monitoring function F3, marketplace function F4, and camera service function F5. It is also possible to configure these functions to be shared and realized by multiple information processing devices. For example, a configuration in which each of the above multiple functions is performed by a single information processing device may be used. Furthermore, a single function among the above functions may be shared by multiple information processing devices. For example, a single function may be shared between the cloud server 100 and the management server 500.
[0174] 19 is an information processing device used by a developer of an AI model. Also, the software developer terminal 700 is an information processing device used by a developer of an AI application.
[0175] (4) Configuration of the Imaging Device Fig. 24 shows an example of the internal configuration of a camera 300. As shown in Fig. 24, the camera 300 includes an imaging optical system 310, an optical system driving unit 320, an image sensor IS, a control unit 330, a memory unit 340, and a communication unit 350. The image sensor IS, the control unit 330, the memory unit 340, and the communication unit 350 are each connected via a bus 360, which enables data communication between them.
[0176] The imaging optical system 310 includes lenses such as a cover lens, a zoom lens, and a focus lens, as well as an iris mechanism. Light (incident light) from a subject is guided by the imaging optical system 310, and the light is collected on the light receiving surface of the image sensor IS.
[0177] The optical system driving unit 320 collectively refers to the driving units for the zoom lens, focus lens, and diaphragm mechanism of the imaging optical system 310. Specifically, the optical system driving unit 320 has actuators and actuator driving circuits for driving the zoom lens, focus lens, and diaphragm mechanism, respectively.
[0178] The control unit 330 is configured with, for example, a microcomputer having a CPU, ROM, and RAM, and performs overall control of the camera 300 by the CPU executing various processes in accordance with programs stored in the ROM or programs loaded into the RAM.
[0179] Furthermore, the control unit 330 issues drive instructions to the optical system drive unit 320 to drive the zoom lens, focus lens, diaphragm mechanism, etc. In response to these drive instructions, the optical system drive unit 320 moves the focus lens and zoom lens, opens and closes the diaphragm blades of the diaphragm mechanism, etc.
[0180] The control unit 330 also controls the writing and reading of various data to and from the memory unit 340. The memory unit 340 includes a non-volatile storage device such as a hard disk drive (HDD) or a flash memory device. The memory unit 340 is used as a storage destination (recording destination) for image data output from the image sensor IS.
[0181] Furthermore, the control unit 330 performs various data communications with external devices via the communication unit 350. The communication unit 350 in the embodiment is capable of data communications with at least the fog server 400 (or the cloud server 100) shown in FIG.
[0182] The image sensor IS is configured as, for example, a CCD type, a CMOS type, or the like image sensor.
[0183] The image sensor IS includes an imaging unit 410, an image signal processing unit 420, an internal sensor control unit 430, an AI image processing unit 440, a memory unit 450, and a communication I / F 460. These are connected via a bus 470 and are capable of mutual data communication.
[0184] The imaging unit 410 includes a pixel array unit in which a plurality of pixels are arranged two-dimensionally, and a readout circuit. The pixels include photoelectric conversion elements such as photodiodes. The readout circuit reads out electrical signals obtained by photoelectric conversion from each pixel in the pixel array unit. The imaging unit 410 outputs the obtained electrical signals as captured image signals.
[0185] The readout circuit performs, for example, CDS (Correlated Double Sampling) processing, AGC (Automatic Gain Control) processing, etc. on the electrical signal obtained by photoelectric conversion, and further performs A / D (Analog / Digital) conversion processing.
[0186] The image signal processing unit 420 performs pre-processing, synchronization processing, YC generation processing, resolution conversion processing, codec processing, etc. on the captured image signal as digital data after A / D conversion processing.
[0187] Preprocessing includes clamping the R, G, and B black levels of the captured image signal to a predetermined level, and correction between the R, G, and B color channels. Synchronization processing involves color separation processing so that image data for each pixel contains all R, G, and B color components. For example, in the case of an image sensor using a Bayer array color filter, demosaic processing is performed as the color separation processing. YC generation processing involves generating (separating) a luminance (Y) signal and a color (C) signal from the R, G, and B image data. Resolution conversion processing involves executing resolution conversion on image data that has undergone various signal processes.
[0188] In codec processing, the image data that has undergone the various processes described above is subjected to encoding processing for recording or communication, and a file is generated. In codec processing, moving image file formats such as MPEG-2 (Moving Picture Experts Group) and H.264 can be generated. Still image file formats such as JPEG (Joint Photographic Experts Group), TIFF (Tagged Image File Format), and GIF (Graphics Interchange Format) can be generated.
[0189] The sensor control unit 430 issues instructions to the imaging unit 410 and controls the execution of imaging operations. Similarly, the sensor control unit 430 also controls the image signal processing unit 420 to execute processing.
[0190] The AI image processing unit 440 performs image recognition processing as AI image processing on the captured image. The image recognition function using AI can be realized using a programmable arithmetic processing device such as a CPU, a field programmable gate array (FPGA), or a digital signal processor (DSP).
[0191] The image recognition functions that can be realized in the AI image processing unit 440 can be switched by changing the algorithm of the AI image processing. In other words, the function type of the AI image processing can be switched by switching the AI model used in the AI image processing. The function types of the AI image processing are, for example, as follows: Class identification Semantic segmentation Person detection Vehicle detection Target tracking Optical character recognition (OCR)
[0192] Of the above function types, class identification is a function that identifies the class of a target. This "class" is information that represents the category of an object. For example, classes distinguish between "people," "cars," "airplanes," "ships," "trucks," "birds," "cats," "dogs," "deer," "frogs," and "horses." Target tracking is a function that tracks a targeted subject. In other words, target tracking is a function that obtains historical information about the subject's position.
[0193] The memory unit 450 is used as a storage destination for various data such as captured image data obtained by the image signal processing unit 420. In the embodiment, the memory unit 450 is also used for temporary storage of data used by the AI image processing unit 440 in the process of AI image processing.
[0194] The memory unit 450 also stores information on AI applications and AI models used in the AI image processing unit 440. The information on the AI applications and AI models may be deployed in the memory unit 450 as a container or the like using the container technology described below. The information on the AI applications and AI models may also be deployed using microservice technology. By deploying the AI models used for AI image processing in the memory unit 450, it is possible to change the function type of the AI image processing, or to change to an AI model whose performance has been improved by relearning.
[0195] Note that the above-described embodiments have been described based on examples of AI models and AI applications used for image recognition. The present technology is not limited to this, and may also be applicable to programs executed using AI technology. Furthermore, if the capacity of the memory unit 450 is small, information on the AI application or AI model may be stored in a memory outside the image sensor IS using container technology, such as the memory unit 340, as a container, and then only the AI model may be stored in the memory unit 450 within the image sensor IS via the communication I / F 460 described below.
[0196] The communication I / F 460 is an interface for communicating with the control unit 330, memory unit 340, etc., which are external to the image sensor IS. The communication I / F 460 communicates to acquire from the outside the program executed by the image signal processing unit 420, the AI application used by the AI image processing unit 440, the AI model, etc. This information is stored in the memory unit 450 provided in the image sensor IS. As a result, the AI model, etc., is stored in part of the memory unit 450 provided in the image sensor IS and can be used by the AI image processing unit 440.
[0197] The AI image processing unit 440 performs predetermined image recognition processing using the AI application and AI model obtained in this manner, thereby recognizing the subject according to the purpose. Recognition result information from the AI image processing is output to the outside of the image sensor IS via the communication I / F 460. That is, the communication I / F 460 of the image sensor IS outputs not only the image data output from the image signal processing unit 420 but also the recognition result information from the AI image processing. Note that the communication I / F 460 of the image sensor IS can output only either the image data or the recognition result information.
[0198] For example, when using the re-learning function of the AI model described above, the captured image data used for the re-learning function is uploaded from the image sensor IS to an information processing device on the cloud side via the communication I / F 460 and the communication unit 350.
[0199] In addition, when inference is performed using an AI model, recognition result information from the AI image processing is output from the image sensor IS to another information processing device outside the camera 300 via the communication I / F 460 and the communication unit 350.
[0200] The image sensor IS may have various structures. In the examples, the structure of an image sensor IS having a two-layer stacked structure will be described. Fig. 25 shows an example of the structure of an image sensor IS as an imaging device. As shown in Fig. 25, the image sensor IS is formed by a semiconductor device in which two semiconductor chips, dies D1 and D2, are stacked together to form a single semiconductor chip.
[0201] The die D1 has the function of the imaging unit 410 shown in Fig. 24. The die D2 has the functions of an image signal processing unit 420, an internal sensor control unit 430, an AI image processing unit 440, a memory unit 450, and a communication I / F 460.
[0202] The die D1 and the die D2 each have terminals on their opposing surfaces. The terminals are formed using, for example, copper (Cu) as a wiring material. That is, the die D1 and the die D2 are electrically connected by Cu-Cu bonding that joins the terminals together.
[0203] An example using container technology will be described below as a method for deploying AI models, AI applications, and the like in the camera 300. Fig. 26 shows an example of the software configuration of an imaging device.
[0204] As shown in FIG. 26, the camera 300 has an operation system 510 installed on various hardware 501 such as a CPU, a GPU (Graphics Processing Unit), a ROM, and a RAM, which constitute the control unit 330 shown in FIG.
[0205] The operation system 510 is basic software that performs overall control of the camera 300 in order to realize various functions of the camera 300. On the operation system 510, a general-purpose middleware 520 is installed.
[0206] The general-purpose middleware 520 is software for realizing basic operations such as a communication function using the communication unit 350 as the hardware 501 and a display function using a display unit (monitor, etc.) as the hardware 501.
[0207] On the operation system 510, not only the general-purpose middleware 520 but also an orchestration tool 530 and a container engine 540 are installed.
[0208] The orchestration tool 530 and the container engine 540 deploy and execute the container 550 by constructing a cluster 560 as an operating environment for the container 550. Note that the edge runtime shown in FIG. 22 corresponds to the orchestration tool 530 and the container engine 540 shown in FIG. 26.
[0209] The orchestration tool 530 has a function for causing the container engine 540 to appropriately allocate resources of the above-described hardware 501 and operation system 510. The orchestration tool 530 groups the containers 550 into predetermined units (pods, which will be described later), and deploys each pod to a worker node (which will be described later) that is set as a logically different area.
[0210] The container engine 540 is one of the middlewares installed in the operation system 510, and is an engine that operates the container 550. Specifically, the container engine 540 has a function of allocating resources (memory, computing power, etc.) of the hardware 501 and the operation system 510 to the container 550 based on a configuration file or the like provided in the middleware in the container 550.
[0211] In addition, in the embodiment, the resources to be allocated include not only resources such as the control unit 330 provided in the camera 300, but also resources such as the sensor control unit 430, memory unit 450, and communication I / F 460 provided in the image sensor IS.
[0212] The container 550 includes an application for realizing a predetermined function and middleware such as a library. The container 550 operates to realize the predetermined function using the resources of the hardware 501 and the operation system 510 allocated by the container engine 540.
[0213] In the embodiment, the AI application and AI model shown in Fig. 22 correspond to one of the containers 550. That is, one of the various containers 550 deployed in the camera 300 realizes a predetermined AI image processing function using the AI application and AI model.
[0214] 27 shows an example of a specific configuration of a cluster 560 constructed by the container engine 540 and the orchestration tool 530. Note that the cluster 560 may be constructed across multiple devices so that functions are realized using not only the hardware 501 provided in one camera 300 but also other hardware resources provided in other devices.
[0215] The orchestration tool 530 manages the execution environment of the container 550 for each worker node 570. The orchestration tool 530 also constructs a master node 580 that manages the entire worker node 570.
[0216] A plurality of pods 590 are deployed in the worker node 570. Each pod 590 includes one or more containers 550 and realizes a predetermined function. The pod 590 is used as a management unit for managing the containers 550 by the orchestration tool 530.
[0217] The operation of the pod 590 in the worker node 570 is controlled by a pod management library 601. The pod management library 601 is configured to have a network proxy and the like. The network proxy includes a container runtime that allows the pod 590 to use the resources of the logically allocated hardware 501, and a network proxy that performs communication between agents and pods 590 controlled by the master node 580, and communication with the master node 580. In other words, the pod management library 601 enables the multiple pods 590 to realize predetermined functions using each resource.
[0218] The master node 580 includes an application server 610, a manager 620, a scheduler 630, and a data sharing unit 640. The application server 610 deploys a pod 590. The manager 620 manages the deployment status of a container 550 by the application server 610. The scheduler 630 determines a worker node 570 on which to place the container 550. Then, the data sharing unit 640 shares data.
[0219] By utilizing the configurations shown in Figures 26 and 27, it is possible to deploy the above-mentioned AI applications and AI models to the image sensor IS of the camera 300 using container technology.
[0220] As mentioned above, the AI model may be stored in the memory unit 450 in the image sensor IS via the communication I / F 460 shown in Fig. 24, and the AI image processing may be executed in the image sensor IS. Also, the AI application and AI model may be deployed in the memory unit 450 and the in-sensor control unit 430 in the image sensor IS using the configurations shown in Figs. 26 and 27, and executed in the image sensor IS using container technology.
[0221] Furthermore, as will be described later, when the AI application and / or the AI model are deployed on the fog server 400 or a cloud-side information processing device, container technology may be used. In this case, the information of the AI application and the AI model is deployed as a container or the like in a memory such as the non-volatile memory unit 740, the storage unit 790, or the RAM 730 shown in FIG. 28 to be described later, and executed.
[0222] (5) Hardware Configuration of Information Processing Device FIG. 28 shows an example of the hardware configuration of information processing devices such as the cloud server 100, the user terminal 200, the fog server 400, and the management server 500 included in the information processing system 1000.
[0223] The information processing device includes a CPU 710. The CPU 710 functions as an arithmetic processing unit that performs the various processes described above. The CPU 710 executes the various processes in accordance with programs stored in the ROM 720 or the nonvolatile memory unit 740, or programs loaded from the storage unit 790 to the RAM 730. The nonvolatile memory unit 740 may be, for example, an EEPROM (Electrically Erasable Programmable Read Only Memory). The RAM 730 also stores data and the like required for the CPU 710 to execute the various processes, as appropriate.
[0224] In addition, the CPU 710 provided in the information processing device serving as the cloud server 100 functions as a license authorization unit, an account service providing unit, a device monitoring unit, a marketplace function providing unit, a camera service providing unit, etc. in order to realize each of the above-mentioned functions.
[0225] The CPU 710, ROM 720, RAM 730, and nonvolatile memory unit 740 are interconnected via a bus 830. In addition, an input / output interface (I / F) 750 is connected to the bus 830.
[0226] An input unit 760 including operators and operation devices is connected to the input / output interface 750. The input unit 760 is, for example, various operators and operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, a remote controller, etc. The input unit 760 detects user operations, and the CPU 710 interprets signals corresponding to the input operations.
[0227] Furthermore, a display unit 770 including an LCD, an organic EL panel, or the like, and an audio output unit 780 including a speaker, or the like, are connected to the input / output interface 750 either integrally or separately. The display unit 770 is a display unit that displays various information. The display unit 770 is configured, for example, by a display device provided in the housing of the computer device, a separate display device connected to the computer device, or the like.
[0228] The display unit 770 displays images for various types of image processing, moving images to be processed, etc. on the display screen based on instructions from the CPU 710. The display unit 770 also displays various operation menus, icons, messages, etc. based on instructions from the CPU 710. A GUI (Graphical User Interface) is used for these displays.
[0229] The input / output interface 750 may be connected to a storage unit 790 configured with a hard disk, solid-state memory, etc., and a communication unit 801 configured with a modem, etc.
[0230] The communication unit 801 performs communication processing via a transmission path such as the Internet, and communication with various devices via wired / wireless communication, bus communication, and the like.
[0231] A drive 810 is connected to the input / output interface 750 as needed. A removable storage medium 820 is appropriately attached to the input / output interface 750 via the drive 810. The removable storage medium 820 includes a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like.
[0232] Drive 810 can read data files such as programs used in various processes from removable storage medium 820. The read data files are stored in storage unit 790. Images included in the data files are output on display unit 770, and audio included in the data files are output on audio output unit 780. Computer programs and the like read from removable storage medium 820 are installed in storage unit 790 as necessary.
[0233] In this computer device, for example, software for the processing of the embodiment can be installed via network communication by the communication unit 801 or via the removable storage medium 820. The software may also be stored in advance in the ROM 720, the storage unit 790, etc. Furthermore, captured images captured by the camera 300 and processing results of AI image processing may be received, and the captured images and processing results may be stored in the storage unit 790 or the removable storage medium 820.
[0234] The CPU 710 performs processing operations based on various programs, thereby executing the necessary information processing and communication processing of the cloud server 100, which is an information processing device equipped with the above-described arithmetic processing unit. The cloud server 100 is not limited to being configured by a single computer device as shown in FIG. 23 , but may also be configured by a system of multiple computer devices. The multiple computer devices are systemized using, for example, a LAN (Local Area Network) or the like. Furthermore, multiple computer devices located in remote locations may be systemized using a VPN (Virtual Private Network) using the Internet or the like. The multiple computer devices may include computer devices serving as a server group (cloud) available through a cloud computing service.
[0235] An imaging device and an imaging device system to which the present technology can be further applied will be described with reference to Fig. 29. Fig. 29 shows an example of a schematic configuration illustrating a processing flow when updating an AI model or an AI application in an imaging device and an imaging device system.
[0236] After the SW components and AI models of the AI application described in the embodiments are deployed, operations by a service provider or user trigger re-learning of the AI model and updating of the AI model (hereinafter referred to as the "edge-side AI model") deployed to multiple cameras 300 and the AI application. FIG. 29 shows one camera 300 of interest among the multiple cameras 300. In the following description, the edge-side AI model to be updated is deployed to the image sensor IS included in the camera 300, as an example. Note that the edge-side AI model may be deployed outside the image sensor IS in the camera 300.
[0237] First, in processing step PS1, a service provider or user issues an instruction to retrain the AI model. This instruction is issued using an API (Application Programming Interface) function provided by an API module in the cloud-side information processing device. The instruction also specifies the amount of images (e.g., number) to be used for training. Hereinafter, the amount of images to be used for training may be referred to as a "predetermined number."
[0238] Upon receiving this instruction, the API module transmits a re-learning request and image volume information to the Hub (see FIG. 22) in processing step PS2.
[0239] In processing step PS3, the Hub transmits an update notification and information on the image volume to the camera 300 as an edge-side information processing device.
[0240] In processing step PS4, the camera 300 transmits the captured image data obtained by capturing an image to an image database (DB) in the storage group. This capturing and transmitting process is repeated until a predetermined number of images required for re-learning is captured.
[0241] In addition, when the camera 300 obtains an inference result by performing inference processing on the captured image data, in processing step PS4, the inference result may be stored in the image DB as metadata for the captured image data.
[0242] By storing the inference results from the camera 300 as metadata in the image DB, it is possible to carefully select the data necessary for relearning the AI model executed on the cloud side. Specifically, relearning can be performed using only image data where the inference results from the camera 300 differ from the inference results executed on the cloud side using abundant computer resources. This makes it possible to shorten the time required for relearning.
[0243] After capturing and transmitting the predetermined number of images, the camera 300 notifies the Hub in processing step PS5 that transmission of the predetermined number of captured image data has been completed.
[0244] Upon receiving the notification, the Hub notifies the orchestration tool in processing step PS6 that the preparation of the re-learning data has been completed.
[0245] In processing step PS7, the orchestration tool sends an instruction to execute the labeling process to the labeling module.
[0246] The labeling module acquires the image data to be subjected to labeling processing from the image DB (processing step PS8) and performs labeling processing.
[0247] Here, the labeling process is a process of performing the class identification described above. The labeling process is also a process of estimating the gender and age of the subject of the image and assigning a label. The labeling process is also a process of estimating the pose of the subject and assigning a label. The labeling process is also a process of estimating the behavior of the subject and assigning a label.
[0248] The labeling process may be performed manually or automatically, and may be completed in an information processing device on the cloud side, or may be realized by using a service provided by another server device.
[0249] After completing the labeling process, the labeling module stores the labeling result information in the dataset DB in processing step PS9. Here, the information stored in the dataset DB may be a combination of label information and image data, or may be image ID (Identification) information for identifying the image data instead of the image data itself.
[0250] The storage management unit, which detects that the labeling result information has been stored, notifies the orchestration tool in processing step PS10.
[0251] The orchestration tool, which has received this notification, confirms that labeling processing has been completed for a predetermined number of image data, and in processing step PS11, transmits a re-learning instruction to the re-learning module.
[0252] Upon receiving the re-learning instruction, the re-learning module acquires the dataset to be used for learning from the dataset DB in processing step PS12, and acquires the AI model to be updated from the trained AI model DB in processing step PS13.
[0253] The re-learning module re-learns the AI model using the acquired data set and the AI model. The updated AI model obtained in this way is stored again in the trained AI model DB in processing step PS14.
[0254] When the storage management unit detects that the updated AI model has been stored, it notifies the orchestration tool in processing step PS15.
[0255] Upon receiving the notification, the orchestration tool sends an instruction to convert the AI model to the conversion module in processing step SP16.
[0256] Upon receiving the conversion instruction, the conversion module retrieves the updated AI model from the trained AI model DB and performs conversion processing of the AI model in processing step PS17. This conversion processing involves conversion to match the specification information of the camera 300, which is the device to which the AI model will be deployed. This processing involves downsizing the AI model so as to minimize degradation in performance, and converting the file format so that the AI model can run on the camera 300.
[0257] The AI model converted by the conversion module is the edge-side AI model described above, and is stored in the converted AI model DB in processing step PS18.
[0258] When the storage management unit detects that the converted AI model has been stored, it notifies the orchestration tool in processing step PS19.
[0259] Upon receiving the notification, the orchestration tool sends a notification to the Hub to update the AI model in processing step PS20. This notification includes information for identifying the location where the AI model to be used for the update is stored.
[0260] The Hub, which has received the notification, transmits an instruction to update the AI model to the camera 300. The update instruction also includes information for specifying the location where the AI model is stored.
[0261] In processing step PS22, the camera 300 performs processing to acquire the target converted AI model from the converted AI model DB and deploy it, thereby updating the AI model used in the image sensor IS of the camera 300.
[0262] After the camera 300 has completed updating the AI model by deploying the AI model, it transmits an update completion notification to the Hub in processing step PS23. Upon receiving the notification, the Hub notifies the orchestration tool in processing step PS24 that the AI model update process for the camera 300 has been completed.
[0263] Note that, here, an example is described in which the AI model is deployed and used within the image sensor IS of the camera 300 (e.g., the memory unit 450 shown in FIG. 24 ). The present technology may also be applicable to a case in which the AI model is deployed and used outside the image sensor IS of the camera 300 (e.g., the memory unit 340 shown in FIG. 24 ) or in a storage unit (not shown) within the fog server 400. Similarly, the AI model can be updated. In this case, when the AI model is deployed, the device (location) where the AI model is deployed is stored in a storage management unit or the like on the cloud side. The Hub reads the device (location) where the AI model is deployed from the storage management unit and transmits an instruction to update the AI model to the device where the AI model is deployed. In processing step PS22, the device that receives the update instruction performs a process of retrieving the target converted AI model from the converted AI model DB and deploying it. This updates the AI model of the device that received the update instruction.
[0264] If only the AI model is updated, the process is completed up to this point. If an AI application that uses the AI model is updated in addition to the AI model, the process described below is further executed.
[0265] Specifically, in processing step PS25, the orchestration tool sends a download instruction for an AI application such as updated firmware to the deployment control module.
[0266] The deployment control module sends an AI application deployment instruction to the Hub in process step PS26, which includes information specifying where the updated AI application is stored.
[0267] In processing step PS27, the Hub transmits the deployment instruction to the camera 300. In processing step PS28, the camera 300 downloads the updated AI application from the container DB of the deployment control module and deploys it.
[0268] The above describes an example in which an AI model running on the image sensor IS of the camera 300 and an AI application running outside the image sensor IS of the camera 300 are updated sequentially. The AI application is simply described here. As described above, an AI application is defined by multiple SW components, such as SW components B1, B2, B3, ..., Bn. When an AI application is deployed, a storage management unit or the like on the cloud side stores the locations of the multiple SW components. When performing processing step PS27, the Hub reads from the storage management unit the device (location) where each SW component is deployed and sends a deployment instruction to the deployed device. In processing step PS28, the device that receives the deployment instruction downloads the updated SW components from the container DB of the deployment control module and deploys them. Here, the AI application is a SW component other than the AI model.
[0269] Furthermore, when both an AI model and an AI application run on a single device, they may be updated together as a single container. In this case, the AI model and the AI application may be updated simultaneously, rather than sequentially. The updates can be achieved by executing the processes of processing steps PS25, PS26, PS27, and PS28.
[0270] For example, if it is possible to deploy containers of both an AI model and an AI application to the image sensor IS of the camera 300, the AI model and the AI application can be updated by executing the processes of processing steps PS25, PS26, PS27, and PS28 as described above.
[0271] By performing the above-described processing, the AI model is retrained using image data captured in the user's usage environment, thereby generating an edge-side AI model that can output highly accurate recognition results in the user's usage environment.
[0272] Furthermore, even if the user's usage environment changes, such as when the store layout is changed or the installation location of the camera 300 is changed, the AI model can be appropriately retrained each time. This makes it possible to maintain the recognition accuracy of the AI model without any degradation. Note that the above-described processes may be executed not only when the AI model is retrained, but also when the system is operated for the first time in the user's usage environment.
[0273] An imaging device and an imaging device system to which the present technology can be applied will be further described with reference to Figures 30 to 32. Figure 30 shows an example of a display presented to a user regarding a marketplace in an imaging device and an imaging device system.
[0274] The login screen G1 displays an ID input field 910 for inputting a user ID and a password input field 920 for inputting a password.
[0275] Below the password input field 920, a login button 930 for logging in and a cancel button 940 for canceling the login are arranged.
[0276] Further below that, operators for transitioning to a page for users who have forgotten their password, operators for transitioning to a page for new user registration, and the like are appropriately arranged.
[0277] After entering an appropriate user ID and password, when the user presses the login button 930, a process of transitioning to a page specific to the user is executed in each of the cloud server 100 and the user terminal 200.
[0278] FIG. 31 shows an example of a display presented to, for example, an AI application developer using the application developer terminal 200A, an AI model developer using the AI model developer terminal 200C, etc.
[0279] Developers can purchase training datasets, AI models, and AI applications for development purposes through the marketplace, and can also register their own AI applications and models on the marketplace.
[0280] 31, available training datasets, AI models, AI applications, etc. (hereinafter collectively referred to simply as "data") are displayed on the left side. Although not shown, when purchasing a training dataset, the developer can prepare for training by simply displaying an image of the training dataset on a display, using an input device such as a mouse to frame only the desired portion of the image, and entering a name.
[0281] For example, if you want to use an image of a cat for AI learning, you can prepare an image with a cat annotation for AI learning by drawing a frame around only the cat part of the image and entering "cat" as text input. Furthermore, to make it easier to find the desired data, you can select objectives such as "traffic monitoring," "traffic flow analysis," and "customer count." That is, a display process is executed on both the cloud server 100 and the user terminal 200 so that only data that matches the selected objective is displayed.
[0282] The developer screen G2 may also display the purchase price of each piece of data.
[0283] In addition, on the right side of the developer screen G2, there are provided input fields 950 for registering learning datasets collected or created by the developer, AI models developed by the developer, and AI applications.
[0284] The developer screen G2 has an input field 950 for inputting the name and storage location of each piece of data. Also, for the AI model, a check box 960 is provided for setting whether or not retraining is required.
[0285] It should be noted that a price setting field (shown as an input field 950 in FIG. 31) may be provided in which the price required to purchase the data to be registered can be set.
[0286] Furthermore, at the top of the developer screen G2, as part of the user information, the user name, the last login date, etc. In addition to this, the amount of currency that the user can use when purchasing data, the number of points, etc. may also be displayed.
[0287] 32 shows an example of a display presented to an application user. This display example is an example of a user screen G3 that is presented to an application user who performs various analyses by deploying an AI application or an AI model in the camera 300 as an edge-side information processing device that the application user manages.
[0288] A user can purchase a camera 300 to be placed in a space to be monitored via the marketplace. Therefore, on the left side of the user screen G3, radio buttons 970 are provided that allow the user to select the type and performance of the image sensor IS to be installed in the camera 300, as well as the performance of the camera 300.
[0289] Furthermore, a user can purchase an information processing device as the fog server 400 via the marketplace. Therefore, radio buttons 970 for selecting each performance of the fog server 400 are arranged on the left side of the user screen G3. Furthermore, a user who already owns a fog server 400 can register the performance of the fog server 400 by inputting the performance information of the fog server 400 here.
[0290] A user can realize the desired functions by installing the purchased camera 300 in a location of their choice, such as a store run by the user. In the marketplace, information about the installation locations of the cameras 300 can be registered in order to maximize the functionality of multiple cameras 300. Note that the camera 300 may also be purchased without going through the marketplace.
[0291] On the right side of the user screen G3, there are arranged radio buttons 980 for selecting environmental information about the environment in which the camera 300 is installed. By appropriately selecting the environmental information about the environment in which the camera 300 is installed, the user can set the above-mentioned optimal image capture settings for the target camera 300.
[0292] Furthermore, when purchasing a camera 300 and the installation location of the camera 300 has already been decided, the user can purchase a camera 300 with optimal imaging settings pre-configured for the planned installation location by selecting each item on the left side and each item on the right side of the user screen G3.
[0293] The user screen G3 has an execute button 990. Pressing the execute button 990 can transition to a confirmation screen for confirming the purchase or a confirmation screen for confirming the setting of environmental information. This allows the user to purchase the desired camera 300 and fog server 400 and set the environmental information for the camera 300.
[0294] In the marketplace, it is possible to change the environmental information of multiple cameras 300 in case the installation location of the camera 300 is changed. By re-entering the environmental information about the installation location of the camera 300 on a change screen (not shown), it is possible to reset the optimal imaging settings for the camera 300.
[0295] The present technology is not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology. For example, two or more solid-state imaging devices according to the embodiments may be combined.
[0296] [Application Example of Information Processing Apparatus] Next, an application example of the cloud-side information processing apparatus will be described. The cloud-side information processing apparatus according to the embodiment can realize the functions of the information processing apparatus described in each embodiment.
[0297] (First Application Example) Figure 33 is a diagram for explaining a first application example of a cloud-side information processing device. In this example, a user terminal 200 (for example, an AI model developer terminal 200C) downloads an image dataset for learning from a cloud-side information processing device. In this example, the cloud-side information processing device functions as a dataset improvement tool that improves bias contained in the dataset.
[0298] The cloud-side information processing device reads out a dataset before the bias of the attribute is improved (hereinafter referred to as the dataset before improvement as appropriate) from a storage unit (for example, the storage unit 790). The dataset to be read out may be a dataset specified by the user terminal 200, or may be a dataset automatically selected by the cloud-side information processing device according to the task content of the AI model to be developed.
[0299] The information processing device on the cloud side processes the unimproved dataset using the attribute extraction unit 2, labeling unit 3, and bias detection unit 4. The processing described in each embodiment can be applied to the processing content.
[0300] Display information showing the labeling results by the labeling unit 3 is transmitted to the user terminal 200. Then, the number of image data for each attribute is displayed on a display unit or the like of the user terminal 200 as an example of the display information. The user of the user terminal 200 inputs, for example, that they want to eliminate the influence of the attribute "race", that is, that they want to improve the bias of the attribute "race". This input is transmitted from the user terminal 200 to an information processing device on the cloud side.
[0301] The cloud-side information processing device adjusts the contents of the pre-improvement dataset based on the information transmitted from the user terminal 200. For example, the dataset adjustment unit 5 adjusts the number of image data so that the number of image data for the attribute "race" specified by the user terminal 200 is equal between men and women. As a result, the cloud-side information processing device generates an improved dataset in which at least one attribute present in the pre-improvement dataset has been improved. The improved dataset is transmitted by the user terminal 200.
[0302] The result of the bias detection unit 4 may be displayed on the user terminal 200. The number of images in the dataset after improvement may be displayed on the user terminal 200. Bias improvement of the dataset may be performed for a fee. Furthermore, if the user terminal 200 does not specify an attribute for which bias is to be improved, all biases may be improved.
[0303] By applying the information processing device as a dataset improvement tool as described above, biases that humans may overlook can be detected. Also, attributes for which biases should be eliminated can be specified on the user terminal 200. It is also possible to consider attribute combinations that involve bias and attribute combinations that eliminate bias.
[0304] (Second Application Example) Next, a second application example will be described, which is also an example in which the cloud-side information processing device is used as a dataset improvement tool.
[0305] 34 is a diagram illustrating a second application example of the cloud-side information processing device. In this example, the process described in the first application example is first performed. That is, when a user specifies the attribute "race" for a data set before improvement, the bias present in the attribute "race" is improved, and an improved data set is generated.
[0306] Furthermore, the user can check whether there is bias for attributes that have not been presented. For example, if the user wants to check whether there is bias for the attribute "makeup," the user specifies the attribute "makeup" and transmits it from the user terminal 200 to the cloud-side information processing device.
[0307] In the cloud-side information processing device, the attribute extraction unit 2 extracts whether the attribute "makeup" is included in image data constituting the dataset (in this case, a dataset in which bias for the attribute "race" has been improved). Then, as shown in Fig. 34, the labeling results for the attribute "makeup" by the labeling unit 3 and the detection results by the bias detection unit 4 are displayed on the user terminal 200. If bias exists in the attribute "makeup", the user terminal 200 may request the cloud-side information processing device to improve the bias.
[0308] According to this example, the user can define any attribute even if it was not discovered by the dataset improvement tool, and can then check whether or not the attribute is biased. Based on the results of the check, the user can also specify another attribute. In this way, the contents of the dataset can be checked and improved until the user is satisfied with the presence or absence of bias in the attribute.
[0309] (Third Application Example) Next, a third application example will be described, which is an example in which an information processing device on the cloud side is used as a dataset assessment tool.
[0310] 35A and 35B are diagrams illustrating a third application example of a cloud-side information processing device. As shown in Fig. 35A, the cloud-side information processing device processes a data set before bias improvement using an attribute extraction unit 2, a labeling unit 3, and a bias detection unit 4. Then, attributes in which bias has been detected by the bias detection unit 4 are transmitted to a user terminal 200.
[0311] The user terminal 200 displays a list of attributes detected by the cloud-side information processing device (in the illustrated example, the attributes "glasses," "whether or not makeup is worn," "age," etc.). The listed attributes can be compared with attributes that the user has defined as having a bias. In other words, according to this example, attributes that the user has defined in advance using their expert knowledge (attributes that they believe to have a bias) can be compared with attributes that have been detected by the cloud-side information processing device as having a bias. As a result of the comparison, the degree of match between the two can be obtained.
[0312] 35B , the degree of imbalance calculated when the bias is detected by the bias detection unit 4 may be displayed. In this case, the degree of imbalance for each attribute may be displayed, or the average of the degrees of imbalance for each attribute may be displayed, or both may be displayed.
[0313] (Fourth Application Example) Next, a fourth application example will be described. As described above, a learning dataset can be downloaded from the cloud-based information processing device. In this example, the user can recognize the bias of the dataset via the user terminal 200.
[0314] For example, when the user terminal 200 accesses the cloud-side information processing device, the screen illustrated in FIG. 36 is displayed on the display unit or the like of the user terminal 200. For example, information indicating a first dataset and information indicating a second dataset are displayed as examples of multiple datasets. Of course, the number of datasets may be other than two. The information indicating the first dataset includes, for example, information such as text information identifying the first dataset, the labeling results for the first dataset by the labeling unit 3 and / or the bias detection results for each attribute by the bias detection unit 4 (e.g., numerical values indicating the degree of imbalance). The information indicating the second dataset includes, for example, information such as text information identifying the second dataset, the labeling results for the second dataset by the labeling unit 3 and / or the bias detection results for each attribute by the bias detection unit 4 (e.g., numerical values indicating the degree of imbalance).
[0315] As described above, depending on the task content of the AI model, it may be preferable for some attributes to have bias. By displaying the labeling results by the labeling unit 3 and / or the bias detection results for each attribute by the bias detection unit 4 together with information identifying the data, the user can select and download a dataset suitable for the AI model to be developed. Note that the content shown in FIG. 36 may be displayed on the display unit 6A or the like instead of the user terminal 200.
[0316] As shown in Fig. 37, it is also possible to display only information about attributes defined by the user for each data set, which allows the user to easily compare data sets to see whether there is bias in the attributes required by the user.
[0317] Furthermore, the cloud-side information processing device may be configured to search for a dataset using an attribute and whether or not the attribute is biased as a key. For example, if a user requests a dataset in which bias exists only in the attribute "whether or not there is a smile," the cloud-side information processing device may search for a dataset that suits the request. Furthermore, if a dataset that suits the request does not exist, the dataset adjustment unit 5 may perform processing so that bias exists only in the attribute "whether or not there is a smile," thereby allowing the cloud-side information processing device to generate a dataset that suits the request.
[0318] (Fifth Application Example) Next, a fifth application example will be described. In this example, an information processing device on the cloud side performs quality assurance on a data set used for training an AI model.
[0319] For example, as shown in FIG. 38 , consider a gender classification model and a human detection model as examples of AI models to be made public by a cloud-side information processing device. The cloud-side information processing device inspects the datasets used to train each AI model. Specifically, the cloud-side information processing device inspects whether or not bias exists with respect to at least ethical attributes in the datasets used to train each AI model. The presence or absence of bias with respect to ethical attributes is determined by applying the method described in the embodiment.
[0320] When the AI model is made public on the cloud-side information processing device, test result information corresponding to the test result is also presented. For example, the bias detection unit 4 or the like detects that a bias exists with respect to the attribute "race," which is an example of an ethical attribute, in the dataset used to train the gender classification model. In this case, when the gender classification model is made public on the cloud-side information processing device, it is clearly indicated that a bias exists with respect to the ethical attribute. In other words, it is clearly indicated that the content of the dataset used to train the gender classification model is inappropriate. Note that, if a bias exists with respect to the ethical attribute, the AI model trained using the dataset including the bias may not be made public on the cloud-side information processing device.
[0321] Furthermore, for example, assume that no bias with respect to ethical attributes was detected in the dataset used to train the human detection model. In this case, when the human detection model is made public on the cloud-side information processing device, it is clearly indicated that there is no bias with respect to ethical attributes. In other words, it is clearly indicated that the content of the dataset used to train the human classification model is appropriate. The absence of bias with respect to ethical attributes may be displayed by the words "No Bias" as shown in FIG. 38 , or may be displayed by a mark or the like that guarantees the quality of the dataset used by the AI model for training.
[0322] In this example, by associating test result information corresponding to the test results of the dataset with the AI model, the quality of the dataset used to train the AI model can be guaranteed when the AI model is made public. It is also possible to prevent an AI model obtained by training a dataset that is biased with respect to ethical attributes from being made public. Note that test result information may be associated with the dataset. This allows the dataset to be treated in the same way as an AI model. For example, when a dataset is made public on an information processing device on the cloud side, it is also possible to prevent a dataset that is biased with respect to ethical attributes from being made public. The test result information may be associated with both the AI model and the dataset used to train the AI model.
[0323] Although multiple embodiments and examples of the present technology have been specifically described above, the content of the present technology is not limited to the above-described embodiments, etc., and various modifications based on the technical concept of the present technology are possible. Modifications will be described below.
[0324] The contents of the training dataset may be other than image data (such as audio data). In this case, for example, a model that automatically estimates emotions and environmental sounds (surrounding sounds) from audio data may be considered, and the emotions and environmental sounds may be considered as examples of attributes. Furthermore, the dataset may not be pre-annotated with "male" and "female."
[0325] For example, the configurations, methods, processes, shapes, materials, and numerical values of the above-described embodiments can be combined or substituted with each other without departing from the spirit of the present technology. Furthermore, one thing can be divided into two or more parts, and some parts can be omitted. Furthermore, the matters described in the embodiments and modified examples can be combined with each other. Furthermore, each of the above-described processes may be performed in a distributed manner by multiple information processing devices.
[0326] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0327] The present technology may also be configured as follows. (1) An information processing device comprising: an attribute extraction unit that extracts attributes from an input dataset; a labeling unit that labels the attributes for each piece of data constituting the dataset; a bias detection unit that detects bias for each attribute based on a result of the labeling; and a dataset adjustment unit that adjusts the dataset to improve at least one bias detected by the bias detection unit. (2) The information processing device according to (1), in which the dataset adjustment unit improves the bias by expanding the dataset. (3) The information processing device according to (2), in which the dataset adjustment unit expands the dataset by adding new data to the dataset. (4) The information processing device according to (2), in which the dataset adjustment unit expands the dataset by adding generated new data to the dataset. (5) The information processing device according to (4), in which the generated new data includes data based on a virtual viewpoint generated based on real data. (6) The information processing device according to (1), in which the dataset adjustment unit improves the bias by deleting some data constituting the dataset. (7) The information processing device according to any one of (1) to (6), further comprising an output unit. (8) The information processing device according to (7), wherein the output unit outputs display information that visualizes a result of labeling by the labeling unit. (9) The information processing device according to (8), wherein the display information includes display information in which the result of labeling is graphed, and data of a representative sample corresponding to at least one attribute. (10) The information processing device according to any one of (7) to (9), wherein the output unit outputs display information that visualizes information based on bias for each attribute detected by the bias detection unit. (11) The information processing device according to (10), wherein the information based on bias for each attribute includes at least one of a bias degree for each attribute and a bias degree for the entire dataset.(12) The information processing device according to any of (7) to (11), wherein the output unit outputs display information visualizing a result of the bias improvement by the dataset adjustment unit. (13) The information processing device according to any of (7) to (12), wherein the output unit outputs at least one dataset in which the bias has been improved. (14) The information processing device according to any of (1) to (13), further comprising a learning unit that performs machine learning based on the at least one dataset in which the bias has been improved. (15) The information processing device according to any of (1) to (14), wherein attributes labeled by the labeling unit include attributes specified by a user, the bias detection unit detects bias for at least the attribute specified by the user based on a result of the labeling, and the dataset adjustment unit adjusts the dataset to improve the bias for the attribute specified by the user detected by the bias detection unit. (16) The information processing device according to any of (1) to (15), wherein an AI model trained using the dataset is used to extract attributes by the attribute extraction unit, label the attributes by the labeling unit, detect bias for each attribute by the bias detection unit, and adjust the dataset by the dataset adjustment unit. (17) The information processing device according to any of (1) to (16), wherein the dataset adjustment unit adjusts the dataset to improve bias in ethical attributes. (18) The information processing device according to any of (1) to (17), wherein the dataset adjustment unit adjusts the dataset to improve bias specified by a user from among biases detected by the bias detection unit. (19) The information processing device according to any of (1) to (18), wherein the dataset is a dataset used for training a predetermined AI model, and the bias detection unit detects the presence or absence of bias in at least ethical attributes and associates detection result information corresponding to the detection result with at least one of the dataset and the AI model trained using the dataset.(20) The information processing device according to (7), wherein the output unit outputs display information that visualizes information of a first dataset and information of a second dataset, the information of the first dataset including at least one of a result of attribute labeling by the labeling unit for the first dataset and a result of bias detection for each attribute by the bias detection unit for the first dataset, and the information of the second dataset including at least one of a result of attribute labeling by the labeling unit for the second dataset and a result of bias detection for each attribute by the bias detection unit for the second dataset. (21) An information processing method, wherein an attribute extraction unit extracts attributes from an input dataset, a labeling unit labels the attributes for each piece of data constituting the dataset, a bias detection unit detects bias for each attribute based on the labeling result, and a dataset adjustment unit adjusts the dataset to improve at least one bias detected by the bias detection unit. (22) A program causing a computer to execute an information processing method, comprising: an attribute extraction unit extracting attributes from an input dataset; a labeling unit labeling the attributes for each piece of data constituting the dataset; a bias detection unit detecting bias for each attribute based on a result of the labeling; and a dataset adjustment unit adjusting the dataset to improve at least one bias detected by the bias detection unit.
[0328] DESCRIPTION OF SYMBOLS 1... Information processing device 2... Attribute extraction unit 3... Labeling unit 4... Bias detection unit 5... Data set adjustment unit 6... Output unit 6A... Display unit 100... Cloud server 200C... AI model developer terminal 500... Management server 1000... Information processing system
Claims
1. An information processing device comprising: an attribute extraction unit that extracts attributes from an input dataset; a labeling unit that labels the attributes for each piece of data that constitutes the dataset; a bias detection unit that detects bias for each attribute based on a result of the labeling; and a dataset adjustment unit that adjusts the dataset so as to improve at least one bias detected by the bias detection unit.
2. The information processing device according to claim 1, wherein the dataset adjustment unit improves the bias by expanding the dataset.
3. The information processing device according to claim 2, wherein the data set adjustment unit extends the data set by adding new data to the data set.
4. The information processing device according to claim 2, wherein the dataset adjustment unit extends the dataset by adding the generated new data to the dataset.
5. The information processing device according to claim 4, wherein the generated new data includes data based on a virtual viewpoint generated on the basis of real data.
6. The information processing device according to claim 1, wherein the dataset adjustment unit improves the bias by deleting a portion of data constituting the dataset.
7. The information processing device according to claim 1, further comprising an output section.
8. The information processing device according to claim 7, wherein the output unit outputs display information that visualizes a result of the labeling by the labeling unit.
9. The information processing device according to claim 8, wherein the display information includes display information in which the results of the labeling are graphed, and data of a representative sample corresponding to at least one attribute.
10. The information processing device according to claim 7, wherein the output unit outputs display information that visualizes information based on the bias for each attribute detected by the bias detection unit.
11. The information processing device according to claim 10, wherein the bias-based information for each attribute includes at least one of a bias degree for each attribute and a bias degree for the entire data set.
12. The information processing device according to claim 7, wherein the output unit outputs display information that visualizes a result of the bias improvement performed by the dataset adjustment unit.
13. The information processing device according to claim 7, wherein the output unit outputs at least one of the data sets in which the bias has been improved.
14. The information processing device according to claim 1, further comprising a learning unit that performs machine learning based on at least one of the data sets in which the bias has been improved.
15. The information processing device of claim 1, wherein the attributes labeled by the labeling unit include attributes specified by a user, the bias detection unit detects bias for at least the attributes specified by the user based on a result of the labeling, and the dataset adjustment unit adjusts the dataset so as to improve the bias for the attributes specified by the user detected by the bias detection unit.
16. The information processing device of claim 1, wherein an AI model trained using the dataset is used to extract attributes by the attribute extraction unit, label the attributes by the labeling unit, detect bias for each attribute by the bias detection unit, and adjust the dataset by the dataset adjustment unit.
17. The information processing device according to claim 1, wherein the dataset adjustment unit adjusts the dataset to improve bias in ethical attributes.
18. The information processing device according to claim 1, wherein the data set adjustment unit adjusts the data set so as to improve a bias designated by a user among the biases detected by the bias detection unit.
19. The information processing device of claim 1, wherein the dataset is a dataset used for training a predetermined AI model, and the bias detection unit detects the presence or absence of bias in at least an ethical attribute, and associates detection result information corresponding to the detection result with at least one of the dataset and the AI model trained using the dataset.
20. The information processing device of claim 7, wherein the output unit outputs display information that visualizes information of the first dataset and information of the second dataset, the information of the first dataset including at least one of a result of attribute labeling by the labeling unit for the first dataset and a result of bias detection for each attribute by the bias detection unit for the first dataset, and the information of the second dataset including at least one of a result of attribute labeling by the labeling unit for the second dataset and a result of bias detection for each attribute by the bias detection unit for the second dataset.
21. An information processing method, comprising: an attribute extraction unit extracting attributes from an input dataset; a labeling unit labeling the attributes for each data constituting the dataset; a bias detection unit detecting bias for each attribute based on a result of the labeling; and a dataset adjustment unit adjusting the dataset so as to improve at least one bias detected by the bias detection unit.
22. A program causing a computer to execute an information processing method, in which an attribute extraction unit extracts attributes from an input dataset; a labeling unit labels the attributes for each data constituting the dataset; a bias detection unit detects bias for each attribute based on the labeling results; and a dataset adjustment unit adjusts the dataset to improve at least one bias detected by the bias detection unit.
Citation Information
Patent Citations
Polyvinyl chloride resin composition
JP2020002255A
Training device, model generation method and program
JP2021189553A
Data generation device, method, and program
WO2023171335A1
Information processing device, information processing method, and computer program
WO2023188790A1