Scoring system, method, and program
The image scoring system uses AI to evaluate images by focusing on specific persons within the context of the entire image, addressing the limitations of existing systems by enhancing scoring accuracy through deep learning techniques.
Patent Information
- Application Number
- JP2024007795
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2044-01-23
AI Technical Summary
Existing image scoring systems primarily focus on individual elements of a person, such as facial expressions and face size, without considering the overall context of the image, limiting the accuracy of scoring.
An image scoring system that utilizes an AI model to evaluate images based on a specific person within the image, considering the entire image context, including interactions and background information, using deep learning techniques like CNN and RNN to determine an evaluation score.
Enhances the accuracy of image scoring by considering the overall context, allowing for better selection and classification of images based on the interaction and background information of the subject, improving user convenience.
Smart Images

Figure 2025113571000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image scoring technology.
Background Art
[0002] In kindergartens, schools, etc., photos of many children are taken, and daily reports, albums, etc. are created. For example, Patent Document 1 discloses an image candidate determination device that assists in determining which image should be selected in order to equalize as much as possible the number of images in which each person appears in the disclosed images.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, Patent Document 1 only discloses a technique for scoring based on the number of people included in an image, the expression of the face, etc., and for scoring each person, it is only performed based on individual elements for each person (the expression of each person, the size of the face, the opening of the eyes, etc.). Therefore, there is room for consideration regarding scoring based on the information of the entire image including elements other than the individual elements of each person.
[0005] The non-limiting embodiments in the present disclosure are made in view of the above background, and contribute to providing a technique for more appropriately scoring images including people.
Means for Solving the Problems
[0006] A scoring system according to an aspect of the present disclosure includes a reception unit that receives an image including at least one or more persons as a reception image, and acquires an image including one or more persons as a learning image, and focuses on a specific person included in the learning image. An evaluation score output unit that outputs an evaluation score for the received image using an image scoring model that has been trained to output an evaluation score for the learning image when paying attention to the person.
[0007] A method according to an aspect of the present disclosure includes a reception step of receiving, using a scoring system, an image including at least one or more persons as a reception image, and acquiring an image including one or more persons as a learning image, and focusing on a specific person included in the learning image. An evaluation score output step of outputting an evaluation score for the received image using an artificial intelligence model that has been trained to output an evaluation score for the learning image when paying attention to the person.
[0008] A computer program according to an aspect of the present disclosure is a program for causing a computer to function as the above-described scoring system, and causes the computer to function as each unit.
[0009] These general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium, or may be implemented by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
Advantages of the Invention
[0010] According to an aspect of the present disclosure, it is possible to provide a technique for more appropriately scoring an image including a person.
[0011] Further advantages and effects in an aspect of the present disclosure will be clarified from the specification and drawings. Such advantages and / or effects are provided by some embodiments and the features described in the specification and drawings, respectively, but not all of them are necessarily provided in order to obtain one or more identical features.
Brief Description of Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Modes for Carrying Out the Invention
[0013] Hereinafter, with appropriate reference to the drawings, one embodiment of the present disclosure will be described in detail. However, a more detailed description than necessary may be omitted. For example, a detailed description of already well-known matters or a redundant description of substantially the same configuration may be omitted. This is to avoid making the following description unnecessarily redundant and to facilitate the understanding of those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and it is not intended to limit the subject matter described in the claims by these.
[0014] 〔Embodiment 1〕 <<Overview of Scoring System 1>> First, an overview of the scoring system 1 according to this embodiment will be described. The scoring system 1 according to this embodiment is an information processing system that scores an evaluation score of an image. The scoring system 1 can be preferably introduced in facilities such as nurseries and kindergartens, for example. Also, the scoring system 1 may be used by users such as nursery teachers and teachers working in facilities such as nurseries and kindergartens. However, the scoring system 1 may also be used in various schools, nursing facilities, hospitals, commercial facilities, etc. other than nurseries and kindergartens, and the application scenarios are not particularly limited.
[0015] For example, in facilities such as nurseries and kindergartens, one or more cameras are installed, and images captured by the cameras are transmitted to a server and then transmitted from the server to a terminal device (for example, a smartphone or a personal computer, etc.) owned by a guardian of a child in the facility or the like. The camera installed in the facility may, for example, automatically repeat photographing inside the facility at a predetermined cycle, and the captured images show the children in the facility. The camera is connected to a network and can transmit the captured images to the server.
[0016] The server receives images from the camera via the network and stores and accumulates the received images in a storage device. Also, the server may transmit the accumulated images to a terminal device such as a smartphone or a personal computer associated with a person related to the facility (guardian of a child, nursery teacher, etc.). Thereby, for example, a guardian who has entrusted a child to a facility such as a nursery or a kindergarten can view the state of the child spending time in the facility on his or her own terminal device by means of the sent images.
[0017] Furthermore, for example, a childcare worker at a facility can obtain images sent from a camera installed at the facility and create a photo album showing the children spending time at the facility, post the photos, etc., without having to take photos of the children themselves. Note that the system is not particularly limited and may be configured such that the childcare worker takes photos of the children themselves using a terminal device they own and sends the photos to a server, where the images are stored in a storage unit.
[0018] For example, cameras installed in a facility may automatically capture thousands to hundreds of thousands of images per day, and it is not easy for users (childcare workers, parents, etc.) to visually review each of these images and select appropriate photos for each subject (e.g., children at a nursery school).
[0019] Therefore, the scoring system 1 of this embodiment may be applied, as an example, to an information processing system that performs appropriate and reasonable scoring on images, making it easy to select appropriate photos for each subject from a large number of images taken by cameras installed in facilities, or to assist in the creation of documents using photos (daily reports, albums, etc.).
[0020] <<Configuration of Scoring System 1>> The scoring system 1 according to this embodiment will be described in detail below with reference to the drawings. FIG. 1 is a block diagram showing an example of the functional configuration of the scoring system 1. The scoring system 1 is an information processing system configured by an information processing device. The scoring system 1 may be configured by one device or by multiple devices. Furthermore, when an information processing device is configured by multiple devices, the devices do not need to be installed in the same space such as the same room, but may be installed in different rooms, different buildings, different regions, etc., and are not particularly limited.
[0021] In FIG. 1, the scoring system 1 includes a server 10 and a user terminal 20. The server 10 is an example of the information processing apparatus in the present disclosure. The server 10 and the user terminal 20 are communicably connected via a network N1. The network N1 connecting the server 10 and the user terminal 20 is a wired LAN (Local Area Network), a wireless LAN, the Internet, a public switched telephone network, a mobile data communication network, or a combination thereof.
[0022] (Configuration of User Terminal 20) In FIG. 1, the user terminal 20 includes a storage unit 21 and a control unit 22. The control unit 22 acquires various types of information (input information) input by the user of the scoring system 1 via an input device. In some cases, input by the user via the input device may simply be described as input by the user.
[0023] Further, the control unit 22 presents to the user various types of information (output information) transmitted from the server 10 in response to the transmission of the information input by the user by displaying it on a display device. In some cases, presenting to the user by displaying it on the display device may simply be described as presenting to the user. Also, the storage unit 21 stores information transmitted and received between the server 10.
[0024] (Configuration of Server 10) The server 10 includes a storage unit 11, a control unit 12, and an artificial intelligence model M. The storage unit 11 stores information transmitted and received between the user terminal 20, as well as information transmitted and received between other devices communicably connected via the network N1.
[0025] The control unit 12 is composed of a reception unit 121, an evaluation score output unit 122, a specified information generation unit 123, a person image detection unit 124, a text generation unit 125, an attribute information output unit 126, an image cutting unit 127, a transmission unit 128, a first application processing unit 129, a second application processing unit 130, a third application processing unit 131, a first document generation unit 132, a second document generation unit 133, and a learning unit 134.
[0026] The artificial intelligence model M is composed of an image scoring model M1, a text generation model M2, and an attribute information output model M3.
[0027] Note that this is merely an example of the functional configuration of the scoring system 1 and does not necessarily need to have all the functional blocks. For example, if there are functions that do not need to be provided according to the user's needs, a configuration with only the necessary combinations of functional blocks may be sufficient.
[0028] Also, the configuration where all the functional blocks included in the control unit 12 are arranged in one server 10 is merely an example, and each functional block may be distributed and arranged in a plurality of information processing devices, and is not particularly limited. Furthermore, each functional block is not limited to a configuration realized by one integrated program, and may also be a configuration realized by a plurality of programs, and is not particularly limited.
[0029] Also, the configuration where all the models included in the artificial intelligence model M are arranged in one server 10 is merely an example, and each model may be distributed and arranged in a plurality of information processing devices, or may be a configuration that utilizes functions provided by an external device communicably connected via a network N1 or the like, and is not particularly limited. Furthermore, each model may be a configuration realized as one function of one integrated general artificial intelligence model, or may be a configuration realized as an independent learned model for each model, and is not particularly limited.
[0030] <<Scoring of Images>> The following describes the functions related to the scoring of images by the scoring system 1.
[0031] As an example in this embodiment, in the scoring system 1, the reception unit 121 receives an image including at least one or more persons as a received image, and the evaluation score output unit 122 outputs an evaluation score for the received image using the image scoring model M1.
[0032] (Reception Unit 121) For example, the user inputs an image including at least one or more persons to the scoring system 1 via the user terminal 20. Note that the image input to the scoring system 1 may be input as data of a JPG file, for example, but may also be data of an image file in other formats such as PNG or GIF, and the format of the image file is not particularly limited. The reception unit 121 receives the image input from the user terminal 20 as a received image. The received image received by the reception unit 121 may be stored in the storage unit 11.
[0033] Note that the configuration mode in which the reception unit 121 receives a received image is only an example, and it may be a configuration in which an image is acquired from another functional block included in the control unit 12, or alternatively, a configuration in which an image is acquired from another information processing device. Further, the reception unit 121 may be configured to receive an input of an image during batch processing executed at a predetermined timing for a large number of images stored in the storage unit 11, and is not limited to a configuration that receives a manual input by the user.
[0034] (Overview of Image Scoring Model M1) The image scoring model M1 is an artificial intelligence model (AI model) that is trained to acquire an image including one or more persons as a training image and output an evaluation score for the training image when paying attention to the image of a specific person included in the training image.
[0035] Here, an AI model is a model created by artificial intelligence learning input data. For example, from a huge number of input learning data sets, it learns the laws and relationships in the learning data sets, and through the optimization of parameters inside the model, etc., it can output an output corresponding to the learned content for the input. That is, an AI model is a mechanism for determining what output to return for the input data.
[0036] Note that as the learning algorithm for training the AI model, known learning algorithms can be applied. For example, as an algorithm for training the AI model, deep learning (CNN: Convolutional Neural Network, RNN: Recurrent Neural Network, LSTM: Long short-term memory, GAN: Generative Adversarial Network, etc.) can be used. The AI model can be trained not only by supervised learning, but also by unsupervised learning, semi-supervised learning, etc. Note that the learning model of AI does not necessarily have to be in the structure of a neural network. For example, it can be an SVM (Support Vector Machine), a decision tree, etc., and is not particularly limited.
[0037] Here, a brief supplement about CNN, etc. is as follows. In recent years, the use of deep learning (deep learning) as an artificial intelligence technology has been on the rise. Deep learning is an algorithm that deepens the layers of a neural network, which is one of the machine learning technologies. A neural network is an algorithm modeled after the brain nerves (neurons) of a living being.
[0038] For example, a neural network has input, hidden, and output layers. Each layer may have nodes, and the nodes may be connected by edges. In this case, the hidden layer can have multiple layers. Machine learning using this deep neural network with a deep hidden layer is called deep learning.
[0039] In other words, deep learning is a type of machine learning in which a computer automatically extracts features of input data using a neural network. Deep learning can be applied to both supervised learning and unsupervised learning. The neural network may be trained by methods such as gradient descent, stochastic gradient descent, or backpropagation.
[0040] For example, each layer of a neural network has a function called an activation function, and the edges can have weights. The value of each node is calculated from the values of the nodes in the previous layer. That is, the value of each node is calculated from the node values in the previous layer, the weight values of the edges, and the activation function. There are various such calculation methods, but the details are omitted here.
[0041] CNN, for example, has proven results in the field of image recognition. In CNN, the hidden layer is composed of a convolutional layer and a pooling layer. In the convolutional layer, filtering is performed on the nodes near those in the previous layer. It is an image of sliding a filter represented by a matrix over the original input image. Thereby, a feature map can be obtained.
[0042] The pooling layer reduces the feature map output from the convolutional layer to generate a new feature map. By this feature map generation process, image misalignment can be absorbed. And the pooling layer performs a process of summarizing local features.
[0043] That is, the convolutional layer and the pooling layer mean reducing the image while maintaining the features of the input image. The difference from mere image compression is that it reduces the image while maintaining the features of the image, or in other words, it can be said to be a process of abstraction.
[0044] In a CNN, each layer sequentially passes meaningful data to the next layer. As the layers progress, the network can learn higher-level features. For example, in the first layer, local features such as edges are detected, and in the next layer, they are combined to detect textures, and in the subsequent layer, features of more abstract parts such as ears are detected.
[0045] In a CNN, parameters for extracting such features are automatically learned. In this way, in a CNN, using the abstracted image images stored in the network, the input image can be recognized or classified.
[0046] While the data mainly handled by a CNN is image data (two-dimensional rectangular data), what an RNN handles is, for example, variable-length time-series data such as voice data. In an RNN, in order to handle variable-length data with a neural network, it has a network structure in which the value of the hidden layer is input back into the hidden layer.
[0047] However, an RNN has problems such as the error disappearing or the amount of computation becoming huge when using data from a long time ago. Therefore, it could only process short-term data. Then LSTM emerged. LSTM is a learning model that eliminates the drawbacks of an RNN and can learn long-term time-series data.
[0048] Also, as described above, it can be said that in a CNN, a feature extractor (or discriminator) that captures various features has been completed. In transfer learning, the part of the feature extractor that has already been completed in the CNN (the layer closer to the input side in the neural network) is reused, and only the part of the classifier (the layer closer to the output side in the neural network) is relearned. As a result, additional learning can be performed with less data than when learning normally.
[0049] For example, in image recognition, when a neural network for recognizing dogs and cats has been learned, by reusing the layer that identifies primitive elements and features closer to the input side and only relearning the layer that identifies more abstract elements and features closer to the output side and the layer that classifies the object, it is possible to learn to recognize other animals such as monkeys and chimpanzees. In this way, the amount of data can be supplemented by performing transfer learning that relearns the part of the layer closer to the output side.
[0050] (Learning unit 134) In the scoring system 1, for example, the learning unit 134 may be configured to automate the learning of the image scoring model M1. For example, the learning unit 134 may perform supervised learning of the image scoring model M1 using teacher data.
[0051] As teacher data, for example, a pair of a learning image including one or more persons and an evaluation score for the learning image when focusing on the image of a specific person included in the learning image can be considered. Note that in the scoring system 1, the image scoring model M1 may be manually learned, and is not particularly limited.
[0052] FIG. 2 is a diagram showing an example of a learning image. For example, the learning unit 134 may use, as an example of teacher data, a pair of the learning image shown in FIG. 2 and an evaluation score for the learning image to learn the image scoring model M1.
[0053] As shown in FIG. 2, the learning image P is an image taken by three children (kindergarten children A, B, and C) and one adult (teacher D). The image scoring model M1 may obtain, for example, the image shown in FIG. 2 from the learning unit 134.
[0054] Further, the image scoring model M1 obtains, for example, from the learning unit 134, together with the image P, a numerical value representing an evaluation score for the learning image when paying attention to the image of a specific person included in the image P.
[0055] Note that the image scoring model M1 is not limited to a configuration automatically learned by the learning unit 134, and may be a configuration in which a set of the image P and the above-described evaluation score is input to the scoring system 1 by the user instead of the learning unit 134.
[0056] Then, the image scoring model M1 repeatedly updates (that is, repeats learning) so that the parameters inside the model are optimized according to, for example, their laws and relationships, etc., from a huge number of sets of images and evaluation scores, thereby constructing a mechanism for determining what evaluation score output to return for the input image.
[0057] (Evaluation Score Output Unit 122) The evaluation score output unit 122 outputs an evaluation score for the received image using the learned image scoring model M1 described above. The evaluation score may be output as a numerical value within the range of 0 to 1, for example.
[0058] In addition, the evaluation score may be evaluated according to predefined subdivided grades such as other numerical ranges, A judgment, B judgment, etc., and is not particularly limited. As will be described later, the evaluation score may be used, for example, for image classification, sorting, selection, etc. Further, the evaluation score may be presented to the user via, for example, the display or speaker of the user terminal 20. Further, the evaluation score may be stored in the storage unit 11 together with information for identifying the received image to be evaluated (for example, received image ID, etc.).
[0059] When the evaluation score output unit 122 outputs an evaluation score, information indicating the person to be focused on when outputting the evaluation score among the persons included in the received image may be specified in advance by the user. Alternatively, the evaluation score output unit 122 may autonomously identify the image area of the person to be focused on when outputting the evaluation score among the persons included in the received image. For example, when outputting an evaluation score for a received image, the evaluation score output unit 122 uses an AI model or the like that detects persons included in the image to identify the image area of the person to be focused on, and then outputs an evaluation score for the received image.
[0060] In addition, when the evaluation score output unit 122 outputs an evaluation score for a received image, for the evaluation score to be output, information indicating which person among the persons included in the received image the evaluation score is for when focusing on that person (for example, the identification ID of the person, etc.) may be output simultaneously, and stored in the storage unit 11 in association with the received image ID and its evaluation score.
[0061] For example, the evaluation score output unit 122 may output a child ID for identifying the child focused on when outputting the evaluation score among the persons included in the image. The evaluation score output unit 122 extracts, for example, the facial feature amounts of the children included in the image, and compares them with the facial images of the children registered in the storage unit 11 or the feature amounts extracted from this image. Then, the evaluation score output unit 122 may search for the data of the children whose feature amounts that match or are similar to the facial feature amounts of the children included in the image are registered in the storage unit 11, and output the corresponding child ID.
[0062] The evaluation score output unit 122 may, for example, receive a face image as input and extract the feature amount of the face of the person included in the image using a learning model that outputs multi-dimensional vector information as the feature amount of this face. Also, when registering the feature amount of a person's face in the storage unit 11, the same learning model can be used. The evaluation score output unit 122 can, for example, calculate the distance or the like of multi-dimensional vectors corresponding to a plurality of feature amounts, and determine that the face features match or are similar when this distance is less than or equal to a threshold value and the distance is the smallest.
[0063] In this way, the evaluation score output by the evaluation score output unit 122 is the evaluation score for the received image when focusing on a specific person. In other words, this is the evaluation score grasped from the entire received image including the image of the part other than the image of the person being focused on, and it is a score representing an evaluation according to the situation in the received image of the person being focused on.
[0064] For example, as an example, the evaluation score output by the evaluation score output unit 122 will output different evaluation scores according to the communication situation (for example, talking with the teacher, listening to the teacher's speech, or looking around and not listening to the teacher's speech at all) between the person being focused on and the people other than the person being focused on among the people included in the received image.
[0065] For example, as will be described later, the evaluation score output unit 122 may output an evaluation score of a high score when the specific person being focused on is talking with the teacher, a medium score when listening to the teacher's speech, or a low score when looking around and not listening to the teacher's speech at all.
[0066] Conventionally, there have been image scoring methods. However, as evaluation items for image scoring, for example, regarding the people shown in a photo, elements such as having open eyes, having a smiling face, being shown in an up position, etc., are merely listed as individual elements related to the person themselves as evaluation items. To evaluate these evaluation items, only the elements of the image of the person themselves are analyzed, and the total value is calculated by, for example, weighting based on the scores of each element.
[0067] In contrast, in the scoring system 1 according to the present disclosure, as described above, not only the information related to the person themselves, but also information grasped from the entire image including other people than that person or the image of the background (such as being indoors in a garden, playing games outdoors, or participating in an event such as a field trip, etc.) is considered. When focusing on a specific person, an evaluation score for the received image will be output according to the information.
[0068] <<Designation of the person to be focused on in image scoring>> Hereinafter, in the image scoring by the scoring system 1, a configuration for designating the image area of the person to be focused on will be described.
[0069] (Designation information) In the scoring system 1, the reception unit 121 receives designation information for designating the image area of the person included in the received image. The reception unit 121 may receive the designation information and the received image simultaneously, or, for example, may first receive only the received image and then receive the designation information at another timing, and is not particularly limited.
[0070] The designation information may be, for example, information on the coordinates of four points representing the range of the image of the person included in the received image, or information representing the range of the image area of the person by the coordinates of the center point of the image area of the person and the vertical and horizontal distances from the center point.
[0071] Alternatively, in the reception image, for example, a rectangular figure surrounding the image area of the person of interest may be added, and the reception unit 121 may be configured to acquire the information of the rectangular figure as the designation information. That is, the reception unit 121 is not limited to a configuration that directly receives coordinate information as the designation information, and the designation information may be any information that can specify the range of the image area of the person of interest in the reception image, and is not particularly limited. Further, in the case of a configuration in which a figure surrounding the image area of a person is added to the reception image, the shape of the figure is not limited to a rectangle, and may be a curved figure such as an ellipse, for example.
[0072] Note that regarding the input of the designation information, the reception unit 121 may be configured to receive what is manually designated by the user, or alternatively, may be configured to receive what is generated by, for example, the designation information generation unit 123 described later, and is not particularly limited.
[0073] (Designation information for learning) Here, the image scoring model M1 has been learned to acquire the designation information for learning that designates the image area of the person included in the learning image and output an evaluation score for the learning image. That is, as described above, the image scoring model M1 learns to output an evaluation score for an image when focusing on a specific person, and acquires the designation information for learning that designates the image area of the specific person to be focused on during learning.
[0074] The designation information for learning may be any information that can specify the range of the image area of the person of interest in the learning image, similar to the above-described designation information, and may be information representing the range of the image in coordinates, or may be information of a figure added to surround the person in the learning image.
[0075] Note that the specified information received by the reception unit 121 is preferably the same type of specified information as the learning specified information used when the image scoring model M1 used for outputting the evaluation score is trained. For example, in a configuration where the reception unit 121 receives coordinate information as the specified information, it is preferable to use the image scoring model M1 that has acquired coordinate information as the learning specified information and has been trained.
[0076] Also, for example, in a configuration where the reception unit 121 receives information on a rectangular figure surrounding a person as the specified information, it is preferable to use the image scoring model M1 that has acquired information on a rectangular figure surrounding a person as the learning specified information and has been trained. However, for example, since the information on the rectangular figure surrounding a person can also be converted into coordinate information, the specified information received by the reception unit 121 does not necessarily have to be the same type of specified information as the learning specified information used when the image scoring model M1 used for outputting the evaluation score is trained, and is not particularly limited.
[0077] (An example of the training of the image scoring model M1) FIG. 3, FIG. 4, and FIG. 5 are diagrams for explaining an example of the training images. The image scoring model M1 acquires the training images P1, P2, and P3 shown in FIGS. 3, 4, and 5, and trains the evaluation scores for each image.
[0078] Note that in this embodiment, as an example, the evaluation score is represented by a numerical value in the range of 0 to 1 (the closer to 1, the higher the evaluation, and the closer to 0, the lower the evaluation). For the person being focused on, an evaluation score is given based on the surrounding situation grasped from the entire image (for example, the communication situation with teachers or friends, etc.). Note that as described above, the evaluation score is not limited to the configuration represented by a numerical value in the range of 0 to 1.
[0079] The learning image P1 shown in FIG. 3 is an image in which the image of child A is surrounded by a broken-line rectangular figure and focuses on child A. In image P1, although teacher D is speaking, it can be seen that child A is not facing teacher D and is not listening to the speech. Therefore, the image P1 focusing on child A is not an image in which a very good scene is depicted for child A, and thus has a low evaluation score. In this embodiment, the evaluation score of image P1 is set to "0.3". The image scoring model M1 obtains the evaluation score "0.3" for image P1 as part of the learning data, updates the parameters inside the model as described above, and learns about the scoring of the evaluation score.
[0080] The learning image P2 shown in FIG. 4 is an image in which the image of child B is surrounded by a broken-line rectangular figure and focuses on child B. In image P2, although child B is not actually talking to teacher D, it can be seen that child B is facing teacher D and listening attentively. Therefore, the image P2 focusing on child B can be said to be an image in which a relatively good scene is depicted for child B, and thus has a medium-level evaluation score. In this embodiment, the evaluation score of image P2 is set to "0.6". The image scoring model M1 obtains the evaluation score "0.6" for image P2 as part of the learning data, updates the parameters inside the model as described above, and learns about the scoring of the evaluation score.
[0081] The learning image P3 shown in FIG. 5 is an image in which the image of child C is surrounded by a broken-line rectangular figure and is an image focusing on child C. In image P3, it can be grasped that child C is talking to teacher D with gestures. Therefore, the image P3 focusing on child C can be said to be an image in which a very good scene is depicted for child C, and thus has a high evaluation score. In this embodiment, the evaluation score of image P3 is set to "0.9". The image scoring model M1 obtains the evaluation score "0.9" for image P3 as part of the learning data, updates the parameters inside the model as described above, and learns about the scoring of the evaluation score.
[0082] In addition to the above-mentioned images P1, P2, and P3, an enormous number of images are input into the image scoring model M1 as learning data, and the evaluation scores for each image when focusing on a specific person included in those images are also input. Then, the image scoring model M1 repeats learning using an enormous number of images about the scoring of the evaluation scores for images when focusing on a specific person, in the same way as the learning about the scoring using the above-mentioned images P1, P2, and P3, to improve the accuracy of scoring.
[0083] In the example of the learning of the image scoring model M1 described with reference to FIGS. 3, 4, and 5, an example was described in which the information of the rectangular figure surrounding the person images included in the obtained learning images P1, P2, and P3 is obtained as learning designation information and learned. However, as described above, for example, in the case where the scoring system 1 receives coordinate information as designation information and outputs an evaluation score, it may be configured to obtain coordinate information as learning designation information and learn.
[0084] By the way, the image scoring model M1 may be configured to obtain and learn the automatically generated learning designation information for each of a plurality of persons included in the learning image, for example.
[0085] In this case, for example, the learning unit 134 may automatically generate learning specification information for each person included in the image by automatically extracting the persons included in the image using a person image detection AI model or the like, and then focusing on each person one by one in order.
[0086] Then, the image scoring model M1, together with the learning specification information, acquires an evaluation score when focusing on the person specified by the learning specification information, and performs learning using these pairs.
[0087] Note that the learning specification information for designating the image area of the person included in the learning image may be configured to be automatically generated using a person image detection AI model or the like as described above, or may be configured to be manually designated by the user, and is not particularly limited.
[0088] In this way, since the image scoring model M1 has learned to acquire the learning specification information as part of the learning data and output an evaluation score of the image using this information, the evaluation score output unit 122 can output an evaluation score for the received image based on the specification information received by the reception unit 121 by using this image scoring model M1.
[0089] (Specification information generation unit 123) In the image scoring model M1, the specification information generation unit 123 may sequentially generate specification information for each of the images of a plurality of persons included in the received image, one person at a time.
[0090] The specification information generation unit 123 can detect, for example, a person image included in the received image using a person image detection AI model or the like. The specification information generation unit 123 may be configured to detect a person image using, for example, the function of a person image detection unit 124 described later, or may itself have the same function as the person image detection unit 124.
[0091] In addition to deep learning, various image recognition technologies such as SVM, decision trees, k-nearest neighbor methods, Gaussian mixture models, Boston matching, and computer vision may be used for detecting human images, and there is no particular limitation.
[0092] The human image detection AI model is, for example, pre-trained to receive image data as input and output, for the image area in which a person is depicted in this image, the coordinate information and the image surrounded by a rectangular shape. The human image detection AI model is trained, for example, using training data in which image data and data indicating the image area in which a person is depicted in this image are associated. Also, when the target person to be detected is a child based on an image taken at a facility such as a nursery or kindergarten, it is expected that the detection accuracy of the child can be improved by creating training data using an image in which the child is depicted.
[0093] Then, the designated information generation unit 123 generates designated information for each of the detected persons. For example, the designated information generation unit 123 may generate, as designated information, the coordinate information of four points representing the range of the image area of the person included in the received image.
[0094] Alternatively, the designated information generation unit 123 may be configured to add a rectangular shape surrounding the image of the person to the received image. As described above for the designated information, the information of the rectangular shape surrounding the person is a kind of designated information, and this configuration is also included in the configuration for generating designated information.
[0095] In this way, the designated information generation unit 123 automatically extracts the image area of the person included in the received image, focuses on each person included in the received image one by one in order, and automatically generates the designated information, so that the evaluation scores when focusing on specific persons for the received image can be automatically output sequentially and continuously.
[0096] As a result, when scoring the evaluation score when focusing on a specific person in the images of photos containing a large number of people captured daily, there is no need for the user to manually input the specified information one by one, which significantly improves convenience.
[0097] (Person Image Detection Unit 124) In the scoring system 1, the person image detection unit 124 may extract the feature amounts of the received image and detect the image of a person. The specified information for designating the image area of the detected person may be generated by the person image detection unit 124, or may be generated by the above-described specified information generation unit 123. Then, the reception unit 121 receives the specified information for designating the image area of the person detected by the person image detection unit 124.
[0098] For example, the person image detection unit 124 uses a person image detection AI model equipped with a feature amount extractor to detect the image of a person. This person image detection AI model is a learned model for detecting a person image.
[0099] The feature amount extractor extracts the feature amounts that characterize the person image from the image. The feature amounts are represented by, for example, multi-dimensional feature amount vectors. Further, the person image detection AI model may include a discriminator and may discriminate the person image according to the feature amount vector extracted by the feature amount extractor.
[0100] For example, the feature amount extractor constituting the person image detection unit 124 converts the person image into a multi-dimensional vector with a label for identifying the person (or person attributes, etc.). Then, a large number of images in which the person appears are input to the feature amount extractor as learning data, and are converted from the person image into a multi-dimensional vector with a label for identifying the person or person attributes, etc. and arranged in the multi-dimensional vector space. Further, the person image detection AI model may be constructed by repeatedly learning so as to create a boundary surface in the multi-dimensional vector space that enables the discriminator to discriminate each person.
[0101] Note that the person image detection unit 124 is not limited to the above-described configuration. For example, it may detect a person image based on, for example, the distance of feature vectors, etc., or may be a configuration using other image recognition techniques based on feature amounts, and is not particularly limited.
[0102] <<Image Classification / Sorting / Selection Based on Score>> (First Application Processing Unit 129) In the scoring system 1, the first application processing unit 129 may perform at least one or more of the processes of classification, sorting, or selection on the received image based on a predetermined condition regarding the evaluation score output by the evaluation score output unit 122.
[0103] For example, when a condition is set such that, for a received image, in the case of focusing on child A, it is classified into three types: high evaluation (evaluation score greater than 0.7), medium evaluation (0.3 to 0.7), and low evaluation (less than 0.3) according to the evaluation score, the first application processing unit 129 performs a classification process of the received image so as to meet this condition.
[0104] Also, for example, when a condition is set such that, for a received image, in the case of focusing on child A, it is sorted in descending order of the evaluation score, the first application processing unit 129 performs a sorting process of the received image so as to meet this condition.
[0105] Also, for example, when a condition is set such that, among the evaluation scores for the received image in the case of focusing on child A, those with high evaluation (for example, 0.7 or more, etc.) are selected, the first application processing unit 129 performs a selection process of the received image that meets this condition.
[0106] As a result, the user can check the list of image files sorted in descending order of evaluation. For example, when it is desired to preferentially send photos of highly evaluated images to, for example, guardians of kindergarten children, the desired image file can be easily searched for. Note that the predetermined conditions may be set as a combination of a plurality of conditions, such as combining sorting conditions and selection conditions, and are not particularly limited.
[0107] Note that the evaluation score may be stored in the storage unit 11 in association with the received image (or the image ID for identifying the received image) and the information for identifying the person (person identification ID) that was focused on when outputting the evaluation score. In this case, different image identification IDs may be set for the same received image according to the person identification ID. Thereby, the first application processing unit 129 can classify, sort, or select the received images stored in the storage unit 11 according to predetermined conditions regarding the evaluation score.
[0108] <<Text generation>> (Text generation unit 125) In the scoring system 1, the text generation unit 125 may generate text explaining the content of the image included in the received image. For example, the text generation unit 125 uses the learned text generation model M2 to generate text explaining the content of the image included in the received image so as to generate text explaining the content of the image when focusing on a specific person included in the learning image.
[0109] For example, in the case where a received image includes an image of kindergarten child A and an image of kindergarten child B, it is assumed that kindergarten child A is shown facing away from the teacher and talking to a friend, and kindergarten child B is shown facing the teacher and listening to the teacher's speech.
[0110] In this case, the text generation unit 125 may generate, for example, a text such as "Child A who is talking to a friend and not listening to the teacher" as the text describing the content of the image when focusing on Child A. Also, the text generation unit 125 may generate, for example, a text such as "Child B who is listening attentively to the teacher" as the text describing the content of the image when focusing on Child B.
[0111] The text generation model M2 is, for example, pre-trained to receive image data as input and output text information that describes the overall content of the image when focusing on the person shown in the image area specified by the learning specification information or the like. The text generation model may be trained, for example, using a large amount of training data in which image data and learning specification information are associated with the text representing its content.
[0112] The text generation model may be configured as, for example, a Visual Language Model (VLM). The Visual Language Model is a type of artificial intelligence, similar to the large language model. The large language model statistically learns the probability distribution of words and sentences from a vast amount of text data and can generate natural language output like a human for a given input (such as a prompt).
[0113] The large language model is also a type of deep learning model and is based on a computational model that mimics the function of neurons in the human brain called the neural network described above. The neural network has a structure consisting of multiple layers, receives a vast amount of input data, and adjusts its parameters by learning patterns.
[0114] Large language models are composed of huge neural networks consisting of an enormous number of parameters and are models trained using a large amount of text data. As a result, large language models have language knowledge such as grammar, meaning, and context internally and can generate appropriate outputs for given inputs.
[0115] Similar to large language models, vision-language models can also perform pre-training using a large amount of text data and image data pairs, and for a given image input, can generate an appropriate text output that describes the content.
[0116] (Second Application Processing Unit 130) In the scoring system 1, the second application processing unit 130 may perform at least one of classification, sorting, or selection processing on the received image based on a predetermined condition related to at least one of the evaluation score output by the evaluation score output unit 122 and the text generated by the text generation unit 125.
[0117] Classification, sorting, and selection based on the evaluation score have the same functions as the first application processing unit 129 described above, and the description thereof is omitted. The second application processing unit can further perform processing according to conditions related to the text generated by the text generation unit 125. For example, the second application processing unit 130 may select a received image for which a text including positive words such as "smiling face", "happy", and "energetic" is generated. Also, the second application processing unit 130 may classify the received images into those for which a text including negative words is generated and those for which a text including positive words is generated. Furthermore, the second application processing unit 130 may sort the received images in alphabetical order, for example, based on the generated text or the main words included in the text.
[0118] As a result, the user can easily retrieve a set of image files in which text containing positive words is generated (i.e., image files with high priority as candidates for posting in a daily report to be reported to a guardian or a graduation album). Note that the predetermined conditions may be set as a combination of a plurality of conditions, such as combining conditions related to text and conditions related to an evaluation score, and are not particularly limited.
[0119] Note that the evaluation score may be stored in the storage unit 11 in association with the received image (or the image ID for identifying the received image), information for identifying the person focused on when outputting the evaluation score (person identification ID), and the text generated by the text generation unit 125. In this case, different image identification IDs may be set according to the person identification ID for the same received image. Thereby, the second application processing unit 130 can classify, sort, or select the received images stored in the storage unit 11 according to predetermined conditions related to the evaluation score and the text.
[0120] <<Output of Attribute Values>> (Attribute Information Output Unit 126) In the scoring system 1, the attribute information output unit 126 may output the attribute information of the person in the image focused on when outputting the evaluation score. The attribute information output unit 126 outputs the attribute information using the learned attribute information output model M3 so as to output the attribute information related to the attributes of the person included in the image. Further, the attribute information output unit 126 may, for example, use the specified information to extract the image of the person included in the received image and output the attribute information of the person using the attribute information output model M3.
[0121] The attribute information output model M3 may be configured as an AI model that outputs, as attribute information, various attributes of a person (e.g., gender, hair length, length, color, and type of each part of clothing, presence or absence of a hat, presence or absence of a backpack or bag, facial expression, etc.) themselves when an image of the person is input, or alternatively, may be configured as an AI model that outputs, as attribute information, the confidence level (an index value representing the degree of confidence numerically, etc.) for various attributes of the person.
[0122] That is, the attribute information output model M3 can be configured such that, for example, it receives a person's image as input and has been pre-trained to output information representing the attributes themselves or the confidence level for the attributes regarding various attributes of the person depicted in the image (e.g., gender, hair length, length, color, and type of each part of clothing, presence or absence of a hat, presence or absence of a backpack or bag, etc.).
[0123] The attribute information output model M3 may be trained using, for example, learning data (a set of input-output data) in which a person's image is associated with label information (e.g., "1" when corresponding to an attribute item and "0" when not) representing the correct answers for various attributes of the person depicted in the image.
[0124] (Third Application Processing Unit 131) In the scoring system 1, the third application processing unit 131 may perform at least one of classification, sorting, or selection processing on the received image based on a predetermined condition related to at least either the evaluation score output by the evaluation score output unit 122 or the attribute information output by the attribute information output unit 126.
[0125] Regarding classification, sorting, and selection based on the evaluation score, it has the same function as the above-described first application processing unit 129, and the description thereof will be omitted. The third application processing unit can further perform processing according to conditions related to the attribute information output by the attribute information output unit 126. For example, the third application processing unit 131 may select, classify, or sort the received images according to the attributes of male or female. Alternatively, the third application processing unit 131 may select, classify, or sort the received images according to the attributes of smiling faces or crying faces.
[0126] As a result, the user can easily extract a set of desired image files (that is, image files with high priority as candidates for publication in, for example, daily reports to be reported to guardians or graduation albums). Note that the predetermined conditions may be set as a combination of a plurality of conditions, such as a combination of conditions related to attributes and conditions related to evaluation scores, and are not particularly limited.
[0127] Note that the evaluation score may be stored in the storage unit 11 in association with the received image (or the image ID for identifying the received image), the information for identifying the person focused on when outputting the evaluation score (person identification ID), and the attribute information output by the attribute information output unit 126. In this case, different image identification IDs may be set for the same received image according to the person identification ID. As a result, the third application processing unit 131 can classify, sort, or select the received images stored in the storage unit 11 according to predetermined conditions related to the evaluation score and the attribute information.
[0128] <<Document Creation>> (First Document Generation Unit 132) In the scoring system 1, the first document generation unit 132 may generate a predetermined document using the received images on which at least one of the processes of selection, classification, or sorting has been performed by the first application processing unit 129, the second application processing unit 130, or the third application processing unit 131. As the predetermined document, for example, in addition to the daily report sent to the guardians of the kindergarten children every day and the graduation album itself, it may be document data serving as a draft for creating them.
[0129] For example, the first document generation unit 132 may select the received images with a high evaluation score (for example, a score of 0.8 or higher) for each kindergarten child, and generate a document by inserting the photos of the selected received images into a format such as a predetermined daily report. Furthermore, the first document generation unit 132 may generate a document using, in addition to the selected, classified, or sorted received images, the text generated by the text generation unit 125 and the attribute information output by the attribute information output unit 126.
[0130] Also, the first document generation unit 132 may generate a document using a large language model. In this case, for example, the first document generation unit 132 generates a prompt according to a predetermined format based on the text generated by the text generation unit 125, the attribute information output by the attribute information output unit 126, etc., inputs the prompt into the large language model, and may cause the large language model to generate a document such as a daily report.
[0131] <<Trimming of Images>> (Image Cropping Unit 127) In the scoring system 1, the image cropping unit 127 may crop the image including the person who was focused on when the evaluation score was output based on the text generated by the text generation unit 125 and the evaluation score.
[0132] As described above, the text generation unit 125 generates text that describes the content of the image included in the received image. Since the reason for the high or low evaluation score can be grasped from this text, the image cropping unit 127 may crop an image in a range corresponding to the content of the text. For example, for an image with a high evaluation score, if text with the content "playing happily with friends" is generated, the image cropping unit 127 may crop an image in a range where the state of playing with friends can be grasped.
[0133] In this case, the image cropping unit 127 may be configured as, for example, an image cropping AI model, and receives as input an image including a person, information specifying the image area of the person of interest, and text information regarding the person, and is pre-trained to crop and output an image in a range where the content of the text information can be grasped for the person of interest.
[0134] For example, when an image including a person, information specifying the image area of the person, and text information "playing with friends" are input to the image cropping AI model, an image in a range where it can be grasped that the person is playing with friends (for example, an image in a range including the friends playing together) is cropped and output.
[0135] For example, the image cropping AI model may be trained using learning data (a set of input / output data) in which an image in a range where the content of the text can be grasped including the person is cropped from the image, together with the image data of the image, information specifying the image area of the person of interest, and the text information.
[0136] (Second Document Generation Unit 133) In the scoring system 1, the second document generation unit 133 may generate a document using the image of the person cropped by the image cropping unit 127. Examples of the generated document include a daily report sent to the guardians of kindergarten children every day, the graduation album itself, or document data that is a draft for creating them.
[0137] The image (cropped image) cropped by the image cropping unit 127 is stored in the storage unit 11 in association with the text generated by the text generation unit 125. The second document generation unit 133 may generate a document, for example, by selecting a cropped image that matches the information of the text specified by the user and inserting the cropped image into a predetermined format such as a daily report.
[0138] For example, when the user creates a daily report of a certain kindergarten child and inputs information instructing the scoring system 1 to generate a draft of a daily report using an image in which the text "playing with friends" is generated, the second document generation unit 133 reads out from the storage unit 11 a cropped image corresponding to the content of the text "playing with friends" and generates a document such as a daily report with a photo using the read cropped image.
[0139] (Transmission unit 128) The transmission unit 128 may be configured as, for example, a messaging application. That is, the server 10 is connected to the Internet via the network N1, and the transmission unit 10 can perform data communication with a mobile terminal owned by a guardian by using the function of the messaging application.
[0140] In the scoring system 1, the image of the person cropped by the image cropping unit 127 may be transmitted to an information processing terminal owned by a person related to the person. For example, the transmission unit 128 may transmit the image to a mobile terminal or the like owned by a guardian of the kindergarten child who is a person related to the kindergarten child shown in the cropped image or a guardian of a friend shown together.
[0141] <<An example of the processing flow in the scoring system 1>> An example of the operation of the scoring system 1 configured as described above will be described with reference to FIG. 6. FIG. 6 is a flowchart for explaining an example of the operation of the scoring system 1. Note that the flowchart shown in FIG. 6 is merely an example of the processing flow. For example, according to a user's request, other steps may be included, the same step may be repeatedly executed, or the order of the steps may be changed or some steps may not be executed, and there is no particular limitation.
[0142] In step S101, as an example, the control unit 22 of the user terminal 20 transmits an image input instruction from the user to the server 10. The image may be configured to be uploaded from the user terminal 20 to the server 10, or a large number of images stored in the storage unit 11 may be input by batch processing or the like. For example, in the case of input by batch processing, when the user performs an operation to instruct the execution of batch processing via the user terminal 20, an execution command for batch processing is transmitted to the server 10.
[0143] In step S102, as an example, the reception unit 121 of the server 10 receives an image including a person as a received image in response to the image input instruction from the user terminal 20. Regarding the specified information, as described above in detail, it may be configured to separately specify coordinate information, or may be included as a rectangular figure surrounding a person in the received image.
[0144] In step S103, as an example, the evaluation score output unit 122 of the server 10 outputs an evaluation score when paying attention to a specific person for the received image received by the reception unit 121. The evaluation score output unit 122 transmits the output evaluation score to the user terminal 20. For example, the evaluation score output unit 122 may individually output an evaluation score when paying attention to all the children shown in each image for all the images, and store it in the storage unit 11 together with the child ID and the image ID.
[0145] In step S104, as an example, the control unit 22 of the user terminal 20 presents the user with an evaluation score for the received acceptance image. The user can check the presented evaluation score.
[0146] In step S105, as an example, the control unit 22 of the user terminal 20 receives an instruction from the user to select an acceptance image, and transmits information representing the instruction content to the server 10. For example, when the evaluation score range is set in three levels: A (high evaluation), B (medium evaluation), and C (low evaluation), for all children, an instruction to select an image with an evaluation score range of A is sent from the user.
[0147] In step SIO6, as an example, in the server 10, the first application processing unit 129 selects an acceptance image that meets the conditions specified by the user regarding the evaluation score. Note that the second application processing unit 130 may select an acceptance image that meets the conditions specified by the user regarding the text information indicating the image content in addition to the conditions regarding the evaluation score. Also, the third application processing unit 131 may select an acceptance image that meets the conditions specified by the user regarding the personal attribute information in addition to the conditions regarding the evaluation score.
[0148] In step S107, as an example, the control unit 22 of the user terminal 20 presents the user with the acceptance image selected by the first application processing unit 129 (or the second application processing unit 130, the third application processing unit 131). The user checks the presented acceptance image.
[0149] In step S108, as an example, the control unit 22 of the user terminal 20 transmits information instructing the generation of a draft of a document (such as a daily report for each child to be sent to the guardian) input by the user to the server 10.
[0150] In step S109, as an example, the first document generation unit 132 (or the second document generation unit 133) of the server 10 generates a draft of a document (such as a daily report for each child to be sent to the guardian).
[0151] In step S110, as an example, the control unit 22 of the user terminal 20 presents a draft of the received document (such as a daily report for each child to be sent to the guardian) to the user. The user checks the presented draft of the document.
[0152] In step S111, as an example, the user may, if necessary, correct the draft of the document (such as the daily report to be sent to the guardian), and the control unit 22 of the user terminal 20 transmits a document reflecting the correction content input by the user to the server 10. At the server 10, the corrected document is stored in the storage unit 11. Thereafter, for example, the document may be distributed to an information terminal such as the guardian's smartphone at an appropriate timing by the transmission unit 128.
[0153] Thus, the scoring system 1 ends its operation. As described above, the flowchart shown in FIG. 6 only explains an example of the operation of the scoring system 1, and the operation of the scoring system 1 is not limited thereto.
[0154] [Embodiment 2] The configuration described in Embodiment 1 is merely an example, and it is not necessarily required that the server side has all of the above-described functional blocks. For example, a configuration in which the client side has all of the above-described functional blocks may be used.
[0155] Also, for example, the server side may have a part of the above-described functional blocks, and the client side may have the remaining functional blocks. The arrangement of the functional blocks on the server side and the client side can be arbitrarily combined according to the specifications required in each use case, and is not particularly limited.
[0156] Hereinafter, the scoring system 1A according to this embodiment will be described in more detail. For the sake of convenience of explanation, members having the same functions as the members described in the above embodiment are denoted by the same reference numerals, and the description thereof will not be repeated.
[0157] <<Configuration of Scoring System 1A>> FIG. 7 is a block diagram showing an example of the functional configuration of the scoring system 1A. As shown in FIG. 7, the scoring system 1A is different from the scoring system 1 according to Embodiment 1 in that it includes a user terminal 20A instead of the user terminal 20 and includes a server 10A instead of the server 10.
[0158] The server 10A is different from the server 10 in that it includes a control unit 12A instead of the control unit 12, includes a storage unit 11A instead of the storage unit 11, and further does not include the artificial intelligence model M.
[0159] The user terminal 20A is different from the user terminal 20 in that it includes a storage unit 21A instead of the storage unit 21, includes a control unit 22A instead of the control unit 22, and further includes an artificial intelligence model M' and an imaging unit 23A.
[0160] The control unit 22A is configured to include a reception unit 221, an evaluation score output unit 222, a designated information generation unit 223, a person image detection unit 224, a text generation unit 225, an attribute information output unit 226, an image cutting unit 227, a transmission unit 228, a first application processing unit 229, a second application processing unit 230, a third application processing unit 231, a first document generation unit 232, a second document generation unit 233, a learning unit 234, an alert unit 235, and a storage processing unit 236.
[0161] The artificial intelligence model M' is configured to include an image scoring model M1', a text generation model M2', and an attribute information output model M3'.
[0162] Here, the reception unit 221 is the same as the reception unit 121, the evaluation score output unit 222 is the same as the evaluation score output unit 122, the specified information generation unit 223 is the same as the specified information generation unit 123, the person image detection unit 224 is the same as the person image detection unit 124, the text generation unit 225 is the same as the text generation unit 125, the attribute information output unit 226 is the same as the attribute information output unit 126, the image cutting unit 227 is the same as the image cutting unit 127, the transmission unit 228 is the same as the transmission unit 128, the first application processing unit 229 is the same as the first application processing unit 129, the second application processing unit 230 is the same as the second application processing unit 130, the third application processing unit 231 is the same as the third application processing unit 131, the first document generation unit 232 is the same as the first document generation unit 132, the second document generation unit 233 is the same as the second document generation unit 133, the learning unit 234 is the same as the learning unit 134, the image scoring model M1' has the same function as the image scoring model M1, the text generation model M2' has the same function as the text generation model M2, and the attribute information output model M3' has the same function as the attribute information output model M3.
[0163] That is, in this embodiment, each functional block included in the control unit 12 on the server side in Embodiment 1 is included in the control unit 22A on the client side. However, the only difference is that the information processing of each functional block is executed in the information processing device on the client side, and the functions of each functional block are the same as those described in Embodiment 1. The alert unit 235 and the storage processing unit 236 will be described later.
[0164] <<An example of the processing flow in the scoring system 1A>> In addition, in this embodiment, the user terminal 20A is provided with an imaging unit 23A. An example of a case where a user (for example, a teacher in a nursery school) takes pictures of children in the nursery school using the user terminal 20A by himself / herself will be described with reference to FIGS. 8 and 9.
[0165] Note that the flowcharts shown in FIGS. 8 and 9 are merely examples of the processing flow. For example, depending on the user's request, other steps may be included, the same step may be repeatedly executed, or the order of steps may be changed or some steps may not be executed, without any particular limitation.
[0166] FIG. 8 is a flowchart for explaining an example of the operation of the scoring system 1A. In step S201, as an example, in response to a user operation, the imaging unit 23A of the user terminal 20A captures an image including a person (for example, an image showing the state of children in a nursery school), and the image is input to the scoring system 1A.
[0167] In step S202, as an example, the reception unit 221 of the user terminal 20A receives the image including the person captured by the imaging unit 23A as a received image. Regarding the designation information, as described in the above embodiment, for example, it may be generated by the designation information generation unit 223 and included as a rectangular figure surrounding the person in the received image.
[0168] In step S203, as an example, the evaluation score output unit 222 of the user terminal 20A outputs an evaluation score when focusing on a specific person for the received image received by the reception unit 221. For example, the evaluation score output unit 222 may individually output the evaluation scores when focusing on all the children shown in each of the captured images, and store them in the storage unit 21A together with the child ID and the image ID.
[0169] In step S204, as an example, the control unit 22A of the user terminal 20A presents the evaluation score for the received image to the user. The user can check the presented evaluation score.
[0170] At this time, for example, when the evaluation score is lower than a predetermined score, the alert unit 235 may present an alert to the user. That is, an evaluation score (alert reference score) serving as a reference for presenting a predetermined alert is stored in the storage unit 21A. The alert unit 235 determines whether the evaluation score output by the evaluation score output unit 222 is lower than the alert reference score. If it is lower, the alert unit 235 outputs an alert indicating that the evaluation score is lower than the alert reference score. As a result, the user can re-take a photo of the target child upon receiving this alert.
[0171] FIG. 9 is a flowchart for explaining an example of the operation of the scoring system 1A. Steps S301, S302, and S303 are respectively the same processing steps as steps S201, S202, and S203 shown in FIG. 8, and the description thereof is omitted.
[0172] In step S304, as an example, the control unit 22A of the user terminal 20A presents the evaluation score for the received image to the user. The user can check the presented evaluation score.
[0173] In step S305, as an example, the control unit 22A of the user terminal 20A receives an instruction from the user to select a received image. For example, when the evaluation score range is set in three levels: A (high evaluation), B (medium evaluation), and C (low evaluation), for all children in the kindergarten, an instruction to select an image with an evaluation score range of A is input by the user.
[0174] In step S306, as an example, in the user terminal 20A, the first application processing unit 229 selects a received image that meets the conditions specified by the user regarding the evaluation score. Note that the second application processing unit 230 may select a received image that meets the conditions specified by the user regarding the text information indicating the image content in addition to the conditions regarding the evaluation score. Further, the third application processing unit 231 may select a received image that meets the conditions specified by the user regarding the attribute information of the person in addition to the conditions regarding the evaluation score.
[0175] In step S307, as an example, the control unit 22A of the user terminal 20A presents the received image selected by the first application processing unit 229 (or the second application processing unit 230, the third application processing unit 231) to the user. The user checks the presented received image.
[0176] In step S308, as an example, the control unit 22A of the user terminal 20A receives an instruction to save the received image input from the user. For example, an instruction is input from the user to save only the received images with an evaluation score of A (high evaluation).
[0177] In step S309, as an example, when the save processing unit 236 receives an instruction to save only the received images with an evaluation score of A (high evaluation) from the user, it transmits only the received images with an evaluation score of A (high evaluation) that correspond to the content of the instruction to the server 10A.
[0178] In step S310, as an example, when the control unit 12A of the server 10A receives the received image transmitted from the user terminal 20A by the save processing unit 236, it stores it in the storage unit 11A. In this way, by adopting a configuration in which only the received images that meet the predetermined conditions are transmitted to and saved in the server 10A, the storage capacity of the storage such as the storage unit 11A can be saved.
[0179] Note that the storage processing unit 236 may be configured to store a received image that satisfies a predetermined condition in the storage unit 21A of the user terminal 20A. Further, for example, conditions for an image to be stored in advance may be stored in the storage unit 21A, and the storage processing unit 236 may read out the conditions for the image to be stored from the storage unit 21 and sequentially store only the received images that meet the conditions in the storage unit 11A or the storage unit 21A. There is no particular limitation.
[0180] 〔Example of Realization by Software〕 The control block of the server 10 may be realized by a logic circuit (hardware) formed in an integrated circuit (IC chip) or the like, or may be realized by software. In the latter case, each of the server 10 and the user terminal 20 is configured using, for example, a computer (electronic computer).
[0181] (Physical Configuration of Server 10) FIG. 10 is a block diagram illustrating the physical configuration of a computer used as the server 10 and the user terminal 20. Note that the server 10A has the same physical configuration as the server 10, and the user terminal 20A has the same physical configuration as the user terminal 20.
[0182] As shown in FIG. 10, the server 10 can be configured by a computer including a bus 110, a processor 101, a main memory 102, an auxiliary memory 103, and a communication interface 104. The processor 101, the main memory 102, the auxiliary memory 103, and the communication interface 104 are connected to each other via the bus 110.
[0183] As the processor 101, for example, a CPU (Central Processing Unit), a microprocessor, a digital signal processor, a microcontroller, or a combination thereof is used.
[0184] As the main memory 102, for example, a semiconductor RAM (random access memory) or the like is used.
[0185] As the auxiliary memory 103, for example, a flash memory, an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a combination thereof is used. The auxiliary memory 103 stores a program for causing the processor 101 to execute the operations of the server 10 described above. The processor 101 expands the program stored in the auxiliary memory 103 onto the main memory 102 and executes each instruction included in the expanded program.
[0186] The communication interface 104 is an interface for connecting to the network N1.
[0187] In this example, the processor 101 and the communication interface 104 are an example of hardware elements that implement the control unit 12. Also, the main memory 102 and the auxiliary memory 103 are an example of hardware elements that implement the storage unit 11.
[0188] (Physical Configuration of User Terminal 20) As shown in FIG. 10, the user terminal 20 can be configured by a computer including a bus 210, a processor 201, a main memory 202, an auxiliary memory 203, a communication interface 204, and an input / output interface 205. The processor 201, the main memory 202, the auxiliary memory 203, the communication interface 204, and the input / output interface 205 are connected to each other via the bus 210. An input device 206 and an output device 207 are connected to the input / output interface 205.
[0189] As the processor 201, for example, a CPU, a microprocessor, a digital signal processor, a microcontroller, or a combination thereof is used.
[0190] As the main memory 202, for example, a semiconductor RAM or the like is used.
[0191] As the auxiliary memory 203, for example, a flash memory, an HDD, an SSD, or a combination thereof is used. The auxiliary memory 203 stores a program for operating the computer as the user terminal 20. The processor 201 expands the program stored in the auxiliary memory 203 onto the main memory 202 and executes each instruction included in the expanded program. Also, the auxiliary memory 203 stores various data that the processor 201 refers to for operating the computer as the user terminal 20.
[0192] The communication interface 204 is an interface for connecting to a network.
[0193] As the input / output interface 205, for example, a USB interface, a short-range communication interface such as infrared or Bluetooth (registered trademark), or a combination thereof is used.
[0194] As the input device 206, for example, a keyboard, a mouse, a touch pad, a microphone, or a combination thereof is used. As the output device 207, for example, a display, a printer, a speaker, or a combination thereof is used.
[0195] In this example, the processor 201 and the communication interface 204 are an example of hardware elements that implement the control unit 22. Also, the main memory 202 and the auxiliary memory 203 are an example of hardware elements that implement the storage unit 21.
[0196] Note that instead of being stored in the auxiliary memories 103 and 203 respectively, each of the above-described programs may be recorded on an external recording medium and supplied to the corresponding computer by being read from the external recording medium. As the external recording medium, a "non-transitory tangible medium" readable by a computer, such as a tape, a disk, a card, a semiconductor memory, a programmable logic circuit, etc. can be used. Also, each of the above-described programs may be supplied to a computer via any transmittable transmission medium (such as a communication network or a broadcast wave). Further, one aspect of the present invention can also be realized in the form of a data signal embedded in a carrier wave, in which each program is embodied by electronic transmission.
[0197] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
[0198] <<Summary of the Present Disclosure>> Hereinafter, the operation and effects of the scoring system according to one aspect of the present disclosure will be mainly described. All the configurations described below can also be used as the configurations of the present embodiment.
[0199] A scoring system according to one aspect of the present disclosure includes a reception unit that receives an image including at least one or more persons as a reception image, and acquires an image including one or more persons as a learning image, and uses an image scoring model that has been trained to output an evaluation score for the learning image when paying attention to a specific person included in the learning image, and outputs an evaluation score for the reception image.
[0200] According to the above configuration, an evaluation score for the received image is output when paying attention to a specific person included in the received image. As a result, for the person being focused on, an evaluation score can be output that includes the situation of the person grasped from the entire received image (listening to the teacher, playing happily with friends, etc.), so that the image scoring can be appropriately performed.
[0201] In the scoring system according to one aspect of the present disclosure, the receiving unit receives designation information for designating an image region of a person included in the received image, and the image scoring model has learned to obtain learning designation information for designating an image region of a person included in the learning image and output an evaluation score for the learning image. The evaluation score output unit preferably outputs the evaluation score for the received image based on the designation information using the image scoring model.
[0202] According to the above configuration, since the image region of the person to be focused on is designated when outputting the evaluation score, the accuracy of the evaluation score output from the image scoring model is improved, and the image scoring can be performed more appropriately.
[0203] In the scoring system according to one aspect of the present disclosure, it is preferable that the designation information generation unit sequentially generates the designation information for each person among the plurality of persons included in the received image.
[0204] According to the above configuration, since the designation information can be automatically generated for all the persons included in the received image, the evaluation scores for the received image when focusing on each person can be output without omission, and the image scoring can be easily performed.
[0205] The scoring system according to one aspect of the present disclosure further includes a person image detection unit that extracts a feature amount of the image for the received image and detects a person, and the receiving unit preferably receives the designation information for designating an image region of the person detected by the person image detection unit.
[0206] According to the above configuration, since the designation information for designating the image area of a person can be automatically generated based on the feature amount, it is possible to easily designate the image area of the person to be focused on when outputting the evaluation score for the received image.
[0207] The scoring system according to one aspect of the present disclosure preferably further includes a first application processing unit that performs at least one of classification, sorting, or selection processing on the received image based on a predetermined condition regarding the evaluation score output by the evaluation score output unit.
[0208] According to the above configuration, it is possible to easily perform classification, sorting, selection, etc. of the received image based on the result of scoring.
[0209] The scoring system according to one aspect of the present disclosure preferably further includes a text generation unit that generates text for explaining the content of the received image using a learned text generation model so as to generate text for explaining the content of the image when focusing on a specific person included in the learning image.
[0210] According to the above configuration, since text for explaining the content of the received image for the person who is focused on when outputting the evaluation score is generated, the user can easily grasp the content of the received image including reasons for high or low evaluation scores.
[0211] The scoring system according to one aspect of the present disclosure preferably further includes a second application processing unit that performs at least one of classification, sorting, or selection processing on the received image based on a predetermined condition regarding at least one of the evaluation score output by the evaluation score output unit and the text generated by the text generation unit.
[0212] According to the above configuration, not only the scoring result but also the classification, sorting, selection, etc. of the received image can be easily performed based on the text describing the content of the received image.
[0213] The scoring system according to one aspect of the present disclosure further preferably includes an attribute information output unit that acquires an image of a person and outputs attribute information regarding the attributes of the person using a learned attribute information output model that outputs the attribute information for the person on whom attention is focused when the evaluation score is output.
[0214] According to the above configuration, since the attribute information of the person on whom attention is focused when the evaluation score is output is output, the user can easily grasp the content of the received image.
[0215] The scoring system according to one aspect of the present disclosure further preferably includes a third application processing unit that performs at least one or more processes of classification, sorting, or selection on the received image based on a predetermined condition related to at least any one of the evaluation score output by the evaluation score output unit and the attribute information output by the attribute information output unit.
[0216] According to the above configuration, not only the scoring result but also the classification, sorting, selection, etc. of the received image can be easily performed based on the attribute information of the person on whom attention is focused when the evaluation score is output.
[0217] The scoring system according to one aspect of the present disclosure further preferably includes a first document generation unit that generates a predetermined document using the received image on which at least one or more processes of the selection, the classification, or the sorting have been performed.
[0218] According to the above configuration, since processes such as classification, sorting, and selection of the received image are performed based on the scoring result, etc., a document (such as a daily report or an album) can be automatically generated using an appropriate received image, and the work load of the user is also reduced.
[0219] The scoring system according to one aspect of the present disclosure preferably further includes an image cutting unit that cuts out, including an image of a person who was focused on when the evaluation score was output, based on the text generated by the text generation unit and the evaluation score.
[0220] According to the above configuration, since the image is cut out based on the content of the text indicating the reason why the evaluation score is high, for example, an image that accurately cuts out the range in which it is possible to grasp the state of having fun playing with friends is automatically trimmed to a more appropriate range.
[0221] The scoring system according to one aspect of the present disclosure preferably further includes a second document generation unit that generates a document using the image of the person cut out by the image cutting unit.
[0222] According to the above configuration, a document (such as a daily report or an album) can be automatically generated from an image in which an appropriate range has been trimmed from the received image, and the work load of the user is also reduced.
[0223] The scoring system according to one aspect of the present disclosure preferably further includes a transmission unit that transmits the image of the person cut out by the image cutting unit to an information processing terminal owned by a person related to the person.
[0224] According to the above configuration, since the image in which an appropriate range has been cut out is automatically transmitted to the person related to the person (child) shown in the image (such as a guardian), the guardian can firmly confirm the state of the child in the facility or the like, and thus can entrust and watch over the child at the facility with confidence.
[0225] The scoring system according to one aspect of the present disclosure preferably further includes a learning unit that learns the image scoring model.
[0226] According to the above configuration, it is possible to generate a learned model used to achieve the above effect.
[0227] The scoring system according to one aspect of the present disclosure preferably further includes an alert unit that presents an alert when the evaluation score output by the evaluation score output unit meets a predetermined condition.
[0228] According to the above configuration, for example, when receiving an alert indicating that the evaluation score is low, the user can retake a photo of the target child.
[0229] The scoring system according to one aspect of the present disclosure preferably further includes a storage processing unit that performs a process for storing only the received image whose evaluation score output by the evaluation score output unit meets a predetermined condition in a storage device.
[0230] According to the above configuration, for example, only images with high evaluation scores can be stored in the storage (storage device), so that the storage capacity of the storage can be saved.
[0231] The method according to one aspect of the present disclosure includes a reception step of receiving, as a received image, an image including at least one or more persons using a scoring system, and obtaining, as a learning image, an image including one or more persons, and using an artificial intelligence model trained to output an evaluation score for the learning image when paying attention to a specific person included in the learning image, and an evaluation score output step of outputting an evaluation score for the received image.
[0232] According to the above configuration, the same effect as the above-described scoring system is achieved.
[0233] A program according to one aspect of the present disclosure is a program for causing a computer to function as the above-described scoring system, and causes the computer to function as each of the above units.
[0234] According to the above configuration, the same effects as those of the above-described scoring system are achieved.
[0235] One aspect of the present disclosure is useful for an information processing system for scoring images.
Explanation of Signs
[0236] 1 Scoring system 10, 10A Server 20, 20A User terminal 11, 21, 11A, 21A Storage unit 12, 22, 12A, 22A Control unit 23A Imaging unit 101, 201 Processor 102, 202 Main memory 103, 203 Auxiliary memory 104, 204 Communication interface 110, 210 Bus 121, 221 Reception unit 122, 222 Evaluation score output unit 123, 223 Designated information generation unit 124, 224 Person image detection unit 125, 225 Text generation unit 126, 226 Attribute information output unit 127, 227 Image cutting unit 128, 228 Transmission unit 129, 229 First application processing unit 130, 230 Second application processing unit 131, 231 Third application processing unit 132, 232 First document generation unit 133, 233 Second document generation unit 134, 234 Learning unit 235 Alert unit 236 Save processing unit 205 Input / output interface 206 Input device 207 Output device M, M’ Artificial intelligence model M1, M1' Image Scoring Model M2, M2' Text Generation Model M3, M3' Attribute Information Output Model
Claims
1. A receiving unit that receives, as a received image, an image including at least one or more persons, An image scoring system comprising: an evaluation score output unit that obtains an image including one or more persons as a learning image, and outputs an evaluation score for the received image using an image scoring model trained to output an evaluation score for the learning image when focusing on a specific person included in the learning image.
2. The receiving unit receives designation information for designating an image region of a person included in the received image, The image scoring model has been trained to obtain learning designation information for designating an image region of a person included in the learning image and output an evaluation score for the learning image, The evaluation score output unit outputs the evaluation score for the received image based on the designation information using the image scoring model. The scoring system according to claim 1.
3. The scoring system according to claim 2, further comprising a designation information generation unit that sequentially generates the designation information for each person for a plurality of persons included in the received image.
4. The scoring system according to claim 2, further comprising a person image detection unit that extracts a feature amount of an image of the received image to detect a person, The receiving unit receives the designation information for designating an image region of a person detected by the person image detection unit.
5. The scoring system according to claim 1, further comprising a first application processing unit that performs at least one or more of classification, sorting, or selection processing on the received image based on a predetermined condition regarding the evaluation score output by the evaluation score output unit.
6. The scoring system according to claim 1, further comprising a text generation unit that generates text for explaining the content of the received image using a text generation model trained to generate text for explaining the content of an image when focusing on a specific person included in the learning image.
7. A second application processing unit that performs at least one or more processes of classification, sorting, or selection on the received image based on a predetermined condition related to at least either the evaluation score output by the evaluation score output unit or the text generated by the text generation unit. The scoring system according to claim 6, further comprising:
8. An attribute information output unit that uses a learned attribute information output model that acquires an image of a person and outputs attribute information related to the attributes of the person, and outputs the attribute information for the person who is focused on when the evaluation score is output. The scoring system according to claim 1, further comprising:
9. A third application processing unit that performs at least one or more processes of classification, sorting, or selection on the received image based on a predetermined condition related to at least either the evaluation score output by the evaluation score output unit or the attribute information output by the attribute information output unit. The scoring system according to claim 8, further comprising:
10. A first document generation unit that generates a predetermined document using the received image on which at least one or more processes of the selection, classification, or sorting have been performed. The scoring system according to any one of claims 5, 7, or 9, further comprising:
11. An image clipping unit that clips, based on the text generated by the text generation unit and the evaluation score, an image including the image of the person who is focused on when the evaluation score is output. The scoring system according to claim 6, further comprising:
12. A second document generation unit that generates a document using the image of the person clipped by the image clipping unit. The scoring system according to claim 11, further comprising:
13. A transmission unit that transmits the image of the person clipped by the image clipping unit to an information processing terminal owned by a person related to the person. The scoring system according to claim 11, further comprising:
14. A learning unit that learns the image scoring model. The scoring system according to claim 1, further comprising:
15. An alert unit that presents an alert when the evaluation score output by the evaluation score output unit meets a predetermined condition. The scoring system according to claim 1, further comprising:
16. The scoring system according to claim 1, further comprising a storage processing unit that performs a process for storing only the received images for which the evaluation score output by the evaluation score output unit meets a predetermined condition in a storage device.
17. Using a scoring system, a reception step of receiving, as a received image, an image including at least one or more persons; an evaluation score output step of outputting an evaluation score for the received image using an artificial intelligence model trained to acquire an image including one or more persons as a learning image and output an evaluation score for the learning image when paying attention to a specific person included in the learning image.
18. A program for causing a computer to function as the scoring system according to claim 1, the program for causing a computer to function as each unit.
Citation Information
Patent Citations
Image candidate determination device, image candidate determination method, program for controlling the image candidate determination device, and recording medium storing the program
JP2023001178A