Group-specific model generation system, server, and group-specific model generation program

The group-specific model generation system addresses the challenge of accurate object detection and recognition in edge devices by clustering and fine-tuning neural networks for each facility's imaging conditions, enhancing detection and recognition accuracy.

JP7762410B2Active Publication Date: 2025-10-30AWL INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2021175859
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2025-10-30
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

Existing edge devices struggle with highly accurate object detection and recognition due to the use of extremely lightweight trained DNN models, and fine-tuning these models across numerous facilities is impractical and ineffective due to diverse imaging conditions.

Method used

A group-specific model generation system that collects and processes images from multiple facilities, removes human images, extracts features, clusters images based on characteristics, and generates tailored, lightweight neural network models for each group using fine-tuning or transfer learning.

Benefits of technology

Enables highly accurate object detection and recognition across numerous facilities by generating group-specific models that adapt to each facility's imaging conditions, improving inference accuracy and reducing the need for extensive model retraining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007762410000001
    Figure 0007762410000001
  • Figure 0007762410000002
    Figure 0007762410000002
  • Figure 0007762410000003
    Figure 0007762410000003
Patent Text Reader

Abstract

To perform a highly accurate object detection process and a recognition process by using an extremely light learned neural network model even when captured images to be subjected to object detection processes and recognition processes of entire edge-side apparatuses are captured images by cameras of a large number of facilities, in a group-specific model generation system, a server, and a group-specific model generation program.SOLUTION: Captured images collected from built-in cameras of signages installed in a plurality of stores are grouped by a Gaussian mixture model on the basis of feature vectors of these captured images (S5), grouping of the built-in cameras that captured these captured images is performed on the basis of grouping results of the captured images (S9), and fine-tuning of an original learned DNN model is performed by using the captured images by the built-in cameras in each group subjected to the grouping (S10). Accordingly, it is possible to generate a group-specific learned DNN model that is suitable for the captured images by the built-in cameras in each group.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a group-specific model generation system, a server, and a group-specific model generation program. [Background technology]

[0002] Conventionally, systems have been known that analyze images (such as object detection and object recognition) captured by a camera installed in a facility such as a store using a device (a so-called edge device) located on the facility side where the camera is installed (see, for example, Patent Document 1). When performing object detection or object recognition using such an edge device, a trained deep neural network model (DNN model) with a low processing load (a so-called "light") is implemented in the edge device, and this trained DNN model is used to perform object detection processing and object recognition processing on images captured by a camera connected to the edge device. Here, due to the vulnerability of computer resources in edge devices, it is desirable that the trained DNN model implemented in the edge device be an extremely light DNN model (with a very low processing load). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6178942 Summary of the Invention [Problem to be solved by the invention]

[0004] However, when implementing such an extremely lightweight (very low processing load) trained DNN model on edge devices located in numerous facilities to perform object detection and object recognition processing on images captured by cameras in numerous facilities, the following problems arise: First, it is difficult to perform highly accurate object detection and object recognition processing with an extremely lightweight trained DNN model.

[0005] Furthermore, when using an extremely lightweight trained DNN model such as the one described above, it is desirable to perform fine-tuning or transfer learning of the original trained DNN model for each camera in each facility using images captured by that camera to ensure accuracy. However, in the case of a large chain store (e.g., convenience store), which has thousands of stores, performing fine-tuning or transfer learning of the trained DNN model for each camera installed in thousands of stores using images captured by that camera would take an enormous amount of time. Therefore, it is not realistic to perform fine-tuning or transfer learning of the trained DNN model for each camera installed in each facility using images captured by that camera, as described above. However, even if an extremely lightweight trained DNN model were to be fine-tuned or transferred using images captured by all cameras installed in thousands of stores, due to the diversity of images acquired (collected) from cameras in thousands of stores (such as the layout, lighting conditions, presence or absence of people, and interior design of each store), an extremely lightweight DNN model would often be unable to fully train.

[0006] The present invention is devised to solve the above-mentioned problems, and aims to provide a group-specific model generation system, server, and group-specific model generation program that enable highly accurate object detection processing and object recognition processing even when the captured images that are the subject of object detection processing and object recognition processing of the entire edge-side apparatus (edge-side device) are images captured by the imaging means of a large number of facilities, such as thousands of stores, and even when the trained neural network model used is an extremely lightweight trained neural network model. [Means for solving the problem]

[0007] In order to solve the above problems, a group-specific model generation system according to a first aspect of the present invention includes: a captured image collection means for collecting captured images from each of image capture means installed in a plurality of facilities; a person image removing means for removing photographed images in which people are reflected from the photographed images collected by the photographed image collecting means; and a person image removing means for removing photographed images in which people are reflected from the photographed images of the facility remaining after the photographed images in which people are reflected are removed by the person image removing means. image feature extraction means for extracting features from each of the captured images; The facility The captured image is extracted by the image feature extraction means. Facilityan image clustering means for grouping the captured images based on the characteristics of each of the captured images; Facility A photographing means classification means for classifying the photographed images into groups of photographing means that have taken these photographed images based on the grouping results of the photographed images, and a classification means for classifying the photographed images into groups of the photographing means that have taken the photographed images based on the grouping results of the photographing means classification means. Human Detection For use or person recognition and a group-specific model generation means for generating a group-specific trained neural network model suitable for the image captured by the imaging means of each group by performing fine tuning or transfer learning of the trained neural network model for the group. The present invention further comprises a person image extraction means for extracting images in which people are captured from the images captured by the image capture means of each group after grouping by the image capture means classification means, and the group-specific model generation means performs fine tuning or transfer learning of the trained neural network model for detecting or recognizing people using the captured images in which people are captured extracted by the person image extraction means. .

[0008] In this group-specific model generation system, the group-specific trained neural network model generated by the group-specific model generation means and suitable for the images captured by the imaging means of each group is transmitted to and stored in an edge-side device disposed in a facility where the imaging means of each group is installed, and the edge-side device performs a calculation for the images captured by the imaging means of each group. people Detection or people Recognition may also be performed.

[0009] In this group-specific model generation system, people For detection or people It is possible to make more accurate inferences than pre-trained neural network models for recognition. people For detection or people The system further includes a pseudo-labeling means for performing inference on the images captured by the imaging means of each group using a trained high-precision neural network model for recognition, and assigning pseudo-labels based on the inference results to the images captured by the imaging means of each group as correct labels, and the group-specific model generating means generates a model based on the original images captured by the imaging means of each group based on the images captured by the imaging means of each group and the correct labels assigned to the images captured by the imaging means of each group by the pseudo-labeling means. people For detection or people It is desirable to perform fine tuning or transfer learning of trained neural network models for recognition.

[0010] In this group-based model generation system, the image clustering means may change the number of clusters, which is the number of groups of the captured images, to determine the value of the information criterion for each number of clusters, and determine the number of clusters appropriate for the distribution of the features of the captured images extracted by the image feature extraction means, based on the value of the information criterion corresponding to each determined number of clusters.

[0011] In this group-specific model generation system, the image feature extraction means uses a trained neural network model to: The image of the facility remaining after the photographed image in which the person is reflected is removed by the human image removal means extracting a feature vector from each of the captured images, and the image clustering means The facility The captured image is extracted by the image feature extraction means. Facility The captured images may be grouped using a Gaussian mixture model based on the feature vectors of each captured image.

[0012] In this group-specific model generation system, the image clustering means, while changing the number of clusters, which is the number of groups of the photographed images, obtains the value of the Bayes information criterion for each number of clusters using the Gaussian mixture model, and based on the value of the Bayes information criterion corresponding to each obtained number of clusters, extracts the Facility The number of clusters may be determined to be suitable for the distribution of feature vectors of the captured image.

[0015] A server according to a second aspect of the present invention includes: a captured image collection means connected via a network to edge-side devices disposed in each of a plurality of facilities in which image capture means are installed, and which collects captured images from each of the image capture means; a person image removing means for removing photographed images in which people are reflected from the photographed images collected by the photographed image collecting means; and a person image removing means for removing photographed images in which people are reflected from the photographed images of the facility remaining after the photographed images in which people are reflected are removed by the person image removing means. image feature extraction means for extracting features from each of the captured images; The facility The captured image is extracted by the image feature extraction means. Facilityan image clustering means for grouping the captured images based on the characteristics of each of the captured images; Facility A photographing means classification means for classifying the photographed images into groups of photographing means that have taken these photographed images based on the grouping results of the photographed images, and a classification means for classifying the photographed images into groups of the photographing means that have taken the photographed images based on the grouping results of the photographing means classification means. Human Detection For use or person recognition and a group-specific model generation means for generating a group-specific trained neural network model suitable for the image captured by the image capture means of each group by performing fine tuning or transfer learning of the trained neural network model for the group. The present invention further comprises a person image extraction means for extracting images in which people are captured from the images captured by the image capture means of each group after grouping by the image capture means classification means, and the group-specific model generation means performs fine tuning or transfer learning of the trained neural network model for detecting or recognizing people using the captured images in which people are captured extracted by the person image extraction means. .

[0016] In this server, the group-specific trained neural network model generated by the group-specific model generation means and suitable for the images captured by the imaging means of each group may be transmitted to and stored in an edge-side device located in the facility where the imaging means of each group is installed.

[0017] A group-specific model generation program according to a third aspect of the present invention includes a computer including: a captured image collection means for collecting captured images from each of image capture means installed in a plurality of facilities; a person image removing means for removing photographed images in which people are reflected from the photographed images collected by the photographed image collecting means; and a person image removing means for removing photographed images in which people are reflected from the photographed images of the facility remaining after the photographed images in which people are reflected are removed by the person image removing means. image feature extraction means for extracting features from each of the captured images; The facility The captured image is extracted by the image feature extraction means. Facility an image clustering means for grouping the captured images based on the characteristics of each of the captured images; Facility A photographing means classification means for classifying the photographed images into groups of photographing means that have taken these photographed images based on the grouping results of the photographed images, and a classification means for classifying the photographed images into groups of the photographing means that have taken the photographed images based on the grouping results of the photographing means classification means. Human Detection For use or person recognitiona group-specific model generation program for causing the program to function as a group-specific model generation means for generating a group-specific trained neural network model suitable for an image captured by the imaging means of each group by performing fine tuning or transfer learning of a trained neural network model for the group; and causing the computer to further function as a person image extraction means for extracting images in which people are captured from images captured by the image capture means of each group after grouping by the image capture means classification means, and the group-specific model generation means performs fine tuning or transfer learning of the trained neural network model for detecting or recognizing people using the captured images in which people are captured extracted by the person image extraction means. . [Effects of the Invention]

[0018] According to the group-specific model generation system according to the first aspect of the present invention, the server according to the second aspect, and the group-specific model generation program according to the third aspect, photographed images collected from each of the photographing means installed in a plurality of facilities are Then, the photographed images of the facility are extracted based on the extracted features of each of the photographed images of the facility. Divide into groups, This facility Based on the grouping results of the photographed images, the photographing means that captured these photographed images are grouped, and the photographed images by the photographing means of each group are Of these, images that include people Using Human Detection For use or person recognition This allows for fine tuning or transfer learning of the trained neural network model for each group. This allows for the model to be tailored to the images captured by each group's photography method (specialized for the images captured by each group's photography method). 、 By group For human detection or recognition It is possible to generate a trained neural network model for each group. For human detection or recognition Even if the trained neural network model is extremely lightweight, it can provide highly accurate results for images captured by each group of imaging methods. people Detection process and people It is also possible to perform recognition processing on the entire edge device. people Detection process and people Even if the images to be recognized are images taken by the camera means of a large number of facilities such as thousands of stores, these camera means can be grouped, and the images taken by a limited number of camera means (for example, several hundred cameras) after grouping can be used to recognize the original images. For human detection or recognition It is possible to fine-tune or transfer learn a trained neural network model, so that the original For human detection or recognitionEven if the trained neural network model is extremely light, it is possible to increase the possibility of performing appropriate machine learning (reduce the possibility of the model not being able to learn properly). people Detection process and people The images to be recognized are taken by the camera of many facilities such as thousands of stores, and the original For human detection or recognition The trained neural network model and each group generated above For human detection or recognition Even if the trained neural network model is an extremely lightweight trained neural network model, the above generated group-specific For human detection or recognition Using a trained neural network model, we can accurately measure the images captured by each group's photography method. people Detection process and people It becomes possible to perform recognition processing. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a block diagram showing a schematic configuration of a group-specific model generation system according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram showing a schematic hardware configuration of the signage in FIG. [Figure 3] FIG. 2 is a block diagram showing the hardware configuration of the signage learning management server in FIG. 1. [Figure 4] FIG. 2 is a functional block diagram of the signage learning management server. [Figure 5] FIG. 5 is an explanatory diagram of the data flow between the functional blocks in FIG. 4. [Figure 6] A flowchart of the group-specific trained DNN model generation process in the above-mentioned group-specific model generation system. [Figure 7] FIG. 7 is an explanatory diagram of the grouping process of the built-in cameras shown in S9 in FIG. 6. [Figure 8] FIG. 7 is an explanatory diagram of the process of estimating an appropriate number of clusters using a Gaussian mixture model shown in S5 in FIG. 6. [Figure 9]7A and 7B are diagrams showing examples of photographed images included in each group as a result of grouping processing of photographed images included in the "group of photographed images not showing people" shown in S7 in FIG. 6; [Figure 10] A diagram showing inference accuracy evaluation indicators before and after fine-tuning of a group-specific trained DNN model generated by the above-mentioned group-specific model generation system. DETAILED DESCRIPTION OF THE INVENTION

[0020] A group-specific model generation system, a server, and a group-specific model generation program according to an embodiment of the present invention will be described below with reference to the drawings. FIG. 1 is a block diagram showing a schematic configuration of a group-specific model generation system 10 according to this embodiment. As shown in FIG. 1, the group-specific model generation system 10 mainly includes signage devices 2a and 2b ("edge-side devices" in the claims), which are tablet terminals for digital signage, installed in each store ("facility" in the claims) Sa, Sb, etc. of a chain store, and a signage learning management server 1 (corresponding to a "server" and a "computer" in the claims) connected to these signage devices 2a and 2b via the Internet. In the following description, "signage 2" collectively refers to the signage devices 2a and 2b, etc., and "store S" collectively refers to stores Sa, Sb, etc. The group-specific model generation system 10 includes one or more signage devices 2 and a wireless LAN router 4 within each store S. Each signage device 2 includes a built-in camera 3 ("photographing means" in the claims).

[0021] The signage 2 displays content such as advertisements on its touch panel display 14 (see Figure 2) to customers who visit the store S (who are in front of the signage 2), and also detects customers who appear in the frame image based on the frame image from its built-in camera 3, and performs image analysis processing such as estimating the attributes of the detected customers.

[0022] The signage learning management server 1 is a server installed in the management department (head office, etc.) of the store S. As will be described in detail later, the signage learning management server 1 generates group-specific trained DNN (Deep Neural Networks) models suitable for images captured by the built-in camera 3 of each signage 2, and transmits the generated group-specific trained DNN models to each signage 2 for installation.

[0023] Next, the hardware configuration of the tablet-type signage 2 will be described with reference to Figure 2. In addition to the built-in camera 3, the signage 2 includes an SoC (System-on-a-Chip) 11, a touch panel display 14, a speaker 15, memory 16 for storing various data and programs, a communication unit 17, a secondary battery 18, and a charging terminal 19. The SoC 11 includes a CPU 12 for controlling the entire device and performing various calculations, and a GPU 13 used for inference processing of various trained DNN (Deep Neural Networks) models.

[0024] The memory 16 stores a group-specific trained DNN model 20 (referred to as a "group-specific trained neural network model" in the claims) suitable for images captured by the built-in camera 3 of the signage 2. The group-specific trained DNN model 20 includes multiple types of trained DNN models, such as a trained DNN model for customer (person) detection (including a trained DNN model for detecting a customer's face or head) and a trained DNN model for customer (person) recognition, such as customer attribute estimation. The communication unit 17 includes a communication IC and an antenna. The signage 2 is connected to the signage learning management server 1 on the cloud via the communication unit 17 and the Internet. The secondary battery 18 is a rechargeable battery, such as a lithium-ion battery, that stores power from a commercial power source after it has been converted to DC power by an AC / DC converter and supplies it to each component of the signage 2.

[0025] Next, the hardware configuration of the signage learning management server 1 will be described with reference to Fig. 3. The signage learning management server 1 includes a CPU 21 that controls the entire device and performs various calculations, a hard disk 22 that stores various data and programs, a RAM (Random Access Memory) 23, a display 24, an operation unit 25, and a communication unit 26. The programs stored on the hard disk 22 include a group-specific model generation program 27.

[0026] Figure 4 mainly shows the functional blocks of the signage learning management server 1. The following explanation of Figure 4 will explain the correspondence between each functional block in the figure and each constituent element (means) in the claims, as well as an overview of the function of each functional block. The signage learning management server 1 includes, as functional blocks, a captured image collection unit 31, a frame image extraction unit 32, a human image removal unit 33, an image feature vector extraction unit 34, an image clustering unit 35, a camera classification unit 36, and an automatic fine-tuning unit 37. The automatic fine-tuning unit 37 also includes a human image extraction unit 38, a pseudo-labeling unit 39, and a group-specific model generation unit 41. The captured image collection unit 31, human image removal unit 33, image feature vector extraction unit 34, image clustering unit 35, camera classification unit 36, human image extraction unit 38, pseudo-labeling unit 39, and group-specific model generation unit 41 correspond to the captured image collection means, human image removal means, image feature extraction means, image clustering means, camera classification means, human image extraction means, pseudo-labeling means, and group-specific model generation means in the claims, respectively. The captured image collection unit 31 is primarily realized by the communication unit 26, CPU 21, and group-specific model generation program 27 in Figure 3. The frame image extraction unit 32, human image removal unit 33, image feature vector extraction unit 34, image clustering unit 35, camera classification unit 36, automatic fine-tuning unit 37, human image extraction unit 38, pseudo-labeling unit 39, and group-specific model generation unit 41 are realized by the CPU 21 and group-specific model generation program 27 in Figure 3.

[0027] The captured image collection unit 31 collects captured images (in this embodiment, videos (captured videos) captured by each built-in camera 3) from each of the built-in cameras 3 of the signage 2 installed in multiple stores S. The frame image extraction unit 32 extracts frame images from the videos captured by each built-in camera 3. The human image removal unit 33 removes captured images that show people from the frame images extracted by the frame image extraction unit 32 (all captured images), thereby extracting a "group of captured images that do not show people" (i.e., a group of captured images of the store). The image feature vector extraction unit 34 extracts a feature vector from each of the captured images of the store ("captured images of the facility" in the claims) using a trained DNN model for vector extraction. Then, the image clustering unit 35 groups the group of captured images of the store using a Gaussian Mixture Model (GMM) based on the feature vectors of each captured image extracted by the image feature vector extraction unit 34.

[0028] Furthermore, as will be described in detail later, the camera classification unit 36 ​​groups the built-in cameras 3 that captured the captured images of the store based on the grouping results of the image clustering unit 35. The person image extraction unit 38 extracts captured images in which people are captured from the captured images by the built-in cameras 3 of each group after grouping by the camera classification unit 36. More precisely, the person image extraction unit 38 targets all frame images extracted by the frame image extraction unit 32 (the entire group of captured images including captured images in which people are captured and captured images in which people are not captured) and extracts captured images in which people are captured from the captured images by the built-in cameras 3 of each group after grouping by the camera classification unit 36 ​​(captured images in which people are captured and captured images in which people are not captured).

[0029] In addition, the pseudo-labeling unit 39 uses a trained high-precision DNN model 40 (corresponding to the "trained high-precision neural network model" in the claims) for detecting or recognizing customers, which is capable of making more accurate inferences than the trained DNN model (hereinafter referred to as the "original trained DNN model") for detecting or recognizing customers (people or people's faces or heads) that is the basis for the group-specific trained DNN model 20 stored in the memory 16 of the signage 2, to make inferences on the captured images (captured images in which people are reflected) extracted by the human image extraction unit 38 from the images captured by the built-in cameras 3 of each group, and assigns pseudo-labels based on this inference result as correct labels to the captured images extracted by the human image extraction unit 38. The group-specific model generation unit 41 generates group-specific trained DNN models 20 (corresponding to the "group-specific trained neural network model" in the claims) suitable for the images captured by the built-in cameras 3 of each group by fine-tuning the original trained DNN model based on the images captured by the built-in cameras 3 of each group and the correct labels assigned to these captured images by the pseudo-labeling unit 39. The CPU 21 of the signage learning management server 1 uses the communication unit 26 to transmit and store the group-specific trained DNN models 20 suitable for the images captured by the built-in cameras 3 of each group to the signage 2 having the built-in cameras 3 of the groups corresponding to the group-specific trained DNN models 20. The trained DNN model for customer detection or recognition that is the basis for the group-specific trained DNN model 20 corresponds to the "original trained neural network model for object detection or object recognition" in the claims.

[0030] Next, the data flow in the group-specific model generation system 10 will be described with reference to the flowcharts of FIGS. 5 and 6. FIG. 5 shows input / output data for each functional block of the signage learning management server 1 described in FIG. 4 above. FIG. 6 is a flowchart of the group-specific trained DNN model generation process in the group-specific model generation system 10. First, the captured image collection unit 31 of the signage learning management server 1 prompts each signage 2 to transfer selected video from the video captured by its built-in camera 3 (video captured in a specified time period) to the signage learning management server 1. In response, each signage 2 transfers the video captured in the time period specified by the captured image collection unit 31 of the signage learning management server 1 from the video captured by its own built-in camera 3 (photographed video) to the signage learning management server 1 (S1 in FIG. 6). Next, the frame image extraction unit 32 of the signage learning management server 1 extracts frame images from the video captured by each built-in camera 3 (S2 in FIG. 6). As shown in Figure 5, the frame image extraction unit 32 performs an extraction process to create frame images extracted from the footage captured by all built-in cameras 3 (a collection of all captured images (a collection of "captured images that include people" and "captured images that do not include people").

[0031] Next, as shown in S3 of FIG. 6, the human image removal unit 33 of the signage learning management server 1 performs human head detection on each of all frame images extracted by the frame image extraction unit 32 (all captured images (a collection of "captured images with people in them" and "captured images without people in them")), and uses the results of this head detection to remove "captured images with people in them" from all captured images, thereby extracting "captured images without people in them" (i.e., a group of captured images showing only the store (hereinafter referred to as "captured images of the store"). More specifically, the human image removal unit 33 of the signage learning management server 1 uses a trained DNN model for (human) head detection to detect "captured images with people in them" from all captured images (all captured images) extracted by the frame image extraction unit 32, and extracts "captured images without people in them" from all captured images. Then, all detected "photographed images in which people are captured" ("group of photographed images in which people are captured") are removed to extract a "group of photographed images in which people are not captured." The extraction process for the "group of photographed images in which people are not captured" by the human image removal unit 33 extracts, for example, 100 "photographed images in which people are not captured" for each built-in camera 3. Therefore, for example, if the number of built-in cameras 3 in the group-specific model generation system 10 (i.e., the number of signages 2 connected to the signage learning management server 1) is 500, the extraction process by the human image removal unit 33 collects (100 x 500) "photographed images in which people are not captured" ("images of the store"). This collection of photographed images ("group of photographed images in which people are not captured" ("group of photographed images of the store")) is used to group (classify) the built-in cameras 3, as described below.

[0032] Next, as shown in S4 of FIGS. 5 and 6, the image feature vector extraction unit 34 of the signage learning management server 1 extracts feature vectors for each captured image included in the "group of captured images not showing people" ("group of captured images of a store") using the pretrained ResNet 50. As a result, as shown in FIG. 5, it is possible to obtain feature vectors (2048-dimensional feature vectors) for each captured image included in the "group of captured images not showing people." For example, as described above, when (100 × 500) "captured images not showing people" ("store images") are collected through the extraction process by the human image removal unit 33, it is possible to obtain (100 × 500) 2048-dimensional feature vectors.

[0033] Next, the image clustering unit 35 of the signage learning management server 1 groups the captured images included in the "group of captured images not showing people" using a Gaussian mixture model based on the feature vectors (2048-dimensional feature vectors) of each captured image. Specifically, the image clustering unit 35 first automatically estimates an appropriate number of clusters k using a Gaussian mixture model based on the feature vectors (2048-dimensional feature vectors) of each captured image extracted by the image feature vector extraction unit 34 (S5). A method for estimating an appropriate number of clusters k using this Gaussian mixture model will be described in detail later.

[0034] Next, the image clustering unit 35 of the signage learning management server 1 checks whether the estimated number of clusters k is equal to or less than the planned (assumed upper limit) number of clusters j (S6). As a result, if the estimated number of clusters k is equal to or less than the planned number of clusters j (YES in S6), the image clustering unit 35 classifies the photographed images included in the "photographed image group not showing people" extracted by the human image removal unit 33 into k "photographed image groups A1 to A2 not showing people." kIn the determination of S6, if the number of clusters k estimated using the Gaussian mixture model is a number that exceeds the planned (assumed upper limit) number of clusters j (NO in S6), the image clustering unit 35 groups the photographed images included in the "group of photographed images in which no people are shown" extracted by the human image removal unit 33 into j "groups of photographed images A1 to A2 in which no people are shown," which is the planned (assumed upper limit) number of clusters. j In FIG. 5, the image clustering unit 35 groups the photographed images included in the "photographed image group not showing people" extracted by the human image removal unit 33 into k "photographed image groups A1 to A2" (S8). k " is shown as an example of grouping.

[0035] Next, the camera classification unit 36 ​​of the signage learning management server 1 groups the built-in cameras 3 that captured these captured images based on the grouping results of the captured images by the image clustering unit 35 (S9).

[0036] The grouping of the built-in camera 3 will be described with reference to Fig. 7. The image clustering unit 35 groups k "photographed images A1 to A2 that do not show people" into groups. k Each image in the group of "photographed images A1 to A2" is assigned the camera ID of the built-in camera 3. This camera ID is used for the group of "photographed images A1 to A2" in which no person is photographed. k Each of the photographed images in the above "photographed image group A1 to A2" is information taken over from each of the photographed images extracted by the frame image extraction unit 32 (each of the photographed images in the entire photographed image group). k By referring to the camera ID assigned to each image in the "Images of the store" section, it is possible to easily determine which built-in camera 3 with which camera ID the "images of the store" ("images that do not show people") in each group grouped by the image clustering unit 35 were taken (the correspondence between each camera ID and each group).

[0037] For example, for ease of explanation, it is assumed that the number of clusters k estimated by the image clustering unit 35 is 2, and the image clustering unit 35 divides the captured images included in the "group of captured images not showing people" extracted by the human image removal unit 33 into group 1 and group 2 as shown in FIG. 7. Looking at FIG. 7, most of the captured images included in group 1 ("store captured images") are images captured by the built-in cameras 3 with camera IDs 1 to 21. For example, group 1 includes 45 images captured by camera ID 0 and just under 100 images captured by camera ID 1, while group 2 does not include any images captured by camera ID 0 or camera ID 1. From this, it can be seen that camera ID 0 and camera ID 1 correspond to group 1, not group 2. Similarly, the captured images of cameras IDs 2 to 21 are included only in group 1, not in group 2, and therefore camera IDs 2 to 21 correspond to group 1, not group 2.

[0038] 7, the images captured by camera ID 31 are included in both group 1 and group 2, but group 2 contains just under 80 images, while group 1 contains only a few. Therefore, by majority vote, camera ID 31 corresponds to group 2, not group 1. In this way, when the images captured by each camera ID are included in multiple groups, the camera ID in question corresponds to the group that contains the most images captured by that camera ID. Note that when clustering by the image clustering unit 35 is successful, it is rare for the images captured by each camera ID to be included in multiple groups, and even when the images captured by each camera ID are included in multiple groups, there will be a large difference in the number of images (number of images) that belong to each group.

[0039] As described above, the correspondence between each camera ID and each group is known, and the camera classification unit 36 ​​shown in FIG. 5 performs grouping of the built-in cameras 3 shown in S9 above in accordance with this correspondence, and classifies the built-in cameras 3 into k (or j) groups. The automatic fine-tuning unit 37 shown in FIG. 4 performs automatic fine-tuning of the original trained DNN model (the trained DNN model for detecting or recognizing customers (people) that is the basis for the group-specific trained DNN model 20) using images captured by the built-in cameras 3 of each group grouped by the camera classification unit 36 ​​(S10 in FIG. 6). Note that the trained DNN model for detecting customers (people) includes a trained DNN model for detecting customers' faces and heads.

[0040] The details of the automatic fine tuning by the automatic fine tuning unit 37 are as follows: First, the automatic fine tuning unit 37 classifies all frame images (all captured images including captured images with and without people) extracted by the frame image extraction unit 32 into k groups of captured images C1 to C2 of the built-in cameras 3, which are grouped by the camera classification unit 36 ​​with reference to the camera IDs assigned to each of these frame images. k (Hereinafter, the set of images C1 to C2 captured by k camera groups will be referred to as k Here, the "camera group" refers to a group of built-in cameras 3 that has been grouped by the camera classification unit 36. Then, as shown in FIG. 5, the automatic fine tuning unit 37 uses the human image extraction unit 38 to extract the captured image groups C1 to C2 from the k camera groups. k (including photographed images with and without people) and extracting photographed images with people in them, and classifying them into k (of camera groups) “photographed image groups B1 to B2 with people in them.” k The extraction of "photographed images in which people are reflected" by the human image extraction unit 38 also uses a trained DNN model for human head detection, similar to the one used by the human image removal unit 33 to detect "photographed images in which people are reflected."

[0041] The human image extraction unit 38 extracts k “photographed images B1 to B2 containing people” k When the process of creating " is completed, the automatic fine-tuning unit 37 uses the pseudo-labeling unit 39 to create k "groups of photographed images B1 to B2 containing people" using a trained high-precision DNN model 40 (for detecting or recognizing customers) that can perform inference with higher accuracy than the original trained DNN model, as shown in FIG. k ", and the pseudo-label based on this inference result is used as the correct label, and the group of photographed images B1 to B k " is assigned to each captured image included in the "

[0042] Then, as shown in FIG. 5, the automatic fine tuning unit 37 generates a group of photographed images B1 to B2 containing people, to which the pseudo labels have been assigned by the group-specific model generation unit 41. k Each of the captured image groups B1 to B k By using this, fine-tuning of the original trained DNN model is performed to generate k group-specific trained DNN models 20 suitable for images captured by the built-in cameras 3 of each of the k groups. That is, for example, by fine-tuning the original trained DNN model based on each captured image included in the captured image group B1 and the correct labels assigned to these captured images, a group-specific trained DNN model 20 suitable for images captured by the built-in cameras 3 of the (first) camera group corresponding to the captured image group B1 is generated, and by fine-tuning the original trained DNN model based on each captured image included in the captured image group B2 and the correct labels assigned to these captured images, a group-specific trained DNN model 20 suitable for images captured by the built-in cameras 3 of the (second) camera group corresponding to the captured image group B2 is generated. Here, fine-tuning means re-training the weights of the entire trained DNN model to be newly generated using the weights of the original (existing) trained DNN model as initial values.

[0043] Next, the CPU 21 of the signage learning management server 1 evaluates the inference accuracy of each group-specific trained DNN model 20 after fine tuning by the group-specific model generation unit 41 (S11). The evaluation of the inference accuracy of each group-specific trained DNN model 20 after fine tuning will be described in detail in the description of FIG. 10 described later. If the inference accuracy (F1 value, etc.) of each group-specific trained DNN model 20 after fine tuning is significantly improved compared to the original trained DNN model before fine tuning, it can be evaluated that the result of grouping by the image clustering unit 35 is appropriate. If the result of grouping by the image clustering unit 35 is appropriate, the group-specific trained DNN model 20 after fine tuning is evaluated as being appropriate based on the captured image groups B1 to B2 used for the respective fine tuning. k The image data is transmitted to and stored in the signage 2 having the built-in camera 3 belonging to each camera group corresponding to the image data.

[0044] If necessary, by periodically repeating the processes S1 to S11 in Figure 6, even if the layout or environment (lighting conditions, interior design, etc.) of each store S changes, it is possible to maintain sufficient accuracy of each group-specific trained DNN model 20 generated by this group-specific model generation system 10.

[0045] Next, with reference to FIG. 8, the method for estimating the appropriate number of clusters k using the Gaussian mixture model described in S5 above will be described in detail. The diagram on the left side of FIG. 8 is a distribution diagram of the two-dimensional feature vectors of each captured image. The image feature vector extraction unit 34 extracts the 2048-dimensional feature vectors of each captured image included in the "group of captured images without people" captured by the built-in cameras 3 of the signage 2 in all stores using a pretrained ResNet 50. The 2048-dimensional feature vectors of each captured image are then reduced to two dimensions using the t-distributed Stochastic Neighbor Embedding (tSNE) algorithm, visualized as follows: Since the 2048-dimensional feature vectors extracted by the image feature vector extraction unit 34 themselves (their distribution) cannot be visualized, the distribution diagram shows the distribution of the feature vectors reduced to two dimensions using tSNE. However, in the clustering process using the Gaussian mixture model in the image clustering unit 35, the 2048-dimensional feature vectors of each captured image extracted by the image feature vector extraction unit 34 are used.

[0046] That is, while changing the number of clusters, which is the number of groups of captured images (i.e., while changing the number of Gaussian distributions included in the Gaussian mixture model), the image clustering unit 35 calculates a BIC (Bayesian information criterion) value for each number of clusters using the Gaussian mixture model based on (the distribution of) the (2048-dimensional) feature vectors of each captured image extracted by the image feature vector extraction unit 34, and then, based on the BIC value corresponding to each calculated number of clusters, calculates a number of clusters appropriate for the distribution of the feature vectors of the captured images extracted by the image feature vector extraction unit 34. That is, first, the image clustering unit 35 sequentially specifies the number of clusters (the number of Gaussian distributions included in the Gaussian mixture model) from 1 to 9, for example, and calculates a BIC value for each number of clusters using the Gaussian mixture model based on (the distribution of) the (2048-dimensional) feature vectors of each captured image extracted by the image feature vector extraction unit 34. The center graph in Figure 8 is a line graph showing the relationship between the number of clusters k calculated as above and the BIC value. In this graph, 1e7 is 1×10 7 Represents.

[0047] The image clustering unit 35 then determines the number of clusters (5 in this example) at the point where the gradient in the line graph stabilizes as the number of clusters appropriate for the distribution of the feature vectors of the captured images extracted by the image feature vector extraction unit 34. The number of clusters at the point where the gradient stabilizes is determined by comparing the change in the BIC value in the previous interval (e.g., the change in the BIC value between cluster number 4 and cluster number 5 in the line graph) with the change in the BIC value in the next interval (e.g., the change in the BIC value between cluster number 5 and cluster number 6), and adopting the number of clusters immediately before the change in the BIC value (decrease) becomes small. This is because, if the number of clusters is too large, the number of fine-tuning processes for the original trained DNN model described in S10 of FIG. 6 increases. Therefore, it is desirable to adopt as small a number of clusters k as possible if the BIC value, which is an index of the optimal model, does not change significantly even if the number of clusters k is increased. Here, generally, a smaller BIC value is preferable.

[0048] In the line graph shown in the middle diagram in FIG. 8, the number of clusters is 5 when the gradient settles. Therefore, the number of clusters appropriate for the distribution of the feature vectors of the captured images, calculated from the BIC value of the Gaussian mixture model, is 5. The diagram on the right side of FIG. 8 is a distribution diagram in which the two-dimensional feature vectors (after dimension reduction) of each captured image in the distribution diagram on the left side of FIG. 8 are grouped by color into groups 1 to 5 in accordance with the appropriate number of clusters (=5). Note that, since color drawings are generally not permitted in patent applications, the diagram on the right side of FIG. 8 shows the groups of each feature vector in grayscale. Furthermore, the diagrams on the left and right sides of FIG. 8 are provided to explain a method for estimating an appropriate number of clusters k using a Gaussian mixture model. In actual clustering processing by the image clustering unit 35, these distribution diagrams of feature vectors reduced to two dimensions are not used. However, the diagram on the left side of Figure 8 may be used to check how many groups the images included in the "group of images not showing people" taken by the built-in cameras 3 of the signage 2 in all stores should be divided into.

[0049] 9 is a diagram showing examples of captured images included in each group when the captured images included in the "group of captured images not showing people" ("group of captured images of a store") are grouped into five "groups of captured images A1 to A5 not showing people" (groups 1 to 5 of captured images) in the process of S7 in FIG. 6 above. By grouping the captured images in S7 above, folders corresponding to the groups 1 to 5 of captured images A1 to A5 are automatically generated, and the captured images of the group corresponding to that folder are stored in each of those folders.

[0050] In the example shown in FIG. 9, the group of captured images of group 1 (group of captured images A1) are images of an area in a store with an aisle in the middle and product shelves and walls on either side of the aisle. The group of captured images of group 2 (group of captured images A2) are images of an area in a store with a slightly narrow aisle and product shelves on either side of the aisle. The group of captured images of group 3 (group of captured images A3) are images of an area in a store with a layout where one side of the aisle is a wall and the other side is a product shelf, with the far end of the aisle blocked by product shelves. The group of captured images of group 4 (group of captured images A4) are images of an area in a store taken from diagonally above the aisle by the built-in camera 3 of signage 2 installed in a corner of the store. The images in the group 5 (group A5) are images of an area inside the store taken from diagonally above the aisle by the built-in camera 3 of the signage 2, and are images in which flare occurs. However, the images in groups 1 to 5 shown in FIG. 9 are merely examples.

[0051] As shown in FIG. 9, each group of photographed images A1 to A5 is a collection of photographed images that have similar characteristics and reflect the layout, lighting conditions, interior decoration, etc. of each store.

[0052] Next, the evaluation of the inference accuracy of the trained DNN model 20 for each group after fine tuning, which was explained in S11 of Fig. 6 above, will be described with reference to Fig. 10. Fig. 10 shows inference accuracy evaluation indexes, such as the F1 value (also referred to as "F value") of the trained DNN model 20 for each group corresponding to the fifth camera group, generated by fine tuning the original trained DNN model based on each captured image included in the fifth captured image group B5 (containing a person) and its correct label (pseudo label), in comparison with the inference accuracy evaluation indexes, etc. of the original trained DNN model before fine tuning. In Figure 10, TP (True Positive) represents "something that was predicted to be true and was actually true" (for example, something that was predicted to be a human head and was actually a human head), FP (False Positive) represents "something that was predicted to be true and was actually false" (for example, something that was predicted to be a human head and was actually not a human head), and FN (False Negative) represents "something that was predicted to be false but was actually true" (for example, something that was predicted not to be a human head and was actually a human head).

[0053] Furthermore, Precision in Figure 10 is the so-called precision rate, which indicates the proportion of data predicted to be correct that is actually correct. Expressed as a formula, Precision = TP / (TP + FP). Recall is the so-called recall rate, which indicates the proportion of data predicted to be correct that is actually correct. Expressed as a formula, Recall = TP / (TP + FN). Furthermore, the F1 value (F-value) is the harmonic mean of Recall (recall rate) and Precision (precision rate), and expressed as a formula, F1 value = (2 × Precision × Recall) / (Precision + Recall).

[0054] The table shown in Figure 10 shows that, compared to the original trained DNN model before fine-tuning, the group-specific trained DNN model 20 after fine-tuning has a significantly improved TP value, which is an index that contributes to the detection rate, and a significantly reduced FN value, which is an index of non-detection. Therefore, compared to the original trained DNN model before fine-tuning, the group-specific trained DNN model 20 after fine-tuning has a significantly improved inference accuracy evaluation indexes, such as Precision, Recall, and F1 value. Note that, in the table shown in Figure 10, the value of FP, which is an index of false positives, increases slightly after fine-tuning, but this does not pose any particular problem from the perspective of practical operation, as evaluation indexes such as Precision, Recall, and F1 value have improved significantly. Furthermore, the table shown in Figure 10 shows the improvement in evaluation indicators such as the F1 value after fine-tuning for the group-specific trained DNN model 20 corresponding to the fifth camera group. However, the evaluation indicators such as the F1 value after fine-tuning also significantly improved for the group-specific trained DNN models 20 corresponding to the first to fourth camera groups, which were generated together with the group-specific trained DNN model 20 corresponding to the fifth camera group.

[0055] In the explanation of Figure 8 above, we mentioned that if the number of clusters (number of (camera) groups) is too large, the number of fine-tuning processes for the original trained DNN model increases. Therefore, it is desirable to adopt as small a number of clusters k as possible if the BIC value, which is an index of the optimal model, does not change significantly even if the number of clusters k is increased. This also applies to inference accuracy evaluation indicators such as the precision, recall, and F1 value. That is, when the image clustering unit 35 increases the number of clusters (number of groups) from (k-1) to k, the values ​​of inference accuracy evaluation indicators such as the BIC value (calculated using a Gaussian mixture model) and the F1 value improve significantly. However, if the inference accuracy evaluation indicators such as the BIC value and the F1 value do not change significantly even if the number of clusters is increased from k to (k+1) or more, k is adopted as the number of clusters (number of groups). The reason for this is that if the number of clusters (number of (camera) groups) is too large, the number of fine-tuning processes for the original trained DNN model increases. That is, the image clustering unit 35 adopts a value of the number of clusters k that strikes a balance between the increase in the value of the evaluation index (decrease in the case of BIC) when the number of clusters (number of groups) is increased and the (smallness of) the number of clusters (number of groups). Note that as the index for determining the number of clusters (number of groups), only the BIC value, which is the index of the optimal model described in the description of FIG. 8, may be used, or only inference accuracy evaluation indexes such as Precision, Recall, and F1 value may be used, or a combination of the BIC value and inference accuracy evaluation indexes such as the F1 value may be used.

[0056] As described above, according to the group-specific model generation system 10, the signage learning management server 1, and the group-specific model generation program 27 of this embodiment, captured images collected from each of the built-in cameras 3 of the signage 2 installed in multiple stores are grouped using a Gaussian mixture model based on the feature vectors of each of these captured images. Based on the grouping results of these captured images, the built-in cameras 3 that captured these captured images are grouped. The images captured by the built-in cameras 3 of each group are used to fine-tune the original trained DNN model (for customer detection or recognition). This makes it possible to generate group-specific trained DNN models 20 that are suitable for the images captured by the built-in cameras 3 of each group (specialized for the images captured by the built-in cameras 3 of each group). Therefore, even if each group-specific trained DNN model 20 is an extremely lightweight trained DNN model, it is possible to perform highly accurate customer detection processing and customer recognition processing on the images captured by the built-in cameras 3 of each group. Furthermore, even if the captured images that are the subject of customer detection processing and customer recognition processing by all signage 2 in the group-based model generation system 10 are images captured by built-in cameras 3 of signage 2 installed in a large number of stores, such as several thousand stores, these built-in cameras 3 can be grouped, and the images captured by the limited number of built-in cameras 3 (e.g., several hundred built-in cameras 3) after grouping can be used to fine-tune the original trained DNN model.Therefore, even if the original trained DNN model is an extremely light trained DNN model, the possibility of being able to perform appropriate machine learning can be increased (the possibility of not being able to learn can be reduced). Therefore, even if the captured images that are the subject of customer detection processing and customer recognition processing by all signage 2 in the group-specific model generation system 10 are images captured by the built-in cameras 3 of signage 2 installed in a large number of stores, such as several thousand stores, and even if the original trained DNN model and the generated group-specific trained DNN models 20 are extremely light trained DNN models, it is possible to perform highly accurate customer detection processing and customer recognition processing on the images captured by the built-in cameras 3 of each group using the generated group-specific trained DNN models 20.

[0057] Furthermore, according to the group-specific model generation system 10 of this embodiment, the group-specific trained DNN model 20 suitable for the images captured by the built-in cameras 3 of each group, generated by the group-specific model generation unit 41, is transmitted to and stored in an edge-side device arranged in a store where the built-in cameras 3 of each group are installed, i.e., the signage 2 having the corresponding built-in camera 3, and customer detection processing or customer recognition processing is performed on the images captured by the built-in cameras 3 of each group by this signage 2. This allows the signage 2 having the built-in cameras 3 of each group to perform highly accurate customer detection processing or customer recognition processing on the images captured by its own built-in cameras 3.

[0058] Furthermore, according to the group-specific model generation system 10 of this embodiment, the trained high-precision DNN model 40 for customer detection or recognition, which is capable of performing inference with higher accuracy than the trained DNN model for the original customer detection or recognition, performs inference on images captured by the built-in camera 3 of each group. Pseudo labels based on this inference result are assigned as correct labels to the images captured by the built-in camera 3 of each group. Then, the trained DNN model for customer detection or recognition is fine-tuned based on the images captured by the built-in camera 3 of each group and the correct labels (pseudo labels) assigned to the images captured by the built-in camera 3 of each group. This automatically assigns correct labels to each of the images captured by the built-in camera 3 of each group, allowing for automatic fine-tuning of the trained DNN model. In other words, the original trained DNN model can be fine-tuned without manual annotation (creating correct labels for each captured image).

[0059] Furthermore, according to the group-specific model generation system 10 of this embodiment, while changing the number of clusters, which is the number of groups of photographed images, the BIC (Bayes Information Criterion) value for each number of clusters is calculated using a Gaussian mixture model, and based on the BIC value corresponding to each calculated number of clusters, the number of clusters appropriate for the distribution of feature vectors of photographed images extracted by the image feature vector extraction unit 34 is calculated. This makes it possible to automatically calculate the number of clusters appropriate for the distribution of feature vectors of photographed images.

[0060] Furthermore, according to the group-specific model generation system 10 of this embodiment, a feature vector is extracted from each of the captured images of the store that remain after removing captured images that show people from the captured images collected from each of the built-in cameras 3 of the signage 2 installed in multiple stores, and the captured images of the store are grouped based on these feature vectors using a Gaussian mixture model, which is unsupervised learning. As described above, by grouping the captured images that are the basis for grouping the built-in cameras 3 based on the feature vectors of the captured images of the store, the captured images of the built-in cameras 3 can be grouped without being affected by people that appear in the captured images.

[0061] Variations: The present invention is not limited to the configurations of the above-described embodiments, and various modifications are possible within the scope of the invention. Next, modifications of the present invention will be described.

[0062] Variation 1: In the above embodiment, an example has been shown in which the image clustering unit 35 groups the group of photographed images of a store using a Gaussian mixture model based on the feature vectors of each photographed image extracted by the image feature vector extraction unit 34. However, the clustering model used to group the group of photographed images is not limited to a Gaussian mixture model, and may be, for example, an unsupervised learning model such as a k-means algorithm or an expectation-maximization (EM) algorithm. Furthermore, the group of photographed images of a store does not necessarily have to be grouped based on the feature vectors of each photographed image as described above, and the group of photographed images may be grouped based on various features of each photographed image.

[0063] Variation 2: In the above embodiment, an example is shown in which the group-specific trained DNN model 20 suitable for the images captured by the built-in camera 3 of each group is generated by fine-tuning the original trained DNN model using the images captured by the built-in camera 3 of each group and the pseudo labels assigned to these images by the pseudo-labeling unit 39. However, the group-specific trained DNN model suitable for the images captured by the built-in camera 3 of each group may be generated by performing transfer learning of the original trained DNN model using the images captured by the built-in camera 3 of each group and the pseudo labels assigned to these images. Here, transfer learning means learning only the weights of the newly added layer while keeping the weights in the original (existing) trained DNN model fixed.

[0064] Variation 3: In the above embodiment, an example is shown in which the group-specific trained DNN models 20 suitable for images captured by the built-in cameras 3 of each group are transmitted to and stored in the signage 2 having the built-in cameras 3 of each group. However, the device that transmits and stores (installs) the group-specific trained DNN models is not limited to signage, and may be an edge-side device installed in a facility such as a store where some kind of camera is installed. Examples of such edge-side devices include an image analysis device that performs object detection or object recognition on images captured by a surveillance camera, and a so-called AI camera.

[0065] Variation 4: In the above embodiment, an example is shown in which the number of clusters, which is the number of groups of captured images, is changed, and the BIC (Bayes Information Criterion) value for each number of clusters is calculated using a Gaussian mixture model, and the number of clusters appropriate for the distribution of feature vectors of the captured images is calculated based on the BIC value corresponding to each calculated number of clusters.However, for example, the AIC (Akaike Information Criterion) value for each number of clusters may be calculated using unsupervised learning such as a Gaussian mixture model, and the number of clusters appropriate for the distribution of feature vectors of the captured images may be calculated based on the AIC value corresponding to each calculated number of clusters.

[0066] Variation 5: In the above embodiment, the group-specific trained DNN model 20 generated by the group-specific model generation unit 41 is a trained DNN model for detecting or recognizing customers. Therefore, the person image extraction unit 38 is used to extract photographed images in which people are reflected, and the extracted photographed images in which people are reflected ("photographed images B1 to B2 in which people are reflected") are used. k Each of the captured image groups B1 to B k) was used to fine-tune the original trained DNN model, thereby generating a group-specific trained DNN model 20 suitable for images captured by the built-in cameras 3 of each group. However, for example, if the group-specific trained DNN model generated by the group-specific model generation unit is a trained DNN model for detecting or recognizing products, or a trained DNN model for detecting or recognizing product shelves, a group-specific trained DNN model suitable for images captured by the built-in cameras of each group can be generated by fine-tuning the original (existing) trained DNN model using a "group of captured images not showing people" captured by the k built-in cameras of each group.

[0067] Variation 6: Furthermore, in the above embodiment, an example has been shown in which signage learning management server 1 is equipped with frame image extraction unit 32 and human image removal unit 33, but each signage may be equipped with functions equivalent to the frame image extraction unit and human image removal unit, and may transmit only captured images (frame images) that do not show people to signage learning management server 1. In this case, the captured image collection unit on the signage learning management server side collects the captured images (frame images) that do not show people from each of the built-in cameras of signage installed in multiple stores. [Explanation of symbols]

[0068] 1 Signage learning management server (server, computer) 2, 2a, 2b Signage (edge ​​device) 3 Built-in camera (photography means) 10 Group-specific model generation system 20 Group-specific trained DNN models (Group-specific trained neural network models) 27 Group-specific model generation program 31 Photographed image collection unit (photographed image collection means) 33 Human image removal unit (human image removal means) 34 Image feature vector extraction unit (image feature extraction means) 35 Image clustering unit (image clustering means) 36 Camera classification unit (photography means classification means) 38 Person image extraction unit (person image extraction means) 39 Pseudo-labeling unit (pseudo-labeling means) 40 Trained High-Precision DNN Models (Trained High-Precision Neural Network Models) 41 Group-specific model generation unit (group-specific model generation means) S, Sa, Sb Stores (facilities)

Claims

1. a captured image collecting means for collecting captured images from each of the image capturing means installed in a plurality of facilities; a person image removing means for removing photographed images in which people are reflected from the photographed images collected by the photographed image collecting means; an image feature extraction means for extracting features from each of the photographed images of the facility that remain after the photographed images in which the person is captured are removed by the human image removal means; an image clustering means for grouping the photographed images of the facility based on the features of each of the photographed images of the facility extracted by the image feature extraction means; an image capturing means classification means for classifying the captured images of the facility into groups based on the grouping results of the image clustering means; a group-specific model generation means for generating a group-specific trained neural network model suitable for images captured by the image capture means of each group grouped by the image capture means classification means by performing fine tuning or transfer learning of an original trained neural network model for person detection or person recognition, The image capturing device further includes a person image extracting means for extracting images in which people are captured from the images captured by the image capturing means of each group after the grouping by the image capturing means classifying means, The group-specific model generation means is a group-specific model generation system that performs fine tuning or transfer learning of a trained neural network model for detecting or recognizing people using photographic images containing people extracted by the human image extraction means.

2. The group-specific model generation system according to claim 1, characterized in that the group-specific trained neural network model generated by the group-specific model generation means and suitable for the images captured by the imaging means of each group is transmitted to and stored in an edge-side device located in a facility where the imaging means of each group is installed, and this edge-side device performs person detection or person recognition on the images captured by the imaging means of each group.

3. The system further comprises a pseudo-labeling means for performing inference on images captured by the image capturing means of each group using a trained high-precision neural network model for human detection or human recognition that is capable of performing inference with higher accuracy than the original trained neural network model for human detection or human recognition, and assigning pseudo-labels based on the inference results to the images captured by the image capturing means of each group as correct labels; 3. The group-specific model generation system according to claim 1, wherein the group-specific model generation means performs fine tuning or transfer learning of the original trained neural network model for person detection or person recognition based on the images captured by the imaging means of each group and the correct labels assigned to the images captured by the imaging means of each group by the pseudo-labeling means.

4. 4. The group-based model generation system according to claim 1, wherein the image clustering means calculates an information criterion value for each number of clusters while changing the number of clusters, which is the number of groups of the photographed images, and calculates a number of clusters appropriate for the distribution of features of the photographed images extracted by the image feature extraction means, based on the information criterion value corresponding to each calculated number of clusters.

5. the image feature extraction means extracts a feature vector from each of the photographed images of the facility that remain after the photographed images in which the people are captured are removed by the human image removal means, using a trained neural network model; 5. The group-based model generation system according to claim 1, wherein the image clustering means groups the photographed images of the facility using a Gaussian mixture model based on the feature vectors of each of the photographed images of the facility extracted by the image feature extraction means.

6. 6. The group-based model generation system according to claim 5, wherein the image clustering means, while changing the number of clusters, which is the number of groups of the photographed images, calculates a value of Bayes Information Criterion for each number of clusters using the Gaussian mixture model, and calculates a number of clusters appropriate for the distribution of feature vectors of the photographed images of the facility extracted by the image feature extraction means, based on the value of Bayes Information Criterion corresponding to each calculated number of clusters.

7. The edge device is connected via a network to an edge device disposed in each of a plurality of facilities where the imaging means is installed, a captured image collecting means for collecting captured images from each of the imaging means; a person image removing means for removing photographed images in which people are reflected from the photographed images collected by the photographed image collecting means; an image feature extraction means for extracting features from each of the photographed images of the facility that remain after the photographed images in which the person is captured are removed by the human image removal means; an image clustering means for grouping the photographed images of the facility based on the features of each of the photographed images of the facility extracted by the image feature extraction means; an image capturing means classification means for classifying the captured images of the facility into groups based on the grouping results of the image clustering means; a group-specific model generation means for generating a group-specific trained neural network model suitable for images captured by the image capture means of each group grouped by the image capture means classification means by performing fine tuning or transfer learning of an original trained neural network model for person detection or person recognition, The image capturing device further includes a person image extracting means for extracting images in which people are captured from the images captured by the image capturing means of each group after the grouping by the image capturing means classifying means, The group-specific model generation means is a server that performs fine tuning or transfer learning of the trained neural network model for detecting or recognizing people using photographic images containing people extracted by the human image extraction means.

8. The server according to claim 7, characterized in that the group-specific trained neural network model generated by the group-specific model generation means and suitable for the images captured by the imaging means of each group is transmitted to and stored in an edge-side device located in a facility where the imaging means of each group is installed.

9. Computer, a captured image collecting means for collecting captured images from each of the image capturing means installed in a plurality of facilities; a person image removing means for removing photographed images in which people are reflected from the photographed images collected by the photographed image collecting means; an image feature extraction means for extracting features from each of the photographed images of the facility that remain after the photographed images in which the person is captured are removed by the human image removal means; an image clustering means for grouping the photographed images of the facility based on the features of each of the photographed images of the facility extracted by the image feature extraction means; an image capturing means classification means for classifying the captured images of the facility into groups based on the grouping results of the image clustering means; a group-specific model generation program for functioning as group-specific model generation means for generating group-specific trained neural network models suitable for images captured by the image capture means of each group grouped by the image capture means classification means by performing fine tuning or transfer learning of an original trained neural network model for person detection or person recognition using images captured by the image capture means of each group grouped by the image capture means classification means, causing the computer to further function as a person image extraction unit that extracts images in which people are captured from the images captured by the image capture unit of each group after the grouping by the image capture unit classification unit; The group-specific model generation means is a group-specific model generation program that performs fine tuning or transfer learning of a trained neural network model for detecting or recognizing people using photographic images containing people extracted by the human image extraction means.

Citation Information

Patent Citations

  • Variable heat insulating house

    JP1986078942A

  • Image recognition device, learning device, image recognition method, learning method and program

    JP2019012426A

  • Image analysis device and image analysis system

    JP2020181488A

  • Information management device, information processing device, control method, and program

    JP2020197995A

  • Classification unit, generation unit, dataset generation device, frame image classification method, and frame image classification program

    JP2022029125A