Model generation device, model generation method, program

The model generation device and method address accuracy issues in object detection by aligning and training the model across varying image angles, enhancing detection precision.

JP7852713B2Active Publication Date: 2026-04-28NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2022-06-02
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing object detection systems face accuracy issues due to varying shooting angles in image capture, leading to misrecognition or omission of objects, especially when furniture information is unavailable.

Method used

A model generation device and method that utilize an object detection model to detect object positions in multiple images with different fields of view, applying geometric deformation parameters to align and train the model, enhancing accuracy across varying angles.

Benefits of technology

The solution effectively suppresses accuracy decreases in object detection caused by shooting angle variations, ensuring precise object localization even in images with different viewpoints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007852713000001
    Figure 0007852713000001
  • Figure 0007852713000002
    Figure 0007852713000002
  • Figure 0007852713000003
    Figure 0007852713000003
Patent Text Reader

Abstract

A model generation device 200 according to the present disclosure comprises: a detection means 201 that detects, using an object detection model, a first position that is an object position in a first image and a second position that is an object position in a second image that is different in angular field from the first image; a generation means 202 that generates from the first position a corresponding position that corresponds, in the second image, to the first position, on the basis of the difference in angular field between the first and second images; and a training means 203 that trains the object detection model on the basis of the second position and the corresponding position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a model generation device that generates a model for detecting an object included in an image.

Background Art

[0002] Techniques for detecting an object from a photographed image of an object are known. For example, as described in Patent Document 1, a system has been proposed that photographs an image of a product shelf in a store and identifies the product position to perform shelf division analysis. In such a system, an object detection model for identifying the product position is learned in advance using a large number of photographed images of product shelves. During operation, the learned object detection model is used to identify the positions of products included in the images of product shelves photographed at each store.

[0003] Here, in Patent Document 1, it is pointed out as a problem that the recognition accuracy decreases, such as misrecognition or omission of recognition of products, because the photographed images of the shelves on which products are displayed are affected by the environment such as the shooting angle at the time of shooting. In response to such a problem, Patent Document 1 describes a method of detecting an area where there is a high possibility of omission of recognition using furniture information.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the method described in Patent Document 1 mentioned above, it is necessary to memorize furniture information in advance, and when such information is not available, omission of recognition of products cannot be detected. Therefore, there still arises a problem that the accuracy of object detection in the image decreases due to the shooting angle of the image.

[0006] The purpose of this disclosure is to solve the aforementioned problem, which is that the accuracy of object detection within an image decreases depending on the image's field of view. [Means for solving the problem]

[0007] A model generation device, which is one form of this disclosure, A detection means that uses an object detection model to detect a first position, which is the position of an object in the first image, and a second position, which is the position of an object in the second image, which has a different field of view from the first image. A generation means that generates a corresponding position which is the position corresponding to the first position within the second image, based on the difference in field of view between the first image and the second image, A learning means for learning the object detection model based on the second position and the corresponding position, Equipped with, This is the structure it takes.

[0008] Furthermore, a model generation method, which is one form of this disclosure, Using an object detection model, the first position, which is the position of the object in the first image, and the second position, which is the position of the object in the second image, which has a different field of view from the first image, are detected. Before or after detecting the second position, a corresponding position is generated from the first position, which is the position within the second image corresponding to the first position, based on the difference in field of view between the first image and the second image. Based on the second position and the corresponding position, the object detection model is trained. This is the structure it takes.

[0009] Furthermore, one form of this disclosure is a program, Using an object detection model, the first position, which is the position of the object in the first image, and the second position, which is the position of the object in the second image, which has a different field of view from the first image, are detected. Based on the difference in field of view between the first image and the second image, a corresponding position is generated from the first position, which is the position corresponding to the first position within the second image. Based on the second position and the corresponding position, the object detection model is trained. The computer is caused to execute the processing. It has the following configuration.

Advantages of the Invention

[0010] According to the present disclosure configured as described above, it is possible to suppress the decrease in the accuracy of object detection in an image due to the shooting angle of the image.

Brief Description of the Drawings

[0011] [Figure 1] It is a diagram showing the overall configuration of the object detection system in the first embodiment of the present disclosure. [Figure 2] It is a diagram showing an example of a store environment where the object detection device disclosed in FIG. 1 is used. [Figure 3] It is a block diagram showing the hardware configuration of the object detection device disclosed in FIG. 1. [Figure 4] It is a block diagram showing the configuration of the object detection device disclosed in FIG. 1. [Figure 5] It is a diagram showing the state of image processing by the object detection device disclosed in FIG. 1. [Figure 6] It is a diagram showing the state of image processing by the object detection device disclosed in FIG. 1. [Figure 7] It is a flowchart showing the operation during the learning process of the object detection model by the object detection device disclosed in FIG. 1. [Figure 8] It is a block diagram showing the configuration of the model generation device in the second embodiment of the present disclosure.

Modes for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. <First Embodiment> [Overall Configuration] FIG. 1 shows the overall configuration of an object detection system according to the first embodiment. As shown in FIG. 1, the object detection system includes an object detection device 100 and an image database (hereinafter, the "database" is referred to as "DB") 2. The object detection device 100 acquires image data from the image DB 2 and performs object detection. Accordingly, the object detection device 100 of the present disclosure also has a function as a model generation device that generates an object detection model used when performing object detection by learning. When the object detection model of the object detection device 100 is learned, a learning dataset is stored in the image DB 2. On the other hand, when the object detection device 100 is applied and used in an actual store or the like, that is, at the time of inference for detecting an object from an image, an image taken in the store is stored in the image DB 2.

[0013] [Example of store environment] FIG. 2 shows an example of a store environment in which the object detection device 100 is used. In the store, a product shelf 3 is installed, and various products are displayed on the product shelf 3. A surveillance camera 4 is installed in the store and is photographing the product shelf 3. The image taken by the surveillance camera 4 is sent to a terminal device 6 and stored in the image DB 2 connected to the terminal device 6. In front of the product shelf 3, a store clerk is using the camera 5 of a mobile terminal to take a front image of the product shelf 3. The image taken by the camera 5 of the mobile terminal is sent to the terminal device 6 and recorded in the image DB 2. The object detection device 100 is realized by, for example, the terminal device 6 or another terminal device.

[0014] [Hardware configuration] FIG. 3 is a block diagram showing the hardware configuration of the object detection device 100. As shown in the figure, the object detection device 100 includes a communication unit 101, a processor 102, a memory 103, and a recording medium 104.

[0015] The communication unit 101 communicates with the image DB3 via wired or wireless connection to acquire pre-prepared training datasets and images captured by the store's camera 4. The processor 102 is a computer such as a CPU (Central Processing Unit) and controls the entire object detection device 100 by executing pre-prepared programs. The processor 102 may be a GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating point number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof. Specifically, the processor 102 performs the pre-training process and additional training process described later.

[0016] Memory 103 consists of ROM (Read Only Memory), RAM (Random Access Memory), and other components. Memory 103 is also used as working memory while the processor 102 is executing various processes.

[0017] The recording medium 104 is a non-volatile, non-temporary recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from the object detection device 100. The recording medium 104 stores various programs that the processor 102 executes. When the object detection device 100 performs various processes, the programs stored in the recording medium 104 are loaded into the memory 103 and executed by the processor 102.

[0018] [Configuration of the object detection device] Next, the configuration of the object detection device 100 will be described. As shown in Figure 4, the object detection device 100 connected to the image DB2 described above comprises an object position estimation unit 20, a loss calculation unit 30, a geometric deformation estimation unit 50, and an automatic correct answer assignment unit 60. The object position estimation unit 20, the loss calculation unit 30, the geometric deformation estimation unit 50, and the automatic correct answer assignment unit 60 can be realized by the processor 102 executing a program stored in the memory 103 or recording medium 104. In this case, the geometric deformation estimation unit 50 further comprises a feature point extraction unit 51, a feature point matching degree calculation unit 52, and a geometric deformation parameter calculation unit 53. The automatic correct answer assignment unit 60 further comprises a detection result transfer unit 61 and a correct rectangle generation unit 62. Each configuration will be described in detail below.

[0019] First, the pre-training function of the object position estimation unit 20 and the loss calculation unit 30 will be explained. Pre-training is a process that first generates a basic object detection model. For this reason, the image DB2 first stores the pre-training image dataset used in pre-training. Specifically, the pre-training image dataset includes pre-training images of a pre-prepared product shelf and ground truth data for product positions. For example, the pre-training images are images of a product shelf taken from the front, and the ground truth data is indicated by the coordinates of the vertices of rectangles that show the positions of objects included in each pre-training image.

[0020] The object position estimation unit 20 detects objects contained in the input image using an object detection model. Specifically, during pre-training, the object position estimation unit 20 uses the object detection model to estimate the rectangular coordinates indicating the positions of objects contained in the pre-training images input from the image DB2. The object detection model is composed of a neural network, such as a CNN (Convolutional Neural Network). The object position estimation unit 20 outputs the rectangular coordinates estimated from the pre-training images input from the image DB2, and the ground truth data of the object positions contained in the pre-training images associated with the input pre-training images, to the loss calculation unit 30.

[0021] The loss calculation unit 30 calculates the loss using the input ground truth data of object positions and the estimation results of the object position estimation unit 20. Specifically, the loss calculation unit 30 calculates the error between the rectangular coordinates of the object positions included in the input ground truth data and the rectangular coordinates of the object positions in the pre-training images estimated by the object position estimation unit 20, and uses this as the loss. The loss calculation unit 30 then updates the parameters of the object detection model of the object position estimation unit 20 so that the calculated loss becomes smaller. The parameters of the object detection model are updated in this way until the value of the loss converges to below a predetermined level, and the pre-training of the object detection model is completed when the value of the loss converges. The object detection model at the time the training is completed is obtained as the pre-trained object detection model.

[0022] As mentioned above, the object detection model generated by pre-training is trained primarily on pre-training images of product shelves taken from the front. Therefore, while the object detection accuracy for images of product shelves taken from the front is high, the accuracy for object detection in images with different angles of view from the front image of the product shelves, such as images taken from the store's surveillance camera 4 as shown in Figure 2, is expected to decrease. For this reason, the object detection device 100 of this disclosure has a function to further train the object detection model so that the accuracy of object detection is also high for images with different angles of view from the front image, such as images taken by the surveillance camera 4. The object detection model generated by the pre-training described above is not necessarily limited to being generated by the object detection device 100; it may also be generated by another device or be one that has been prepared in advance. The configuration for performing additional training by the object detection device 100 will be described below.

[0023] Image DB2 includes, for additional training purposes, two image pairs (image pairs) of the same product shelf 3, i.e., the same object, taken simultaneously in an area where there is no product movement, by the surveillance camera 4 and the mobile device camera 5, respectively. Here, the mobile device camera 5 takes an image of the product shelf 3 from the front, and such an image is referred to as the "front image" (first image). However, the front image is not limited to an image of the product shelf 3 taken strictly from the front, but can be an image taken from approximately the front. Also, the surveillance camera 4 is installed, for example, on the ceiling or wall of the store, and the field of view of the image taken by the surveillance camera 4 will be different from that of the front image. The image taken by the surveillance camera 4 is referred to as the "surveillance camera image" (second image). Note that the front image described in this embodiment is not necessarily limited to an image of the product shelf 3 taken from the front by the mobile device camera 5, but may be an image taken from any direction with any shooting device. Also, the surveillance camera image is not necessarily limited to an image taken by the surveillance camera 4, but may be an image taken from any direction with any shooting device. However, the first image, which corresponds to the front view, and the second image, which corresponds to the surveillance camera image, are images with different fields of view.

[0024] The geometric deformation estimation unit 50 (estimation means) estimates geometric deformation parameters between two images using the above-mentioned pair of additional learning images included in the image DB2, namely the pair of frontal image and surveillance camera image. In particular, in this embodiment, the geometric deformation estimation unit 50 estimates affine transformation parameters to match the field of view of the mobile device's camera 5 to the field of view of the surveillance camera 4.

[0025] Specifically, the feature point extraction unit 51 extracts feature points from each of the two input images: the surveillance camera image and the front view image. The extracted feature points, with their coordinate values ​​and feature quantities stored as vectors, are input to the feature point matching degree calculation unit 52.

[0026] The feature point matching calculation unit 52 calculates the similarity between feature points of the two images extracted by the feature point extraction unit 51 and outputs pairs of feature points with high similarity. For example, for each feature point in the front view image, the cosine similarity with all feature points in the surveillance camera image is calculated, and the point with the highest similarity among the points whose similarity exceeds a predetermined value is selected as the pair. In other words, the feature point matching calculation unit 52 extracts pairs of feature points in the front view image (first feature point) and the corresponding feature point in the surveillance camera image (second feature point). The feature point matching calculation unit 52 then outputs the coordinate values ​​of each point in the selected pair to the geometric deformation parameter calculation unit 53.

[0027] The geometric deformation parameter calculation unit 53 uses the coordinates of the matching feature point pairs between the two images selected by the feature point matching degree calculation unit 52 to calculate affine transformation parameters to match the field of view of the front image to the field of view of the surveillance camera image. Specifically, for each pair of feature points, the coordinates of the feature points in the front image are affine transformed, the error with the coordinates of the feature points in the surveillance camera image is calculated, and the affine transformation parameters are determined so that the sum of the errors for each pair of feature points is small. The affine transformation parameters thus obtained are output as geometric deformation parameters to the detection result transfer unit 61.

[0028] The above method for calculating geometric deformation parameters is just one example; any method for associating identical points between two image pairs is acceptable. For example, if the installation positions and field of view of two cameras are given as metadata, the identical points can be estimated by analytically calculating the transformation of the field of view between those cameras.

[0029] The object position estimation unit 20 (detection means) estimates the position of objects contained in the input image using a pre-trained object detection model. In other words, during additional training, the object position estimation unit 20 receives a pair of front-facing images and surveillance camera images as input and estimates the position of objects contained in each image. In this case, for the front-facing image, a pre-trained object detection model generated by learning images with similar field of view in the pre-training is used, so the position of the object can be estimated with high accuracy. On the other hand, for the surveillance camera image, images with similar field of view were not used in the pre-training, so the accuracy of object detection by the pre-trained object detection model decreases. The object position estimation unit 20 outputs the front-facing image coordinates (first position), which represent the position of the product in the front-facing image and are the estimation results for the object position of the front-facing image, to the automatic correction unit 60, and outputs the surveillance camera image coordinates (second position), which represent the position of the product in the surveillance camera image and are the estimation results for the surveillance camera image, to the loss calculation unit 30.

[0030] The automatic correction unit 60 (generation means) uses the geometric deformation parameters, which are the output of the geometric deformation estimation unit 50, and the rectangular coordinates, which are the output of the object position estimation unit 20, to transfer the front image coordinates, which are the estimation results of the object position estimation unit 20 for the front image, to the coordinates in the surveillance camera image. Specifically, the rectangular coordinates, which are the front image coordinates representing the position of the object estimated by the object position estimation unit 20, are transformed according to the geometric deformation parameters estimated by the geometric deformation estimation unit 50 to obtain transformed coordinates (corresponding positions) that correspond to the position of the object included in the surveillance camera image.

[0031] Specifically, the detection result transfer unit 61 transforms the position of the object included in the front image estimated by the object position estimation unit 20 using affine transformation parameters calculated by the geometric deformation parameter calculation unit 53 to calculate the corresponding position of the object on the surveillance camera image. In this embodiment, the detection result transfer unit 61 transforms the coordinates of four points of a rectangle, which is the front image coordinates indicating the position of the object included in the front image estimated by the object position estimation unit 20, using affine transformation parameters calculated by the geometric deformation parameter calculation unit 53, and outputs the coordinate values ​​of the four points, which are the transformed coordinates obtained by the transformation, to the correct rectangle generation unit 62.

[0032] The ground truth rectangle generation unit 62 further converts the transformed coordinates of the object positions on the surveillance camera image calculated by the detection result transfer unit 61 into rectangle coordinates for training the object detection model. For example, the ground truth rectangle generation unit 62 calculates the smallest rectangle enclosing the four points that are the transformed coordinates representing the object positions calculated by the detection result transfer unit 61, and outputs the coordinates of the four points that are the vertices of this smallest rectangle as the ground truth rectangle (position information) to the loss calculation unit 30.

[0033] Here, Figure 5 shows an example of a pair of front view images P1 and surveillance camera images P2. The pair of front view images P1 and surveillance camera images P2 are images of the same product shelf taken at the same time within a range where the objects are not moving, so the same products are displayed in the same arrangement on the product shelf in both images. The rectangle surrounding the products shown in front view image P1 is an example of front view image coordinates that represent the position of the products in the front view image, which is the output result of the object position estimation unit 20 and detected from front view image P1. Surveillance camera image P2 is an image of the product shelf taken at a different field of view than the front view image. Note that the surveillance camera image coordinates that represent the position of the products in the surveillance camera image, which are the output result of the object position estimation unit 20 and detected from surveillance camera image P2, are not shown in the surveillance camera image P2.

[0034] Then, using the pair of frontal images and surveillance camera images described above, the geometric deformation estimation unit 50 estimates geometric deformation parameters so that the field of view of both images matches. The automatic correction unit 60 then uses the geometric deformation parameters, which are the estimation results of the geometric deformation estimation unit 50, to convert the frontal image coordinates, which represent the position of the product in the frontal image and are the estimation results of the object position estimation unit 20 shown in the frontal image P1, into coordinates on the surveillance camera image P2.

[0035] Figure 6 shows the transformation of the rectangular coordinates representing the object position by the automatic correct answering unit 60 described above. In Figure 6, the symbol P11 represents one of the products captured in the front image P1 and the product position rectangle estimated by the object position estimation unit 20. In Figure 6, the symbol P12 is a dotted line representing the rectangle obtained when the product position rectangle in the front image shown in symbol P11 is transformed by geometric deformation parameters in the detection result transfer unit 61 and transferred onto the surveillance camera image P2. At this time, the dotted rectangle shown in symbol P12 is not suitable as input to the loss calculation unit 30 described later, so the correct rectangle generation unit 62 described above converts it into the smallest rectangle that includes this dotted rectangle, that is, a solid rectangle as shown in symbol 13. The coordinates of the solid rectangle shown in symbol 13 are output to the loss calculation unit 30 as the correct rectangle (correct data).

[0036] The loss calculation unit 30 (learning means) calculates the loss using the surveillance camera image coordinates (second position) representing the position of the product in the surveillance camera image, which is the output of the object position estimation unit 20, and the ground truth rectangle (position information based on the corresponding position), which is the output of the automatic ground truth assignment unit 60. It then updates the parameters of the object detection model of the object position estimation unit 20 in the same way as during pre-training and performs learning. Specifically, it calculates the error between the rectangle coordinates, which are the surveillance camera image coordinates that are the estimated result of the object detection model for the surveillance camera image, and the ground truth rectangle coordinates of the object position calculated by the automatic ground truth assignment unit 60, and uses this as the loss. The loss calculation unit 30 then updates the parameters of the object detection model to reduce the loss. The parameters of the object detection model are updated until the loss value converges to below a predetermined value, and pre-training of the object detection model is completed when the loss value converges. The object detection model at the time the learning is completed is obtained as the trained object detection model.

[0037] As described above, the trained object detection model obtained is later used for object detection from images to be inferred. Specifically, the object position estimation unit 20 takes a surveillance camera image to be inferred as input, uses the trained object detection model to estimate the rectangular coordinates representing the positions of objects contained in the input image, and outputs the result.

[0038] In the above configuration, the object position estimation unit 20 is an example of a detection means, the geometric deformation estimation unit 50 and the automatic correct answer assignment unit 60 are examples of generation means, and the loss calculation unit 30 is an example of a learning means.

[0039] [Object detection device operation] Next, the operation of the object detection device 100 will be explained. Figure 7 is a flowchart of the learning process for the object detection model, and in particular, it shows the additional learning operation described above. For this reason, it is assumed that the object detection device 100 has undergone the pre-training described above in advance, and that the basic object detection model has been generated. It is also assumed that the image DB2 stores image pairs of frontal images and surveillance camera images taken at the same time in the range where the object is not moving.

[0040] First, the object detection device 100 inputs image pairs of a front view image and a surveillance camera image from the image DB2 to the geometric deformation estimation unit 50 and the object position estimation unit 20 (step S11). The geometric deformation estimation unit 50 estimates the geometric deformation parameters between the two input images and inputs them to the automatic correction unit 60 (step S12). The object position estimation unit 20 estimates the rectangular coordinates representing the position of each object contained in the two images and inputs the estimation results to the automatic correction unit 60 and the loss calculation unit 30 (step S13). Specifically, the object position estimation unit 20 inputs the ground truth image coordinates, which are the estimation results for the front view image, to the automatic correction unit 60, and inputs the surveillance camera image coordinates, which are the estimation results for the surveillance camera image, to the loss calculation unit 30. Note that the process of estimating the object position in step S13 may be performed before the geometric deformation parameter estimation process in step S12.

[0041] Next, the automatic correction unit 60 uses inputs from the geometric deformation estimation unit 50 and the object position estimation unit 20 to calculate the rectangular coordinates of the object positions included in the surveillance camera image and inputs them to the loss calculation unit 30 (step S14). Specifically, the automatic correction unit 60 transforms the rectangular coordinates, which are front image coordinates representing the position of the object estimated from the front image by the object position estimation unit 20, according to the geometric deformation parameters estimated by the geometric deformation estimation unit 50, to obtain transformed coordinates corresponding to the position on the surveillance camera image, and inputs them to the loss calculation unit 30. Note that the process of generating transformed coordinates from front image coordinates in step S14 may be performed before the process of estimating the object position from the surveillance camera image in step S13. In other words, the process of estimating the object position from the surveillance camera image in step S13 may be performed after step S14.

[0042] The loss calculation unit 30 calculates the loss using the rectangular coordinates input from the automatic correct answering unit 60 and the object position estimation unit 20 (step S15). Specifically, the loss calculation unit 30 calculates the loss using the surveillance camera image coordinates representing the position of the product in the surveillance camera image, which are the output of the object position estimation unit 20, and the correct rectangular coordinates, which are the output of the automatic correct answering unit 60. The loss calculation unit 30 then determines whether the loss has converged to a predetermined value or less (step S16). If the loss has not converged (step S16: No), the loss calculation unit 30 updates the parameters of the object detection model that constitutes the object position estimation unit 20 so that the loss becomes smaller (step S17). Then, the process returns to step S11. On the other hand, if the loss has converged (step S16: Yes), the process ends.

[0043] Subsequently, the object detection device 100 can input a surveillance camera image to be inferred, estimate the coordinates representing the positions of objects included in the input image using a trained object detection model, and output the results.

[0044] As described above, the object detection model generation device of the first embodiment can accurately detect objects in images even for new images such as surveillance camera images with different fields of view, by training an object detection model using pairs of images with different fields of view, such as a front view image and a surveillance camera image. At this time, since the object positions in the front view image and the surveillance camera image are automatically assigned as correct answers, it is possible to generate an object detection model that can handle images with new fields of view while keeping the cost of manual correct answering low.

[0045] In the embodiments described above, the case in which the object to be detected is a product displayed on a shelf was used as an example, but the applications of this disclosure are not limited to product detection. For example, it can be applied to situations where training images can be obtained from multiple angles of view during a period in which the position of the object does not change, such as in surveillance cameras for people, detection of abandoned objects, and item monitoring.

[0046] <Embodiment 2> Next, a second embodiment of the present disclosure will be described with reference to Figure 8. Figure 8 is a block diagram showing the configuration of the model generation device in Embodiment 2. In this embodiment, the configuration of the object detection device described in the above-described embodiment is shown in schematic.

[0047] The model generation device 200 in this embodiment is configured as a general information processing device, and for example, it is equipped with a hardware configuration similar to that of the object detection device described in Embodiment 1. In other words, the model generation device comprises a communication unit, a processor, memory, and a recording medium.

[0048] The model generation device 200 can be equipped with the detection means 201, generation means 202, and learning means 203 shown in Figure 8 by having a processor acquire a program stored in memory or a storage medium and execute it. The program may be supplied to the processor via a communication network, or it may be stored in a storage medium beforehand, and a drive device may read the program and supply it to the processor. However, the detection means 201, generation means 202, and learning means 203 described above may be constructed with dedicated electronic circuits for realizing such means.

[0049] The detection means 201 uses an object detection model to detect the first position, which is the position of the object in the first image, and the second position, which is the position of the object in the second image, which has a different field of view from the first image. At this time, the first image and the second image are images of the same object in which the object is located, but have different fields of view from each other. For example, the images are of a product shelf on which objects are displayed. For example, the first image is an image of the product shelf taken from the front, and the second image is a surveillance camera image of the product shelf taken by a surveillance camera installed on the ceiling or elsewhere. The detection means 201 then detects the positions (first position and second position) of the products displayed on the product shelf from the front image and the surveillance camera image, respectively. At this point, the object detection model has been trained mainly using images taken with the field of view of the first image, so the object detection accuracy from the first image is high, while the object detection accuracy from the second image is low.

[0050] The generation means 202 generates a corresponding position, which is the position in the second image corresponding to the first position, based on the difference in field of view between the first image and the second image. For example, the generation means 202 estimates the difference in field of view between the first image and the second image and generates deformation parameters to transform the first image into the second image. Then, the generation means 202 generates a corresponding position by deforming the position of the object in the first image, for example, the position of the object in the front image (first position), using the generated deformation parameters. As a result, a corresponding position on the second image with a different field of view is generated with high accuracy from the first position detected from the first image.

[0051] The learning means 203 described above learns an object detection model based on the second position and the corresponding position. For example, the learning means 203 updates and learns the parameters of the object detection model using the corresponding position as the correct data for the second position. As a result, the object detection model learns to detect the position of an object from a second image, such as a surveillance camera image, so that it approaches the corresponding position.

[0052] As described above, this disclosure enables accurate detection of the position of an object even in a second image with a different field of view from the first image, using the generated object detection model.

[0053] Although the present disclosure has been described above with reference to the embodiments described above, the present disclosure is not limited to the embodiments described above. Various modifications to the structure and details of the present disclosure can be made that can be understood by those skilled in the art within the scope of the present invention. Furthermore, at least one of the functions of the detection means, generation means, and learning means described above may be performed on an information processing device installed and connected at any location on the network, that is, it may be performed using so-called cloud computing.

[0054] <Note> Some or all of the above embodiments may also be described as follows. The general configuration of the model generation apparatus, model generation method, and program in the present invention will be described below. However, the present invention is not limited to the following configuration. (Note 1) A detection means that uses an object detection model to detect a first position, which is the position of an object in the first image, and a second position, which is the position of an object in the second image, which has a different field of view from the first image. A generation means that generates a corresponding position which is the position corresponding to the first position within the second image, based on the difference in field of view between the first image and the second image, A learning means for learning the object detection model based on the second position and the corresponding position, A model generation device equipped with the following features. (Note 2) The model generation device described in Appendix 1, The generation means calculates deformation parameters for transforming the first image into the second image based on the first image and the second image, and generates the corresponding position from the first position using the deformation parameters. Model generation device. (Note 3) The model generation device described in Appendix 2, The generation means extracts a first feature point in the first image and a second feature point in the second image corresponding to the first feature point, and calculates the deformation parameter based on the first feature point and the second feature point. Model generation device. (Note 4) The model generation device described in Appendix 2, The generation means generates the corresponding position by deforming the rectangular region corresponding to the first position using the deformation parameter. Model generation device. (Note 5) The model generation device described in Appendix 1, The learning means trains the object detection model so that the error between the second position, which is the position of an object in the second image detected using the object detection model, and the position information based on the corresponding position is reduced. Model generation device. (Note 6) The model generation device described in Appendix 5, The learning means trains the object detection model so that the error between the second position, which is the position of an object in the second image detected using the object detection model, and position information consisting of a rectangle including the region of the corresponding position, becomes smaller. Model generation device. (Note 7) Using an object detection model, the first position, which is the position of the object in the first image, and the second position, which is the position of the object in the second image, which has a different field of view from the first image, are detected. Before or after detecting the second position, a corresponding position is generated from the first position, which is the position within the second image corresponding to the first position, based on the difference in field of view between the first image and the second image. Based on the second position and the corresponding position, the object detection model is trained. Model generation method. (Note 8) The model generation method described in Appendix 7, Based on the first image and the second image, deformation parameters are calculated to transform the first image into the second image, and the corresponding position is generated from the first position using these deformation parameters. Model generation method. (Note 9) The model generation method described in Appendix 7, The object detection model is trained to minimize the error between the second position, which is the position of an object in the second image detected using the object detection model, and the position information based on the corresponding position. Model generation method. (Note 10) Using an object detection model, the first position, which is the position of the object in the first image, and the second position, which is the position of the object in the second image, which has a different field of view from the first image, are detected. Based on the difference in field of view between the first image and the second image, a corresponding position is generated from the first position, which is the position corresponding to the first position within the second image. Based on the second position and the corresponding position, the object detection model is trained. A computer-readable storage medium that stores a program for causing a computer to execute a process. [Explanation of Symbols]

[0055] 2 Image Database 3 Product shelves 4 Surveillance cameras 5. Camera on a mobile device 6 Terminal devices 20 Object position estimation section 30 Loss calculation section 50 Geometric Transformation Calculation Unit 51 Feature point extraction unit 52 Feature Point Matching Calculation Unit 53 Geometric deformation parameter estimation unit 60 Automatic Correct Answering Section 61 Detection result transition area 62 Correct Rectangle Generation Unit 100 Object detection devices 101 Communications Department 102 processors 103 memory 104 Recording media

Claims

1. A detection means that uses an object detection model trained to detect the position of an object from an image taken of the object from the front to detect the first position, which is the position of the object in the first image taken from the front, and the second position, which is the position of the object in the second image, which has a different field of view from the first image. A generation means that calculates deformation parameters to transform the first image into the second image based on the difference in field of view between the first image and the second image, and uses these deformation parameters to generate a corresponding position which is the position in the second image corresponding to the first position from the first position, A learning means for training the object detection model so that the error between the second position and the position information based on the corresponding position is reduced, Equipped with, The generation means extracts a first feature point in the first image and a second feature point in the second image corresponding to the first feature point, and calculates the deformation parameter based on the first feature point and the second feature point. Model generation device.

2. A model generation apparatus according to claim 1, The generation means generates the corresponding position by deforming the rectangular region corresponding to the first position using the deformation parameter. Model generation device.

3. A model generation apparatus according to claim 1, The learning means trains the object detection model so that the error between the second position, which is the position of an object in the second image detected using the object detection model, and position information consisting of a rectangle including the region of the corresponding position, becomes smaller. Model generation device.

4. A detection means that uses an object detection model trained to detect the position of an object from images taken of an object at a specific field of view to detect a first position, which is the position of the object in a first image taken at the specific field of view, and a second position, which is the position of the object in a second image with a different field of view from the first image. A generation means that calculates deformation parameters to transform the first image into the second image based on the difference in field of view between the first image and the second image, and uses these deformation parameters to generate a corresponding position which is the position in the second image corresponding to the first position from the first position, A learning means for training the object detection model so that the error between the second position and the position information based on the corresponding position is reduced, Equipped with, The generation means extracts a first feature point in the first image and a second feature point in the second image corresponding to the first feature point, and calculates the deformation parameter based on the first feature point and the second feature point. Model generation device.

5. Using an object detection model trained to detect the position of an object from an image taken from the front, the first position, which is the position of the object in the first image taken from the front, and the second position, which is the position of the object in the second image, which has a different field of view from the first image, are detected. Before and after detecting the second position, a deformation parameter is calculated to transform the first image into the second image based on the difference in field of view between the first image and the second image, and a first feature point in the first image and a second feature point in the second image corresponding to the first feature point are extracted. The deformation parameter is calculated based on the first feature point and the second feature point, and a corresponding position is generated from the first position, which is the position in the second image corresponding to the first position. The object detection model is trained to minimize the error between the second position and the position information based on the corresponding position. Model generation method.

6. Using an object detection model trained to detect the position of an object from images taken at a specific field of view, the first position, which is the position of the object in the first image taken at the specific field of view, and the second position, which is the position of the object in the second image with a different field of view from the first image, are detected. Before and after detecting the second position, a deformation parameter is calculated to transform the first image into the second image based on the difference in field of view between the first image and the second image, and a first feature point in the first image and a second feature point in the second image corresponding to the first feature point are extracted. The deformation parameter is calculated based on the first feature point and the second feature point, and a corresponding position is generated from the first position, which is the position in the second image corresponding to the first position. The object detection model is trained to minimize the error between the second position and the position information based on the corresponding position. Model generation method.

7. Using an object detection model trained to detect the position of an object from an image taken from the front, the first position, which is the position of the object in the first image taken from the front, and the second position, which is the position of the object in the second image, which has a different field of view from the first image, are detected. Based on the difference in field of view between the first image and the second image, a deformation parameter is calculated to transform the first image into the second image, and a first feature point in the first image and a second feature point in the second image corresponding to the first feature point are extracted. Based on the first feature point and the second feature point, the deformation parameter is calculated, and a corresponding position is generated from the first position to the position in the second image that corresponds to the first position. The object detection model is trained to minimize the error between the second position and the position information based on the corresponding position. A program that causes a computer to perform a process.

8. Using an object detection model trained to detect the position of an object from images taken at a specific field of view, the first position, which is the position of the object in the first image taken at the specific field of view, and the second position, which is the position of the object in the second image with a different field of view from the first image, are detected. Based on the difference in field of view between the first image and the second image, a deformation parameter is calculated to transform the first image into the second image, and a first feature point in the first image and a second feature point in the second image corresponding to the first feature point are extracted. Based on the first feature point and the second feature point, the deformation parameter is calculated, and a corresponding position is generated from the first position to the position in the second image that corresponds to the first position. The object detection model is trained to minimize the error between the second position and the position information based on the corresponding position. A program that causes a computer to perform a process.

Citation Information

Patent Citations

  • Gaze point conversion device and method

    JP2014186720A

  • Image processing apparatus, display control apparatus, image processing method, and recording medium

    JP2020061158A