Training data generation system, model learning system, individual identification system, training data generation method, model learning method, individual identification method, program

The system automates the generation of training data for individual identification by detecting feature points and tracking objects, addressing the inefficiencies and errors of manual correction methods, thereby enhancing the accuracy and speed of identification processes.

JP2026053209APending Publication Date: 2026-03-25IWATE PREFECTURAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing individual identification systems require manual input of correction information, increasing time and cost for generating training data and potentially leading to human errors.

Method used

A system that automatically generates training data by detecting feature points in images of moving objects and assigning identification information based on predetermined conditions, using imaging and communication devices to track and label objects within a specified area.

Benefits of technology

Facilitates efficient and accurate generation of training data without manual intervention, reducing human errors and improving the efficiency of individual identification processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053209000001_ABST
    Figure 2026053209000001_ABST
Patent Text Reader

Abstract

The present invention provides a training data generation system, a training data generation method, and a program capable of easily generating training data for individual identification based on images, and an individual identification system, a model training system, an individual identification method, a model training method, and a program capable of easily performing individual identification based on such training data. [Solution] The learning data generation system comprises: a first acquisition means 41 that acquires an image of a moving object located in a predetermined area, captured by an imaging device; a detection means 42 that detects at least one feature point of the moving object based on the acquired image; a second acquisition means 44 that acquires information regarding the position of the moving object within the predetermined area; and a generation means 46 that generates learning data for individual identification of moving objects based on images of moving objects by assigning identification information of the moving object as a label to at least one feature point of the moving object when the position of the moving object within the predetermined area satisfies predetermined conditions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to a learning data generation system, a model learning system, an individual identification system, a method for generating learning data, a model learning method, an individual identification method, and a program. [Background technology]

[0002] In recent years, individual identification systems for identifying each of multiple moving objects have been known (for example, Patent Document 1).

[0003] The individual identification system described in Patent Document 1 is configured to extract features from images of registered pets (moving objects) and store them in association with the pet's identifier. Furthermore, the individual identification system is configured to identify pets using a trained model based on machine learning using training data. Here, the individual identification system is configured to prompt the user for correction if it determines that the identification result is incorrect, and to update the identification result (register training data) when correction information is obtained from the user. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2019-71895 [Overview of the project] [Problems that the invention aims to solve]

[0005] The technology described in Patent Document 1 requires manual input by the user of correction information (accurate identification information (correct label)) when generating training data for individual identification based on images. This increases the time required to generate training data and the burden on the user, potentially leading to higher costs for generating training data.

[0006] The present invention has been made in view of the above problems, and aims to provide a learning data generation system, a learning data generation method, a program that can easily generate learning data for individual identification based on images, and an individual identification system, a model learning system, an individual identification method, a model learning method, and a program that can easily perform individual identification based on such learning data. [Means for solving the problem]

[0007] To solve the above problems, firstly, the present invention provides a learning data generation system comprising: a first acquisition means for acquiring an image of a moving object located in a predetermined area, captured by an imaging device; a detection means for detecting at least one feature point of the moving object based on the acquired image; a second acquisition means for acquiring information regarding the position of the moving object within the predetermined area; and a generation means for generating learning data for individual identification of the moving object based on an image of the moving object, by assigning identification information of the moving object as a label to at least one feature point of the moving object when the position of the moving object within the predetermined area satisfies predetermined conditions. (Invention 1)

[0008] According to the invention (Invention 1), when the position of a moving object within a predetermined area satisfies predetermined conditions, identification information of the moving object is assigned as a label to at least one feature point of the moving object detected from the image of the moving object. This makes it possible to automatically generate training data for individual identification of the moving object based on the image of the moving object, taking into account the position of the moving object within the predetermined area. As a result, training data can be generated more easily compared to, for example, generating training data through manual work. Furthermore, it becomes possible to eliminate human errors when generating training data, thus enabling the generation of appropriate training data.

[0009] In the above invention (Invention 1), each time the detection means detects at least one feature point of one or more moving bodies, it assigns individual tracking information to the detected at least one feature point of the moving body, and the generation means may assign identification information of the moving body as a label to at least one feature point of the moving body and the tracking information if the position of the moving body within the predetermined area satisfies the predetermined conditions (Invention 2).

[0010] According to the present invention (Invention 2), for example, even if the captured image contains multiple moving objects, by using the tracking information of each moving object, it becomes possible to efficiently assign identification information of a predetermined moving object (i.e., a moving object whose position within a predetermined area satisfies predetermined conditions) to at least one feature point of that moving object.

[0011] In the above invention (Invention 2), the predetermined condition may include the condition that no other second communication device exists within a predetermined range from the position of the second communication device provided on the moving body (Invention 3).

[0012] According to the present invention (Invention 3), for example, it becomes possible to suppress the generation of inaccurate training data, such as incorrectly assigning identification information of another moving object corresponding to another second communication device to at least one feature point of a moving object detected from an image of the moving object.

[0013] In the above invention (Invention 2), the predetermined conditions may include the state that the Euclidean distance between the first communication device and the second communication device provided on the moving body is less than a predetermined value for a predetermined period of time (Invention 4).

[0014] According to this invention (Invention 4), for example, the longer the Euclidean distance between the first communication device and the second communication device provided on the moving object remains short, the easier it becomes to assign identification information of the moving object to at least one feature point of the moving object, thereby improving the efficiency of generating accurate training data.

[0015] In the above invention (Invention 2), if the identification information of the moving body has already been assigned to the assigned tracking information, the generating means may assign the already assigned identification information of the moving body as a label to the assigned tracking information (Invention 5).

[0016] According to this invention (Invention 5), if motion identification information has already been assigned to the assigned tracking information, the motion identification information that has already been assigned is automatically added to the assigned tracking information, thereby making it possible to further improve the efficiency of generating training data.

[0017] In the above inventions (inventions 1 to 5), the predetermined conditions may include the shortest Euclidean distance between the first communication device provided in the predetermined area and the second communication device provided on the moving body among a plurality of second communication devices located in the predetermined area (invention 6).

[0018] According to this invention (Invention 6), it becomes possible to automatically generate training data for individual identification of a moving object based on an image of the moving object, taking into account the Euclidean distance between a first communication device provided in a predetermined area and a second communication device provided on the moving object within the predetermined area.

[0019] In the above invention (Invention 1), the detection means may detect the coordinates of each of at least one part of the moving body in the image as at least one feature point of the moving body (Invention 7).

[0020] According to this invention (Invention 7), training data can be generated using the coordinates of each of at least one part of a moving body.

[0021] In the above invention (Invention 7), conversion means is provided for converting the coordinates in the image of each of at least one part of the detected moving object based on the distance between at least two parts of the detected moving object, and when the position of the moving object in the predetermined area satisfies a predetermined condition, the generation means may assign the identification information of the moving object as a label to the coordinates of each of at least one part of the converted moving object (Invention 8).

[0022] According to such an invention (Invention 8), learning data can be generated using the coordinates of each of at least one part of the moving object converted according to the distance between at least two parts of the moving object (for example, it may be the measured value of the distance between at least two parts of the moving object, or it may be the measured value of the distance between at least two parts of the moving object in a pre-captured image).

[0023] In the above invention (Invention 8), the conversion means may convert the coordinates in the image of each of at least one part of the detected moving object based on the distance between the neck part and the waist part of the moving object (Invention 9).

[0024] According to such an invention (Invention 9), by using the distance between the neck part and the waist part of the moving object (for example, it may be the measured value of the distance between the neck part and the waist part of the moving object, or it may be the measured value of the distance between the neck part and the waist part of the moving object in a pre-captured image), the coordinates in the image of each of at least one part of the moving object can be easily converted.

[0025] In the above inventions (Inventions 7 to 9), the detection means may detect at least one joint of the moving object as at least one part of the moving object (Invention 10).

[0026] According to such an invention (Invention 10), learning data can be generated using the coordinates in the image of each of at least one joint of the moving object.

[0027] In the above invention (Invention 1), the second acquisition means may acquire information regarding the position of the moving body measured using a communication device (Invention 11).

[0028] According to this invention (Invention 11), it becomes possible to easily determine whether or not the position of a moving object satisfies predetermined conditions based on the position of the moving object measured by a communication device.

[0029] In the above invention (Invention 11), the second acquisition means may acquire information regarding the received signal strength of signals transmitted and received between the first communication device provided in the predetermined area and the second communication device provided on the moving body, as information regarding the position of the moving body (Invention 12).

[0030] According to this invention (Invention 12), it becomes possible to easily determine whether or not the position of a moving object satisfies predetermined conditions based on the received signal strength of signals transmitted and received between the first communication device and the second communication device.

[0031] In the above invention (Invention 12), estimation means may be provided for estimating the position of the moving object within the predetermined area based on information regarding the received signal strength (Invention 13).

[0032] According to this invention (invention 13), it becomes possible to estimate the position of a moving object based on the received signal strength of signals transmitted and received between a first communication device provided in a predetermined area and a second communication device provided on the moving object.

[0033] In the above invention (Invention 13), the estimation means may select at least two second communication devices, including the second communication device provided on the moving body, if the received signal strength for the second communication device provided on the moving body is higher than the received signal strength for at least one other second communication device located in the predetermined area, and test whether there is a significant difference in the average value of the received signal strength over a predetermined period between the selected second communication devices. If such a significant difference is found, the position of the second communication device provided on the moving body may be estimated as the position of the moving body (Invention 14).

[0034] According to this invention (Invention 14), for example, instead of simply comparing the received signal strengths corresponding to each of a plurality of second communication devices, it becomes possible to estimate the position of a second communication device installed on a moving object (i.e., the position of the moving object) based on whether or not there is a significant difference between the received signal strength corresponding to a specific second communication device and the received signal strength corresponding to another second communication device (i.e., whether or not the difference in received signal strength is accidental). This makes it possible to accurately estimate the position of the moving object while reducing the effects of, for example, radio wave fluctuations.

[0035] Secondly, the present invention provides a model learning system (Invention 15) that includes a learning means for learning a model used to identify individual moving objects based on images of moving objects by machine learning using the learning data generated by the learning data generation system of the above inventions (Inventions 1 to 14).

[0036] According to this invention (Invention 15), it becomes possible to train a model used for identifying individuals by machine learning that uses information about at least one feature point of a moving object in an image as training data. Therefore, by using this model, individual identification can be easily performed.

[0037] Thirdly, the present invention provides an individual identification system (Invention 16) comprising: an image acquisition means for acquiring an image of a moving object located in a predetermined area, captured by an imaging device; and an identification means for performing individual identification of the moving object based on the image of the moving object and a trained model based on machine learning using training data generated by the training data generation system of the above inventions (Inventions 1 to 14).

[0038] According to this invention (invention 16), individual identification of a moving object is performed based on information relating to each of at least one feature points of the moving object detected based on an image of the moving object captured by the imaging device, and a trained model. Therefore, individual identification of the moving object can be easily performed by using an image of the moving object captured by the imaging device.

[0039] Fourthly, the present invention provides a method for generating learning data, wherein a computer performs the following steps: acquiring an image of a moving object located in a predetermined area, captured by an imaging device; detecting at least one feature point of the moving object based on the acquired image; acquiring information regarding the position of the moving object within the predetermined area; and generating learning data for individual identification of the moving object based on the image of the moving object by assigning identification information of the moving object as a label to at least one feature point of the moving object when the position of the moving object within the predetermined area satisfies predetermined conditions (Invention 17).

[0040] Fifth, the present invention provides a model learning method (Invention 18) in which a computer performs the step of learning a model used to identify individual moving objects based on images of moving objects by machine learning using the learning data generated by the learning data generation method of the above invention (Invention 17).

[0041] Sixth, the present invention provides an individual identification method (Invention 19) in which a computer performs the following steps: acquiring an image of a moving object located in a predetermined area, captured by an imaging device; and performing individual identification of the moving object based on the image of the moving object and a trained model based on machine learning using training data generated by the training data generation method of the above invention (Invention 17).

[0042] Seventh, the present invention provides a program for a computer to implement the following functions: acquiring an image of a moving object located in a predetermined area, captured by an imaging device; detecting at least one feature point of the moving object based on the acquired image; acquiring information regarding the position of the moving object within the predetermined area; and generating learning data for individual identification of the moving object based on the image of the moving object by assigning identification information of the moving object as a label to at least one feature point of the moving object when the position of the moving object within the predetermined area satisfies predetermined conditions (Invention 20).

[0043] Eighth, the present invention provides a program for a computer to implement a function that learns a model used for individual identification of moving objects based on images of moving objects by machine learning using training data generated by the program of the above invention (Invention 20) (Invention 21).

[0044] Ninthly, the present invention provides a program for a computer to implement the following functions: acquiring an image of a moving object located in a predetermined area, captured by an imaging device; and performing individual identification of the moving object based on the image of the moving object and a trained model based on machine learning using training data generated by the program of the above invention (Invention 20) (Invention 22). [Effects of the Invention]

[0045] According to the learning data generation system, learning data generation method, and program of the present invention, learning data for individual identification based on images can be easily generated. Furthermore, according to the individual identification system, model learning system, individual identification method, model learning method, and program of the present invention, individual identification can be easily performed based on such learning data. [Brief explanation of the drawing]

[0046] [Figure 1] This figure schematically shows the basic configuration of a learning data generation system, a model learning system, and an individual identification system according to one embodiment of the present invention. [Figure 2] This is a block diagram showing the configuration of the identification device. [Figure 3] This is a functional block diagram illustrating the functions that play a major role in the learning data generation system, model learning system, and individual identification system. [Figure 4] This figure shows an example of the structure of the first acquired data. [Figure 5] This figure shows an example of the detection result for at least one part of a moving object. [Figure 6] This figure shows an example of transforming the coordinates of at least one part of a moving object. [Figure 7] This figure shows an example of the structure of the second set of acquired data. [Figure 8] This figure shows an example of the relationship between the distance from the first communication device and the received signal strength. [Figure 9] This flowchart shows an example of the estimation process for a second communication device located adjacent to a first communication device. [Figure 10] This flowchart shows an example of processing by the generation means. [Figure 11] This figure shows an example of the structure of the first training data. [Figure 12] This flowchart shows another example of processing by the generation means. [Figure 13] This figure shows an example of the structure of the second training data. [Figure 14]This flowchart shows an example of the main processing steps of a learning data generation system according to one embodiment of the present invention. [Figure 15] This flowchart shows an example of the main processing steps of a model generation system according to one embodiment of the present invention. [Figure 16] This flowchart shows an example of the main processing steps of an individual identification system according to one embodiment of the present invention. [Figure 17] This figure shows examples of the division of labor between the identification device and the learning device for the functions of the learning data generation system, the model learning system, and the individual identification system. [Modes for carrying out the invention]

[0047] One embodiment of the present invention will be described in detail below with reference to the accompanying drawings. However, this embodiment is illustrative and the present invention is not limited thereto.

[0048] (1) Basic configuration of the learning data generation system, model learning system, and individual identification system Figure 1 is a schematic diagram showing the basic configuration of a learning data generation system, a model learning system, and an individual identification system according to one embodiment of the present invention.

[0049] In the learning data generation system according to this embodiment, when a moving object (in this embodiment, a subject T, which is a person to be individually identified among multiple people present in the spatial SP) is present in a spatial SP that constitutes a predetermined area (for example, a part of a store or facility), the identification device 30 acquires an image of the subject T captured by an imaging device 10 installed at a predetermined position in the spatial SP (in the example of Figure 1, the upper center of the spatial SP), and information regarding the position of the subject T measured by a communication device 20. Furthermore, the identification device 30 is configured to generate learning data for individual identification of the subject T based on the image of the subject T by detecting at least one feature point of the subject T based on the acquired image of the subject T, and if the position of the subject T satisfies predetermined conditions, by assigning the identification information of the subject T as a label to at least one feature point of the subject T.

[0050] Furthermore, in the model generation system according to this embodiment, the identification device 30 learns a model used for individual identification of subject T based on an image of subject T by machine learning using the training data generated by the above-mentioned training data generation system.

[0051] Furthermore, in the individual identification system according to this embodiment, the identification device 30 acquires an image of a subject T located in a predetermined area, which is captured by the imaging device 10, and performs individual identification of the subject T based on the acquired image of the subject T and a trained model based on machine learning using the training data generated by the training data generation system described above.

[0052] Here, the imaging device 10 and the communication device 20, as well as the identification device 30, are connected to a communication network NW (network), such as the Internet or a LAN (Local Area Network).

[0053] The imaging device 10 may be, for example, an imaging device that captures moving images and / or still images (e.g., a digital camera or a digital video camera), and is configured to capture images of the spatial SP at a predetermined position within the spatial SP. The imaging device 10 may be mounted on a predetermined moving body and capture images of the spatial SP while moving in conjunction with the movement of the moving body. The imaging device 10 is configured to perform imaging processing at a predetermined frame rate (e.g., 30 fps (frames per second)) and transmit the captured images to the identification device 30 via a communication network NW. The imaging device 10 may also perform imaging processing when it receives a predetermined imaging instruction signal from the identification device 30.

[0054] Here, the imaging device 10 may be an imaging device that captures omnidirectional images (for example, images of the surrounding 360° (in the example of Figure 1, the surrounding 360° in the horizontal direction)) (for example, an omnidirectional camera, etc.). In this case, the subject T can be imaged over a wide area.

[0055] Furthermore, the imaging device 10 may be an imaging device that captures infrared images (for example, an infrared camera). This makes it possible to detect the subject T in the captured image even when the subject T is in an environment with poor visibility (for example, at night, in a dark place, in bad weather, etc.).

[0056] Furthermore, the imaging device 10 may be an imaging device that captures stereo images (for example, a stereo camera). Here, a stereo image may be, for example, a set of two images having a predetermined parallax.

[0057] In this embodiment, the case of imaging using one imaging device 10 is described as an example, but spatial SP may be imaged using multiple imaging devices (for example, two imaging devices spaced apart in the horizontal and / or vertical directions).

[0058] The communication device 20 is a device that measures the position of a subject T within space SP continuously or intermittently (for example, at predetermined intervals (e.g., 300 milliseconds or 1 second)). Here, information regarding the position of subject T may be represented by position information measured using position measurement technology such as GPS (Global Positioning System) (for example, at least one of latitude, longitude, and altitude), or by information regarding the received signal strength indicator (RSSI) when a signal transmitted by one device is received by the other device. Here, information regarding the received signal strength may be, for example, the value of the received signal strength, a value obtained by substituting the value of the received signal strength into a predetermined calculation formula, or information representing the degree of the received signal strength.

[0059] Furthermore, the communication device 20 may consist of a first communication device 21 positioned at a predetermined location within the space SP, as shown in Figure 1, and a second communication device 22 that can be possessed by or worn on the body of the subject T, and which communicates wirelessly with the first communication device 21 when the second communication device 22 is present within the space SP. In the example shown in Figure 1, one first communication device 21 is provided within the space SP as an example, but multiple first communication devices 21 may be provided within the space SP. In this case, each of the multiple first communication devices 21 may be provided at a different location within the space SP.

[0060] One of the first communication device 21 and the second communication device 22 may be configured to communicate wirelessly with the other of the first communication device 21 and the second communication device 22 using a predetermined wireless communication method (for example, wireless LAN (for example, Wi-Fi®)) when the second communication device 22 is located within the spatial SP.

[0061] The first communication device 21 may be located in a position within the spatial SP that allows it to wirelessly communicate with the second communication device 22 using a predetermined wireless communication method (e.g., wireless LAN (e.g., Wi-Fi®)). The first communication device 21 may also be a device that relays wireless communication between two or more second communication devices 22 located within the spatial SP, or a device that relays wireless communication between the second communication device 22 and other devices (not shown) located within the spatial SP, or a device that relays communication between the second communication device 22 and other devices (e.g., identification device 30) connected via a communication network NW. The first communication device 21 may also be a packet capture device.

[0062] The second communication device 22 may be configured to transmit a wireless signal (e.g., a probe request) containing its own identification information (e.g., a MAC (Media Access Control) address, etc.) and / or the identification information (e.g., the person's ID and name, etc.) of the person possessing or wearing the second communication device 22 (including the subject T) at predetermined intervals (e.g., every few hundred milliseconds) in order to communicate wirelessly with the first communication device 21. Here, the person's identification information may be stored in the second communication device 22 in advance, for example. Furthermore, the second communication device 22 may be, for example, a device that can be worn by a person (e.g., a wearable device) or a portable device that a person can carry. Moreover, the second communication device 22 may be a communication device operated by an individual user, such as a mobile terminal, smartphone, PDA (Personal Digital Assistant), personal computer, or television receiver with two-way communication capabilities (including so-called multi-functional smart TVs).

[0063] Furthermore, either the first communication device 21 or the second communication device 22 may be provided with an RSSI circuit for detecting the received signal strength (RSSI) when it receives a signal transmitted by the other device. In addition, this signal may include identification information (e.g., MAC address, etc.) of the communication device that transmitted the signal (first communication device 21 or second communication device 22), and when the received signal strength of this signal is detected by the RSSI circuit, the detected received signal strength and the identification information of the device that transmitted this signal may be stored in a memory device (not shown) provided in the communication device that received the signal (first communication device 21 or second communication device 22) in a corresponding state.

[0064] In this explanation, we describe a case where wireless communication is performed between the first communication device 21 and the second communication device 22 using Wi-Fi® as an example, but the communication method is not limited to this case. For example, wireless communication methods such as Bluetooth®, ZigBee®, UWB, optical wireless communication (e.g., infrared) may be used, or wired communication methods such as USB may be used.

[0065] The communication device 20 (either the first communication device 21 or the second communication device 22) is configured to transmit information regarding the measured location of subject T to the identification device 30 via the communication network NW each time it measures the location of subject T (in this case, the received signal strength (RSSI) when one of the communication devices, the first or second communication device 21, transmits a signal to the other communication device).

[0066] Furthermore, the imaging device 10 and the communication device 20 may be configured to communicate directly with the identification device 30 using a wired or wireless connection, or they may be configured to communicate with the identification device 30 via a predetermined relay device (not shown) by transmitting and receiving information using a wired or wireless connection.

[0067] The identification device 30 communicates with the imaging device 10 and the communication device 20 via a communication network NW, and is configured to acquire images captured by the imaging device 10 and information regarding the location of the subject T measured by the communication device 20 over time via the communication network NW. The identification device 30 may be a terminal device operated by an individual user, such as a mobile terminal, smartphone, PDA, personal computer, or television receiver with two-way communication capabilities (including so-called multi-functional smart TVs).

[0068] (2) Configuration of the identification device The configuration of the identification device 30 will be described with reference to Figure 2. Figure 2 is a block diagram showing the internal configuration of the identification device 30. As shown in Figure 2, the identification device 30 comprises a CPU (Central Processing Unit) 31, a ROM (Read Only Memory) 32, a RAM (Random Access Memory) 33, a storage device 34, a display processing unit 35, a display unit 36, an input unit 37, and a communication interface unit 38, and is provided with a bus 30a for transmitting control signals or data signals between each unit.

[0069] When power is supplied to the identification device 30, the CPU 31 loads various programs stored in the ROM 32 or storage device 34 into the RAM 33 and executes them. In this embodiment, the CPU 31 reads and executes programs stored in the ROM 32 or storage device 34 to realize the functions of the first acquisition means 41, detection means 42, conversion means 43, second acquisition means 44, estimation means 45, generation means 46, learning means 47, image acquisition means 48, and identification means 49 (shown in Figure 3), which will be described later.

[0070] The storage device 34 may be a non-volatile storage device such as flash memory, SSD (Solid State Drive), magnetic storage device (e.g., HDD (Hard Disk Drive), floppy disk (registered trademark), magnetic tape, etc.), or optical disk, or it may be a volatile storage device such as RAM, and it stores programs executed by the CPU 31 and data referenced by the CPU 31. The storage device 34 also stores first acquired data (shown in Figure 4) and second acquired data (shown in Figure 7), as well as first learning data (shown in Figure 11) and / or second learning data (shown in Figure 13), which will be described later.

[0071] The display processing unit 35 displays the display data provided by the CPU 31 on the display unit 36. The display unit 36 ​​is, for example, an LCD (Liquid Crystal Display) monitor including thin-film transistors arranged in a matrix on a pixel-by-pixel basis, and displays the data to be displayed on the display screen by driving the thin-film transistors based on the display data.

[0072] If the identification device 30 is a button-input type device, the input unit 37 includes a group of buttons including a plurality of instruction input buttons such as a direction indicator button and a select button for receiving user operation input, and a group of buttons including a plurality of instruction input buttons such as a numeric keypad, and includes an interface circuit for recognizing the pressing (operation) input of each button and outputting it to the CPU 31.

[0073] If the identification device 30 is a touch panel input device, the input unit 37 primarily accepts input via a touch panel, such as touching the display screen with a fingertip or pen. The touch panel input method may be a known method such as a capacitive touch panel.

[0074] Furthermore, if the identification device 30 is a device capable of voice input, the input unit 37 may be configured to include a microphone for voice input, or it may include an interface circuit for outputting voice data input via an external microphone to the CPU 31. In addition, if the identification device 30 is a device capable of inputting moving images and / or still images, the input unit 37 may be configured to include a digital camera or digital video camera for image input, or it may include an interface circuit for receiving image data captured by an external digital camera or digital video camera and outputting it to the CPU 31.

[0075] The communication interface unit 38 includes an interface circuit for communicating with other devices (for example, the imaging device 10 and the communication device 20, etc.) via a communication network NW.

[0076] (3) Overview of each function in the learning data generation system, model learning system, and individual identification system The functions realized in the learning data generation system, model learning system, and individual identification system of this embodiment will be described with reference to Figure 3. Figure 3 is a functional block diagram illustrating the functions that play a major role in the learning data generation system, model learning system, and individual identification system of this embodiment. In the functional block diagram of Figure 3, the first acquisition means 41, detection means 42, second acquisition means 44, and generation means 46 correspond to the main components of the learning data generation system of the present invention, the learning means 47 corresponds to the main components of the model learning system of the present invention, and the image acquisition means 48 and identification means 49 correspond to the main components of the individual identification system of the present invention. Other means (conversion means 43 and estimation means 45) are not necessarily essential components, but are components that make the present invention even more preferable.

[0077] The first acquisition means 41 has the function of acquiring an image of a subject T (moving body) located in a spatial SP (predetermined area) captured by the imaging device 10.

[0078] The function of the first acquisition means 41 is realized, for example, as follows. First, the imaging device 10 performs imaging processing at a predetermined frame rate (e.g., 30fps) when, for example, the subject T is present in space SP, and each time imaging processing is performed, it transmits the image data of the captured image to the identification device 30 via the communication network NW. Here, the image data of the image captured by the imaging device 10 may be transmitted to the identification device 30 in association with the date and time of capture and the identification information of the imaging device 10 (e.g., the serial number or MAC (Media Access Control) address of the imaging device 10).

[0079] Meanwhile, the CPU 31 of the identification device 30 receives (acquires) image data transmitted from the imaging device 10 via the communication interface unit 38, and stores the received image data in the first acquisition data shown in Figure 4, for example, in association with the date and time the image was captured. The first acquisition data is data that describes the image data of each image in association with the date and time the image was captured. In this way, the first acquisition means 41 can acquire an image of the subject T captured by the imaging device 10.

[0080] The CPU 31 may acquire a stereo image of the subject T (moving body) captured by the imaging device 10. In this case, a three-dimensional model of the subject T can be generated based on the stereo image of the subject T, and the detection means 42, which will be described later, will be able to determine the position (coordinates) of at least one part of the subject T in three-dimensional space based on this three-dimensional model. This makes it possible to capture the position of at least one part of the subject T more accurately compared to, for example, using a two-dimensional model of the subject T, thereby improving the accuracy of individual identification of the subject T based on the image of the subject T.

[0081] Furthermore, for example, if two or more imaging devices are provided that are spaced apart in the horizontal and / or vertical directions, the CPU 31 may acquire two images captured by each imaging device at substantially the same time (for example, two images with the same capture date and time, or two images with a time difference of a predetermined range (for example, a few milliseconds to tens of milliseconds, etc.)) as a set of images having a predetermined parallax (i.e., a stereo image).

[0082] The detection means 42 has the function of detecting at least one feature point of the subject T (moving body) based on the acquired image.

[0083] Furthermore, the detection means 42 may detect the coordinates of each image of at least one part of the subject T (moving body) as at least one feature point of the subject T. This allows training data to be generated using the coordinates of at least one part of the subject T.

[0084] Furthermore, the detection means 42 may detect at least one joint of the subject T (moving body) as at least one body part of the subject T. This allows training data to be generated using the coordinates of each of the at least one joints of the subject T within the image.

[0085] Furthermore, the detection means 42 may assign individual tracking information to at least one feature point of each of the detected subject T (moving object) each time it detects at least one feature point of the subject T. This makes it possible to assign individual tracking information to each subject even if the captured image contains multiple subjects.

[0086] The function of the detection means 42 can be implemented as follows, for example. The CPU 31 of the identification device 30 may, for example, based on the function of the first acquisition means 41, store the image data transmitted from the imaging device 10 in the first acquisition data, set up nodes in the image corresponding to the position of at least one part of the subject T, and calculate the coordinates of each node in the image. Here, the setting of each node and the calculation of the coordinates may be performed using, for example, deep learning-based feature point (keypoint) detection technology (e.g., OpenPose or YOLOv8 (https: / / docs.ultralytics.com / ja)). Furthermore, the coordinates of each node may be two-dimensional coordinates, or, for example, three-dimensional coordinates if a three-dimensional model of the subject T is generated from the image data (stereo image).

[0087] For example, if the imaging device 10 captures an omnidirectional image, the CPU 31 may unfold the omnidirectional image captured by the imaging device 10 into a panoramic image and then perform node setting and node coordinate calculation on the unfolded panoramic image.

[0088] Furthermore, if stereo images of subject T are acquired, the CPU 31 may perform preprocessing (e.g., noise reduction) on each image in a set of images having a predetermined parallax, and then perform a well-known stereo matching process to generate a three-dimensional model of subject T.

[0089] As shown in Figure 5, for example, the CPU 31 detects the position (in the example, the X-coordinate and Y-coordinate) of at least one part of the subject T (in the example, "neck," "right shoulder," "left shoulder," "waist," "right buttock," "left buttock," etc.) in the image captured by the imaging device 10, and sets a node (in the example, 14 nodes) corresponding to each detected part. Here, the coordinates of each node may be stored in the first acquired data in a state where each node is associated with the image in which it was set. Furthermore, if, for example, at least one part of the subject T cannot be detected because it is obscured by another object (for example, another moving object or installed object, etc.) during imaging, the CPU 31 may store the coordinates of the node corresponding to that part as NULL data in the first acquired data.

[0090] In this explanation, we have described an example in which at least one part of subject T (e.g., "neck," "right shoulder," "left shoulder," "right buttock," "left buttock") is detected using the image captured by the imaging device 10. However, for example, at least one joint of subject T (e.g., "right shoulder joint," "right elbow joint," "right hip joint," "right knee joint," "left shoulder joint," "left elbow joint," "left hip joint," and "left knee joint") may also be detected using the image captured by the imaging device 10.

[0091] Furthermore, the CPU 31 may, for example, store image data transmitted from the imaging device 10 in the first acquisition data based on the function of the first acquisition means 41, and each time it stores the image data in the first acquisition data, use a well-known tracking algorithm (e.g., ByteTrack or BoT-SORT (Robust Associations Multi-Pedestrian Tracking)) to assign the same tracking information (tracking ID) to at least one feature point of each subject T (here, the coordinates in each image of at least one body part (which may include a joint)) of the subject T that is presumed to be the same person in each image. In addition, the CPU 31 may store the tracking information assigned to one or more subject Ts included in the image in the first acquisition data in association with at least one feature point of the corresponding subject T. In this case, the person tracking function assigns the same tracking information to at least one feature point of the person that is presumed to be the same person, and this same tracking information is maintained until the tracking information of the subject T is changed, for example, by occlusion.

[0092] Furthermore, if the image contains multiple people, the CPU 31 may detect at least one feature point for each of the multiple people. In addition, the CPU 31 may assign the same tracking information (tracking ID) to each person who is presumed to be the same person in each of the multiple images.

[0093] The conversion means 43 has the function of converting the coordinates of at least one part of the detected subject T (moving body) in each image based on the distance between at least two parts of the detected subject T.

[0094] Furthermore, the conversion means 43 may convert the coordinates of each image of at least one part of the detected subject T (moving body) based on the distance between the neck and waist of the subject T. In this case, the coordinates of each image of at least one part of the subject T can be easily converted by using the distance between the neck and waist of the subject T (for example, the measured value of the distance between the neck and waist of the subject T, or the measured value of the distance between the neck and waist of the subject T in a previously captured image).

[0095] The function of the conversion means 43 is realized, for example, as follows. The CPU 31 of the identification device 30 detects the position (coordinates in the image) of at least one part of the subject T based on the function of the detection means 42, and then converts the coordinates of each of the detected at least one part based on the distance between the neck and waist of the subject T. The coordinates of the neck and waist of the subject T may be detected based on the function of the detection means 42 described above, or, for example, the coordinates of the midpoint between the left shoulder and right shoulder of the subject T may be calculated as the coordinates of the neck, and the coordinates of the midpoint between the left waist and right waist of the subject T may be calculated as the coordinates of the waist. Now, an example of this conversion method will be explained with reference to Figure 6. For example, as shown in the left figure of Figure 6, when the coordinates of the waist of the subject T are taken as the origin, the CPU 31 calculates the Euclidean distance r between each of at least one part of the subject T (which may include joints) and the waist. Next, the CPU 31 may calculate a multiplier v (e.g., 1 / r) for at least one part of the subject T such that the length between the neck and waist of the subject T is 1, and transform the original coordinates (r cosθ, r sinθ) of the corresponding part into (vr cosθ, vr sinθ).

[0096] In this example, the conversion means 43 converts the coordinates in the acquired image of each of at least one detected part of the subject T based on the distance between the neck and waist of the subject T measured in advance. However, the present invention is not limited to this case. For example, the conversion means 43 may convert the coordinates in the acquired image of each of at least one detected part of the subject T based on the distance between either the neck or waist of the subject T and other parts other than the neck and waist, or based on the distance between two or more other parts of the subject T other than the neck and waist.

[0097] The second acquisition means 44 has the function of acquiring information regarding the position of a subject T (moving body) within a spatial SP (predetermined area).

[0098] Furthermore, the second acquisition means 44 may acquire information regarding the position of the subject T (moving body) measured using the communication device 20. This makes it possible to easily determine, based on the position of the subject T measured by the communication device 20, whether or not the position of the subject T satisfies predetermined conditions in the function of the generation means 46 described later.

[0099] Furthermore, the second acquisition means 44 may acquire information regarding the received signal strength of signals transmitted and received between the first communication device 21 located in a predetermined area and the second communication device 22 located on the subject T (moving object), as information regarding the position of the subject T. This makes it possible to easily determine whether the position of the subject T satisfies predetermined conditions in the function of the generation means 46 described later, based on the received signal strength (RSSI) of signals transmitted and received between the first communication device 21 and the second communication device 22.

[0100] The function of the second acquisition means 44 is realized, for example, as follows. Hereinafter, we will explain as an example the case in which the second acquisition means 44 acquires information regarding the received signal strength of signals transmitted and received between the first communication device 21 installed in a predetermined area and the second communication device 22 installed on the subject T, as information regarding the position of the subject T.

[0101] The second communication device 22, for example, when a target person T is present in space SP, communicates wirelessly with the first communication device 21 and, each time it receives a wireless signal (e.g., a beacon signal) transmitted from the first communication device 21 at predetermined intervals (e.g., several hundred millisecond intervals), stores the value of the received signal strength (RSSI) of the wireless signal, detected (measured) by the RSSI circuit, in a storage device (not shown) provided in the second communication device 22. Here, the RSSI value may be stored in association with the reception date and time (e.g., the date and time when the second communication device 22 received the wireless signal corresponding to the RSSI value from the first communication device 21) and the identification information of the first communication device 21 that transmitted the wireless signal corresponding to the RSSI value (e.g., the serial number or MAC address of the first communication device 21). The second communication device 22 may also transmit the RSSI value of the wireless signal transmitted from the first communication device 21 to the identification device 30 via the communication network NW each time it receives a wireless signal from the first communication device 21. Here, the RSSI value of the wireless signal may be transmitted to the identification device 30 in association with the date and time of reception of the wireless signal, the identification information of the first communication device 21 that transmitted the wireless signal, the identification information of the second communication device 22 (for example, the serial number or MAC address of the second communication device 22), and the identification information of the person T possessing or wearing the second communication device 22 (for example, the ID or name of person T). The identification information of person T may be stored in advance in the second communication device 22 possessed or worn by person T.

[0102] In this explanation, we have described the case in which the second communication device 22 transmits the RSSI value to the identification device 30 when it receives a radio signal transmitted from the first communication device 21 as an example. However, the first communication device 21 may also transmit the RSSI value to the identification device 30 when it receives a radio signal transmitted from the second communication device 22.

[0103] On the other hand, each time the CPU 31 of the identification device 30 receives (acquires) information transmitted from the communication device 20 (in this case, the second communication device 22) via the communication interface unit 38, it stores the received information in the second acquired data shown in Figure 7, for example. The second acquired data is data that describes the information received from each of the one or more second communication devices 22 (in the example shown in the figure, the date and time of reception of the signal transmitted from the first communication device 21, the received signal strength (RSSI) of the signal, the identification information of the first communication device 21 (first communication device ID), the identification information of the second communication device 22 (second communication device ID), and the identification information of the person T who possesses or wears the second communication device 22).

[0104] Furthermore, the CPU 31 of the identification device 30 can estimate the position of a subject T (moving object) within a predetermined area based on information regarding the received signal strength, using the function of the estimation means 45, which will be described later. Therefore, the CPU 31 can obtain information regarding the position of the subject T within the spatial SP by receiving (acquiring) information transmitted from the communication device 20 (in this case, the second communication device 22).

[0105] The estimation means 45 has the function of estimating the position of a subject T (moving object) within a spatial SP (predetermined area) based on information regarding the received signal strength. This makes it possible to estimate the position of the subject T based on the received signal strength of the signals transmitted and received between the first communication device 21 located within the spatial SP and the second communication device 22 located on the subject T.

[0106] Furthermore, the estimation means 45 may, if the received signal strength for the second communication device 22 installed on the subject T (moving body) is higher than the received signal strength for at least one other second communication device 22 present in the spatial SP (predetermined area), select at least two second communication devices 22, including the second communication device 22 installed on the subject T, and test whether there is a significant difference in the average value of the received signal strength over a predetermined period between the selected second communication devices 22. If a significant difference is found, the location of the second communication device 22 installed on the subject T may be estimated as the location of the subject T. This makes it possible to estimate the location of the second communication device 22 installed on the subject T (i.e., the location of the subject T) based on whether there is a significant difference between the received signal strength corresponding to a specific second communication device 22 and the received signal strength corresponding to other second communication devices 22 (i.e., whether the difference in received signal strength is accidental), rather than simply comparing the magnitudes of the received signal strengths corresponding to each of the multiple second communication devices 22. This makes it possible to accurately estimate the position of the subject T while reducing the effects of, for example, radio wave fluctuations.

[0107] The function of the estimation means 45 is implemented, for example, as follows. The CPU 31 of the identification device 30 uses the received signal strength (RSSI) acquired based on the function of the second acquisition means 44 to determine the distance between the first communication device 21 and the second communication device 22, and consequently, the position of the second communication device 22 in space SP (i.e., the position of the subject T). Specifically, the distance between the first communication device 21 and the second communication device 22 can be calculated, for example, by using the following equations (1) and (2). P r =P t +G r +G t -L …(1)

number

[0108] Thereby, it becomes possible to obtain the distance between the first communication device 21 and the second communication device 22 by using the received signal strength (RSSI) of the signal between the first communication device 21 and the second communication device 22. Further, when a plurality of first communication devices 21 are provided in the space SP, the position of the second communication device 22 in the space SP (that is, the position of the target person T) can be obtained by using the distances between each of the plurality of first communication devices 21 and the second communication device 22.

[0109] Also, in the present embodiment, the CPU 31 of the identification device 30 estimates the position of the target person T according to, for example, the flowchart shown in FIG. 9. Referring to the flowchart of FIG. 9, the CPU 31 of the identification device 30 accesses the second acquired data and selects at least two second communication devices 22 including the second communication device 22 with the highest received signal strength (RSSI) corresponding to the radio signal among the plurality of second communication devices 22 existing in the space SP (step S100). Here, it is assumed that the received signal strength corresponding to the second communication device 22 provided for the target person T (for example, "user A") is the highest.

[0110] Next, the CPU 31 of the identification device 30 determines whether the received signal strength values ​​corresponding to the two selected second communication devices 22 follow a normal distribution (i.e., whether they are normal) by performing a normality test (for example, the Kolmogorv-Smirnov test or the Shapiro-Wilk test) (step S102).

[0111] If the CPU 31 of the identification device 30 determines that the data is normal (step S102: YES), it performs a parametric test. Specifically, the CPU 31 determines whether the received signal strength values ​​corresponding to the two selected second communication devices 22 have equal variances, for example by performing an F-test (step S104). Furthermore, if the CPU 31 determines that the data has equal variances (step S104: YES), it determines whether there is a significant difference in the mean values ​​of the received signal strengths between the two selected second communication devices 22, for example by performing a t-test (step S106). Also, if the CPU 31 determines that the data does not have equal variances (step S104: NO), it determines whether there is a significant difference in the mean values ​​of the received signal strengths within a predetermined period between the two selected second communication devices 22, for example by performing a Welch's t-test (step S108).

[0112] Furthermore, if the CPU 31 of the identification device 30 determines that the signal is not normal (step S102: NO), it determines whether there is a significant difference in the average value of the received signal strength over a predetermined period between the two selected second communication devices 22 by performing a non-parametric test such as the Mann-Whitney U test or the Wilcoxon rank-sum test (step S110).

[0113] Next, in the processing of step S106, step S108, or step S110, if the CPU 31 of the identification device 30 determines that there is a significant difference in the average value of the received signal strength over a predetermined period between the two selected second communication devices 22 (step S112: YES), it may estimate that the second communication device 22 with the highest corresponding received signal strength among the two selected second communication devices 22 (i.e., the second communication device 22 provided to subject T (user A)) is the second communication device 22 located closest to the first communication device 21 (step S114). Alternatively, if the CPU 31 of the identification device 30 determines that there is no significant difference in the average value of the received signal strength over a predetermined period (step S112: NO), it may determine that the second communication device 22 with the highest received signal strength (i.e., the second communication device 22 provided to subject T (user A)) is not located closest to the first communication device 21 and terminate the processing.

[0114] In this way, based on information regarding the received signal strength of the signals transmitted and received between the first communication device 21 and the second communication device 22, the location of the second communication device 22 closest to the first communication device 21 (the second communication device 22 provided to the target person T (user A)) can be estimated.

[0115] The generation means 46 has a function to generate training data for individual identification of a subject T (moving body) based on an image of the subject T, by assigning identification information of the subject T as a label to at least one feature point of the subject T when the position of the subject T (moving body) in a spatial SP (predetermined area) satisfies predetermined conditions.

[0116] Here, the predetermined conditions may include the shortest Euclidean distance between the first communication device 21 located in the spatial SP (predetermined area) and the second communication device 22 located for the target person T (for example, "User A") among a plurality of second communication devices 22 located in the spatial SP. This makes it possible to automatically generate training data for individual identification of target person T based on an image of target person T, taking into account the Euclidean distance between the first communication device 21 located in the spatial SP and the second communication device 22 located for target person T within the spatial SP.

[0117] Furthermore, the generation means 46 may, when the position of the subject T (e.g., "User A") (moving body) within a spatial SP (predetermined area) satisfies predetermined conditions, assign the identification information of subject T (User A) as a label to each coordinate of at least one part of the transformed subject T (User A). This makes it possible to generate training data using the coordinates of each at least one part of subject T that has been transformed according to the distance between at least two parts of subject T (for example, the measured distance between at least two parts of subject T, or the measured distance between at least two parts of subject T in a previously captured image).

[0118] The function of the generation means 46 is realized, for example, as follows. Here, two training data generation methods will be described as examples. First, regarding the first training data generation method, the CPU 31 of the identification device 30 generates the first training data shown in Figure 11, for example, according to the flowchart shown in Figure 10.

[0119] The CPU 31 accesses the second acquired data, for example, every predetermined period (e.g., 1 second), and uses the received signal strength (RSSI) received (acquired) from each of the one or more second communication devices 22 during that predetermined period, along with the above-mentioned equations (1) and (2), to calculate the Euclidean distance between the first communication device 21 and each of the one or more second communication devices 22 during that predetermined period (step S200). Next, the CPU 31 selects the second communication device 22 corresponding to the shortest Euclidean distance from among the calculated Euclidean distances (step S202). Here, it is assumed that the second communication device 22 possessed or worn by the subject T (e.g., "User A") is selected as the second communication device 22 corresponding to the shortest Euclidean distance.

[0120] Next, the CPU 31 determines whether the selected second communication device 22 is estimated to be a second communication device 22 located close to the first communication device 21 (step S204). Here, the processing content of step S204 may be the same as the processing content of steps S100 to S114 shown in Figure 9.

[0121] Then, if the selected second communication device 22 is estimated to be a second communication device 22 adjacent to the first communication device 21 (step S204: YES), the CPU 31 stores the identification information of subject T (user A) as a label for at least one feature point (here, the coordinates of at least one body part) of subject T (user A) in the first training data (step S206). Here, for example, if the image contains multiple people, the CPU 31 may extract at least one feature point of the person closest to the first communication device 21 in the image from the first acquired data as at least one feature point of subject T (user A). Alternatively, the CPU 31 may extract the identification information of subject T associated with the second communication device 22 estimated to be a second communication device 22 adjacent to the first communication device 21 from the second acquired data as the identification information of subject T (user A).

[0122] The first training data shown in Figure 11 is data that describes the relationship between the coordinates of at least one part of subject T (at least one feature point of subject T) and the identification information of subject T that has been assigned as a label. As a result of machine learning, a trained model is constructed that shows the relationship between information about the captured image of subject T and the identification information of subject T.

[0123] Furthermore, if the selected second communication device 22 is not estimated to be a second communication device 22 adjacent to the first communication device 21 (step S204: NO), the CPU 31 does not have to assign a label to at least one feature point (in this case, the coordinates of at least one body part) of the subject T (user A) (i.e., it does not have to store at least one feature point of the subject T (user A) in the first training data).

[0124] In this way, the CPU 31 can generate first training data for individual identification of subject T (user A) based on an image of subject T (user A) by assigning identification information of subject T (user A) as a label to at least one feature point of subject T (user A) when the position of subject T (user A) in spatial SP satisfies a predetermined condition (in this case, the shortest Euclidean distance between the first communication device 21 provided in spatial SP and the second communication device 22 provided for subject T (user A) among the multiple second communication devices 22 present in spatial SP).

[0125] Next, the second method for generating training data will be described. The CPU 31 of the identification device 30 generates the second training data shown in Figure 13, for example, according to the flowchart shown in Figure 12. Note that the processing content of steps S200 to S208 shown in Figure 12 is the same as the processing content of steps S200 to S208 shown in Figure 10. Furthermore, here we assume that tracking information (tracking ID) of each person contained in the image is associated with each image data in the first acquired data.

[0126] Here, the generation means 46 may assign identification information of subject T as a label to at least one feature point and tracking information of subject T when the position of subject T (moving body) within the spatial SP (predetermined area) satisfies predetermined conditions. This makes it possible to efficiently assign identification information of subject T to at least one feature point of a predetermined person (i.e., a person whose position within the spatial SP satisfies predetermined conditions (subject T)) by using the tracking information of each person, even if multiple people are included in the captured image.

[0127] Furthermore, the predetermined conditions may include the condition that no other second communication devices 22 exist within a predetermined range from the position of the second communication device 22 installed on the subject T (moving body). This makes it possible to suppress the generation of inaccurate training data, such as incorrectly assigning identification information of another person corresponding to another second communication device 22 to at least one feature point of subject T detected from the image of subject T.

[0128] Furthermore, the predetermined conditions may include the state in which the Euclidean distance between the first communication device 21 and the second communication device 22 provided on the subject T (moving object) remains below a predetermined value for a predetermined period of time. As a result, for example, the longer the state in which the Euclidean distance between the first communication device 21 and the second communication device 22 provided on the subject T remains short, the easier it becomes to assign identification information of subject T to at least one feature point of subject T, thereby improving the efficiency of generating accurate training data.

[0129] Furthermore, if the identification information of the subject T (moving object) has already been assigned to the assigned tracking information, the generation means 46 may also assign the already assigned identification information of the subject T as a label to the assigned tracking information. This makes it possible to further improve the efficiency of generating training data, as the already assigned identification information of the subject T is automatically assigned to the assigned tracking information when the identification information of the subject T has already been assigned to the assigned tracking information.

[0130] In this case, the function of the generation means 46 is realized, for example, as follows. In Figure 12, when the selected second communication device 22 is estimated to be a second communication device 22 adjacent to the first communication device 21 (step S204: YES), the CPU 31 determines whether the identification information of the subject T (user A) has already been associated (assigned) to the tracking information of the subject T (e.g., "user A") who possesses or wears the estimated second communication device 22 (step S210). Here, the CPU 31 may also access the second learning data, which will be described later, to determine whether the identification information of the subject T has already been associated to the same tracking information (tracking ID) corresponding to the subject T (user A).

[0131] Then, if the CPU 31 has already associated the identification information of subject T (user A) with the tracking information of subject T (user A) who possesses or wears the estimated second communication device 22 (step S210: YES), it may proceed to the processing in step S206. In this way, if the identification information of subject T (user A) has already been assigned to the assigned tracking information, it becomes possible to assign the already assigned identification information of subject T (user A) as a label to the assigned tracking information.

[0132] On the other hand, if the tracking information of the subject T (user A) possessing or wearing the estimated second communication device 22 is not associated with the identification information of subject T (user A) (step S210: NO), the CPU 31 determines whether or not another second communication device 22 exists within a predetermined range (for example, within 1 meter) from the location of the estimated second communication device 22 (step S212). Here, the CPU 31 may, for example, access the second acquired data and calculate the Euclidean distance between the estimated second communication device 22 and each of the other second communication devices 22 based on the received signal strength (RSSI) to determine whether or not another second communication device 22 exists within a predetermined range from the location of the estimated second communication device 22.

[0133] The CPU 31 may proceed to step S216 if there are no other second communication devices 22 within a predetermined range from the estimated location of the second communication device 22 (step S212: YES). On the other hand, the CPU 31 may proceed to step S208 if there are other second communication devices 22 within a predetermined range from the estimated location of the second communication device 22 (step S212: NO).

[0134] Next, the CPU 31 determines whether the state in which the Euclidean distance between the first communication device 21 and the same second communication device 22 (estimated second communication device 22) is less than a predetermined value (e.g., within 2 meters) has been maintained for a predetermined period (e.g., 10 seconds) (step S216). Here, the CPU 31 may, for example, access the second acquired data and calculate the Euclidean distance between the first communication device 21 and the estimated second communication device 22 based on the received signal strength (RSSI), thereby determining whether the state in which the calculated Euclidean distance is less than a predetermined value has been maintained for a predetermined period.

[0135] Then, if the Euclidean distance between the first communication device 21 and the estimated second communication device 22 remains below a predetermined value for a predetermined period of time (step S216: YES), the CPU 31 stores the identification information of subject T (user A) as a label for at least one feature point (here, the coordinates of at least one body part) of subject T (user A) who possesses or wears the estimated second communication device 22 in the second training data (step S218). Here, for example, if the image contains multiple people, the CPU 31 may extract at least one feature point of the person closest to the first communication device 21 in the image from the first acquired data as at least one feature point of subject T (user A). Alternatively, the CPU 31 may extract the identification information of subject T associated with the second communication device 22, which is estimated to be the second communication device 22 adjacent to the first communication device 21, from the second acquired data as the identification information of subject T (user A).

[0136] The second training data shown in Figure 13 is data that describes the relationship between the coordinates of at least one part of subject T (at least one feature point of subject T), the tracking information of subject T (tracking ID), and the identification information of subject T assigned as a label. As a result of machine learning, a trained model is constructed that shows the relationship between information about the captured image of subject T and the identification information of subject T.

[0137] Furthermore, if the CPU 31 does not maintain a state where the Euclidean distance between the first communication device 21 and the estimated second communication device 22 is less than a predetermined value for a predetermined period of time (step S216: NO), it may proceed to the processing in step S208.

[0138] In this way, the CPU 31 can generate second training data for individual identification of subject T (user A) based on an image of subject T (user A) by assigning the identification information of subject T (user A) as a label to at least one feature point of subject T (user A) when the position of subject T (user A) in spatial SP satisfies predetermined conditions (here, no other second communication device 22 exists within a predetermined range from the estimated position of the second communication device 22, and the state in which the Euclidean distance between the first communication device 21 and the estimated second communication device 22 is less than a predetermined value is maintained for a predetermined period of time).

[0139] Furthermore, if the tracking information of subject T is changed due to, for example, occlusion, the CPU 31 does not need to assign the identification information of subject T as a label to at least one feature point of subject T until a new second communication device 22 adjacent to the first communication device 21 is estimated.

[0140] The learning means 47 has the function of learning a model used to identify individual subject T (moving object) based on images of subject T, using machine learning with learning data generated by the learning data generation system.

[0141] The functions of the learning means 47 are realized, for example, as follows. When a predetermined model learning instruction is input to the CPU 31 of the identification device 30 using the first learning data shown in Figure 11 or the second learning data shown in Figure 13, the CPU 31 learns the model using the first learning data shown in Figure 11 or the second learning data shown in Figure 13. The CPU 31 may learn using, for example, a time-series-responsive neural network model. Here, as the time-series-responsive neural network, for example, an RNN (Recurrent Neural Network) or an advanced version of RNN such as LSTM (Long Short-Term Memory) can be applied. The CPU 31 may also learn using any of several models, such as a graph neural network (GNN) model, a convolutional neural network (CNN) model, a support vector machine (SVM) model, a fully connected neural network (FNN) model, a gradient boosting (HGB) model, or a wavenet (WN) model. In this embodiment, it is possible to improve individual identification accuracy by using one of the following GNN derivatives: a graph convolutional neural network (GCN) model, a graph attention network (GAT) model, or a graph convolutional LSTM (GC-LSTM) model.

[0142] The image acquisition means 48 has the function of acquiring images of a subject T (moving body) located in a spatial SP (predetermined area) that have been captured by the imaging device 10.

[0143] The functions of the image acquisition means 48 are realized, for example, as follows. First, the imaging device 10 performs imaging processing at a predetermined frame rate (e.g., 30fps) when, for example, a subject T is present in space SP, and each time imaging processing is performed, it transmits the image data of the captured image to the identification device 30 via the communication network NW. Here, the image data of the image captured by the imaging device 10 may be transmitted to the identification device 30 in association with the date and time of capture and the identification information of the imaging device 10 (e.g., the serial number or MAC (Media Access Control) address of the imaging device 10).

[0144] Meanwhile, the CPU 31 of the identification device 30 stores the received image data in, for example, the RAM 33 or storage device 34, each time it receives (acquires) image data transmitted from the imaging device 10 via the communication interface unit 38, in association with the date and time the image was captured. In this way, the image acquisition means 48 can acquire images of the subject T captured by the imaging device 10.

[0145] Furthermore, each time image data is stored in, for example, RAM 33 or storage device 34, the CPU 31 may detect the position of at least one part (which may include joints) of the subject T in the image data based on the function of the detection means 42 described above, and store information regarding the detected position in, for example, RAM 33 or storage device 34 in association with the image data. In addition, the CPU 31 may transform the coordinates of at least one part of the subject T in the image data based on the function of the conversion means 43 described above.

[0146] The identification means 49 has the function of performing individual identification of subject T (moving object) based on an image of subject T and a trained model based on machine learning using training data generated by the training data generation system.

[0147] The function of the identification means 49 may be realized, for example, as follows: Based on the function of the image acquisition means 48, the CPU 31 of the identification device 30 acquires an image of a subject T located in a predetermined area, and then performs individual identification of the subject T by inputting the coordinates of each of at least one part of the subject T in the image into a trained model learned by the learning system.

[0148] The CPU 31 may also present the individual identification result of the subject T. For example, if the CPU 31 performs individual identification of the subject T based on the function of the identification means 49, it may display information regarding the individual identification result of the subject T on, for example, the display unit 36. Here, the information regarding the individual identification result of the subject T may consist of text data or image data, etc. Also, if the information regarding the individual identification result of the subject T consists of audio data, the CPU 31 may output the information regarding the individual identification result of the subject T from an audio output device such as a speaker. Furthermore, the CPU 31 may transmit the information regarding the individual identification result of the subject T to another computer (for example, a server, etc.) connected to the identification device 30 via a communication network NW.

[0149] (4) Main processing flow of the learning data generation system of this embodiment Next, an example of the main processing flow performed by the learning data generation system of this embodiment will be explained with reference to the flowchart in Figure 14.

[0150] The CPU 31 of the identification device 30 acquires an image of the subject T (moving body) located in a spatial SP (predetermined area) captured by the imaging device 10, based on the function of the first acquisition means 41 (step S300). Next, the CPU 31 detects at least one feature point of the subject T (moving body) based on the acquired image, based on the function of the detection means 42 (step S302).

[0151] Next, the CPU 31 acquires information regarding the position of the subject T (moving body) within the spatial SP (predetermined area) based on the function of the second acquisition means 44 (step S304). Then, based on the function of the generation means 46, the CPU 31 generates training data for individual identification of the subject T based on an image of the subject T by assigning identification information of the subject T as a label to at least one feature point of the subject T when the position of the subject T (moving body) within the spatial SP (predetermined area) satisfies predetermined conditions (step S306).

[0152] In this way, machine learning, using information about at least one feature point of subject T in an image as training data, makes it possible to train a model used to identify individuals.

[0153] (5) Main processing flow of the model learning system of this embodiment Next, an example of the main processing flow performed by the model learning system of this embodiment will be explained with reference to the flowchart in Figure 15.

[0154] The CPU 31 of the identification device 30 learns a model used for individual identification of subject T (moving object) based on an image of subject T, using machine learning with the learning data generated by the learning data generation system, based on the functions of the learning means 47 (step S400).

[0155] In this way, it becomes possible to perform individual identification of subject T based on information about each of the at least one feature points of subject T detected based on the image of subject T captured by the imaging device 10, and a trained model.

[0156] (6) Main processing flow of the individual identification system of this embodiment Next, an example of the main processing flow performed by the individual identification system of this embodiment will be explained with reference to the flowchart in Figure 16.

[0157] The CPU 31 of the identification device 30 acquires an image of a subject T (moving body) located in a spatial SP (predetermined area) captured by the imaging device 10, based on the function of the image acquisition means 48 (step S500). Next, the CPU 31 performs individual identification of the subject T (moving body) based on the image of the subject T (moving body) and a trained model based on machine learning using training data generated by the training data generation system, based on the function of the identification means 49 (step S502).

[0158] In this way, it becomes possible to perform individual identification of subject T based on information about each of the at least one feature points of subject T detected based on the image of subject T captured by the imaging device 10, and a trained model.

[0159] As described above, according to the learning data generation system, learning data generation method, and program of this embodiment, when the position of subject T within a predetermined area satisfies predetermined conditions, identification information of subject T is assigned as a label to at least one feature point of subject T detected from the image of subject T. Therefore, it becomes possible to automatically generate learning data for individual identification of subject T based on the image of subject T, taking into account the position of subject T within a predetermined area. This makes it easier to generate learning data compared to, for example, generating learning data through manual work. Furthermore, it becomes possible to eliminate human errors when generating learning data, thus enabling the generation of appropriate learning data.

[0160] Furthermore, according to the model learning system, model learning method, and program of this embodiment, it becomes possible to learn a model used for individual identification by machine learning using information about at least one feature point of a subject T in an image as training data. Therefore, individual identification can be easily performed by using this model.

[0161] Furthermore, according to the individual identification system, individual identification method, and program of this embodiment, individual identification of subject T is performed based on information regarding each of at least one feature point of subject T detected based on the image of subject T captured by the imaging device 10, and a trained model. Therefore, individual identification of subject T can be easily performed by using the image of subject T captured by the imaging device 10.

[0162] The program of the present invention may be stored on a computer-readable storage medium. The storage medium on which this program is recorded may be the ROM 32, RAM 33, or storage device 34 of the identification device 30 shown in Figure 2. Alternatively, the storage medium may be a CD-ROM or the like that can be read by being inserted into a program reading device such as a CD-ROM drive. Furthermore, the storage medium may be magnetic tape, cassette tape, flexible disk, MO / MD / DVD, or semiconductor memory.

[0163] The embodiments described above are provided to facilitate understanding of the present invention and are not intended to limit it. Accordingly, each element disclosed in the above embodiments is intended to include all design modifications and equivalents that fall within the technical scope of the present invention.

[0164] In the embodiments described above, the case where the moving object is a person (subject T) was explained as an example, but it is not limited to this case. The moving object can be anything that can be the subject of individual identification, for example, it may be a living organism other than a person, or it may be an object such as a robot or a product.

[0165] Furthermore, in the embodiment described above, the case in which the neck, right shoulder, left shoulder, right buttock, and left buttock of subject T (moving body) are detected as body parts of subject T was explained as an example, but other body parts (for example, head, hands, feet, joints, etc.) may also be detected.

[0166] Furthermore, although the above-described embodiment explained the case where individual identification of a single moving object (subject T) is performed as an example, individual identification of multiple moving objects may be performed simultaneously.

[0167] Furthermore, although the above-described embodiment explained the case in which one identification device 30 is provided as an example, it is not limited to this case. For example, multiple identification devices 30 may be provided, in which case the operation content and processing results on any of the identification devices 30 may be displayed in real time on the other identification devices 30, or the processing results on any of the identification devices 30 may be shared among the multiple identification devices 30.

[0168] Furthermore, in the embodiments described above, the case in which the coordinates of each of at least one part of the subject T in the image correspond to the "at least one feature point of a moving body" of the present invention was explained as an example, but the present invention is not limited to this case. The "at least one feature point of a moving body" may be any one that can be used for individual identification of a moving body.

[0169] Furthermore, in the above-described embodiment, the identification device 30 is configured to realize the functions of the first acquisition means 41, detection means 42, conversion means 43, second acquisition means 44, estimation means 45, generation means 46, learning means 47, image acquisition means 48, and identification means 49, but the configuration is not limited to this. For example, a learning device 50 (shown in Figure 17) may be provided, which is composed of a computer (for example, a general-purpose personal computer or server computer) that is connected to the identification device 30 in a communication network such as the Internet or LAN, and which is used to learn a model used to identify individuals. In this case, the identification device 30 and the learning device 50 can adopt substantially the same hardware configuration, so that the learning device 50 can realize the function of at least one of the means 41 to 49 described in the above embodiment. For example, each function of the functional block diagram shown in Figure 3 may be arbitrarily divided between the identification device 30 and the learning device 50, as shown in Figures 17(a) and (b).

[0170] Furthermore, the function of at least one of the above means 41 to 49 may be realized by the imaging device 10 and / or the communication device 20. [Industrial applicability]

[0171] The learning data generation system, model learning system, individual identification system, learning data generation method, model learning method, individual identification method, and program of the present invention described above can easily generate learning data for individual identification based on images, and furthermore, can easily perform individual identification based on such learning data. Therefore, they can be suitably used, for example, in systems for individual identification of moving objects (e.g., people such as workers or living organisms other than people), information provision systems that provide appropriate information according to the identified individual, and monitoring systems for moving objects (e.g., hospitalized patients, residents of facilities, pets, etc.). Accordingly, their industrial applicability is extremely large. [Explanation of symbols]

[0172] 10…Imaging device 20...Communication equipment 21...First communication device 22...Second communication device 30…Identification device 41…First acquisition means 42...Detection means 43...Conversion method 44…Second acquisition means 45...Estimation means 46…Generation means 47…Learning methods 48…Method for acquiring images 49... Identification means 50…Learning device T…Target person

Claims

1. A first acquisition means for acquiring an image of a moving object located in a predetermined area, captured by an imaging device, A detection means for detecting at least one feature point of the moving object based on the acquired image, A second acquisition means for acquiring information regarding the position of the moving body within the predetermined area, The system includes a generation means for generating learning data for individual identification of a moving object based on an image of the moving object, by assigning identification information of the moving object as a label to at least one feature point of the moving object when the position of the moving object within the predetermined area satisfies predetermined conditions. A system for generating training data.

2. Each time the detection means detects at least one feature point of one or more moving bodies, it assigns individual tracking information to the detected at least one feature point of the moving body. The learning data generation system according to claim 1, wherein the generation means assigns identification information of the moving body as a label to at least one feature point of the moving body and corresponding tracking information when the position of the moving body within the predetermined area satisfies the predetermined conditions.

3. The learning data generation system according to claim 2, wherein the predetermined condition includes that no other second communication device exists within a predetermined range from the position of the second communication device provided on the moving body.

4. The learning data generation system according to claim 2, wherein the predetermined condition includes the state that the Euclidean distance between the first communication device and the second communication device provided on the moving body is less than a predetermined value for a predetermined period of time.

5. The learning data generation system according to claim 2, wherein the generation means, when the identification information of the moving body has already been assigned to the assigned tracking information, assigns the identification information of the moving body that has already been assigned to the assigned tracking information as a label.

6. The learning data generation system according to any one of claims 1 to 5, wherein the predetermined condition includes the shortest Euclidean distance between a first communication device provided in the predetermined area and a second communication device provided on the moving body among a plurality of second communication devices present in the predetermined area.

7. The learning data generation system according to claim 1, wherein the detection means detects the coordinates of each of at least one part of the moving body in the image as at least one feature point of the moving body.

8. The system includes a transformation means for transforming the coordinates of each of at least one of the detected moving parts in the image based on the distance between at least two parts of the detected moving part, The learning data generation system according to claim 7, wherein the generation means assigns identification information of the moving body as a label to each coordinate of at least one part of the transformed moving body when the position of the moving body within the predetermined area satisfies predetermined conditions.

9. The learning data generation system according to claim 8, wherein the conversion means converts the coordinates in the image of each of at least one detected part of the moving body based on the distance between the neck portion and the waist portion of the moving body.

10. The learning data generation system according to any one of claims 7 to 9, wherein the detection means detects at least one joint of the moving body as at least one part of the moving body.

11. The learning data generation system according to claim 1, wherein the second acquisition means acquires information regarding the position of the moving body measured using a communication device.

12. The learning data generation system according to claim 11, wherein the second acquisition means acquires information regarding the received signal strength of signals transmitted and received between a first communication device provided in a predetermined area and a second communication device provided on the moving body, as information regarding the position of the moving body.

13. The learning data generation system according to claim 12, further comprising estimation means for estimating the position of the moving object within a predetermined area based on the information regarding the received signal strength.

14. The learning data generation system according to claim 13, wherein the estimation means selects at least two second communication devices, including the second communication device provided on the moving body, when the received signal strength for the second communication device provided on the moving body is higher than the received signal strength for at least one other second communication device located in the predetermined area, tests whether there is a significant difference in the average value of the received signal strength over a predetermined period between the selected second communication devices, and if such a significant difference is found, estimates the position of the second communication device provided on the moving body as the position of the moving body.

15. The system comprises a learning means for learning a model used to identify individual moving objects based on images of moving objects by machine learning using the learning data generated by the learning data generation system of claim 1. Model learning system.

16. Image acquisition means for acquiring images of moving objects in a predetermined area, captured by an imaging device, The system comprises an identification means for performing individual identification of the moving object based on an image of the moving object and a trained model based on machine learning using training data generated by the training data generation system of claim 1. Individual identification system.

17. Computers The steps include acquiring an image of a moving object located in a predetermined area, captured by an imaging device, The steps include detecting at least one feature point of the moving object based on the acquired image, The steps include: obtaining information regarding the position of the moving body within the predetermined area; The steps include generating learning data for individual identification of the moving object based on an image of the moving object by assigning identification information of the moving object as a label to at least one feature point of the moving object when the position of the moving object within the predetermined area satisfies predetermined conditions, Perform each step Method for generating training data.

18. Computers The method for generating training data according to claim 17 involves performing a step of training a model used for individual identification of moving objects based on images of moving objects, using machine learning with the training data generated by the training data generation method of claim 17. Model learning methods.

19. Computers The steps include acquiring an image of a moving object located in a predetermined area, captured by an imaging device, A step of performing individual identification of the moving object based on the image of the moving object and a trained model based on machine learning using the training data generated by the training data generation method of claim 17, Perform each step Individual identification method.

20. On the computer, A function to acquire images of moving objects in a predetermined area, captured by an imaging device, A function to detect at least one feature point of the moving object based on the acquired image, A function to acquire information regarding the position of the moving object within the predetermined area, A function for generating learning data for individual identification of a moving object based on an image of the moving object, by assigning identification information of the moving object as a label to at least one feature point of the moving object when the position of the moving object within the predetermined area satisfies predetermined conditions, A program to achieve this.

21. On the computer, A program for realizing a function to learn a model used for individual identification of moving objects based on images of moving objects by machine learning using training data generated by the program of claim 20.

22. On the computer, A function to acquire images of moving objects in a predetermined area, captured by an imaging device, A function for performing individual identification of the moving object based on the image of the moving object and a trained model based on machine learning using training data generated by the program of claim 20, A program to achieve this.

Citation Information

Patent Citations

  • Pet individual identification system

    JP2019071895A