Data collection device and data collection method

The data collection device addresses the burden of generating learning data for data recognition models by using two recognition units to automatically collect data points with differing recognition results, thereby reducing user intervention and improving efficiency.

JP7690957B2Active Publication Date: 2025-06-11KONICA MINOLTA INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022533823
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-03
Filing Date
2021-06-16
Publication Date
2025-06-11
Estimated Expiration
2041-06-16

AI Technical Summary

Technical Problem

Existing methods for generating learning data for data recognition models, such as image and speech recognition, require significant manual effort and time, especially when dealing with multiple installation sites, leading to a substantial burden on users.

Method used

A data collection device and method that utilize two recognition units with different processing times and calculation scales, where the first unit performs regular recognition and the second unit performs recognition at less frequent intervals, allowing for the collection of learning data with reduced user intervention.

Benefits of technology

This approach reduces the manual effort required for generating learning data by automatically collecting data points where the recognition results differ between the two units, thereby improving the efficiency of data collection for data recognition models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690957000001
    Figure 0007690957000001
  • Figure 0007690957000002
    Figure 0007690957000002
  • Figure 0007690957000003
    Figure 0007690957000003
Patent Text Reader

Abstract

Provided is a data collection device which allows a load for the user for generating learning data for a data recognition model to be reduced. An image recognition device 100 comprises a first CNN 130, a second CNN 140, a recognition result comparison part 150 for comparing a recognition result of the first CNN 130 and a recognition result of the second CNN 140 with respect to an input image, and a data collection part 160 for collecting the input image as learning data depending on a comparison result of the recognition result comparison part 150.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for collecting learning data used for learning a data recognition model.

Background Art

[0002] Conventionally, an image recognition system that recognizes the position and state of an object such as a person or a car from an image using machine learning is known.

[0003] For example, in Patent Document 1, a method has been proposed in which environmental dependence attributes according to the installation site are given to learning data, and learning is performed using the learning data having the specified environmental dependence attributes, thereby performing learning specialized for the installation site.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In Patent Document 1, in order to generate learning data, it is necessary for a user to label the images taken at the installation site. However, in order to improve the recognition performance of the data recognition model, a large amount of learning data is required, and in order to prepare a sufficient amount of learning data for each of a plurality of installation sites, there is a problem that the man-hours required for the user to perform the work become enormous.

[0006] Moreover, the same problem exists even when performing speech recognition or natural language processing using machine learning.

[0007] The present disclosure has been made in view of the above problems, and an object thereof is to provide a data collection device and a data collection method capable of reducing the burden on a user related to the generation of learning data for a data recognition model.

Means for Solving the Problems

[0008] A data collection device according to an aspect of the present disclosure is a data collection device that collects learning data for a data recognition model, and includes a first recognition unit, a second recognition unit different from the first recognition unit, a comparison unit that compares a recognition result of the first recognition unit with a recognition result of the second recognition unit for input data, a collection unit that collects the input data as learning data according to a comparison result of the comparison unit, and a timing determination unit that determines a timing for operating the second recognition unit. The first recognition unit requires less processing time to recognize the same object than the second recognition unit. The first recognition unit performs normal data recognition, and the timing determination unit determines the timing for operating the second recognition unit to be less frequent than the normal data recognition by the first recognition unit such that it becomes a fixed interval equal to or longer than the processing time of the second recognition unit at a timing, and the second recognition unit performs data recognition at the timing determined by the timing determination unit.

[0009] Further, the calculation scale of the first recognition unit may be smaller than the calculation scale of the second recognition unit.

[0010] Further, the collection unit may collect the recognition result of the second recognition unit as learning data indicating correct answer data for the input data.

[0011] Furthermore, a learning unit that performs additional learning of the first recognition unit using the learning data collected by the collection unit may be provided.

[0012] Further, the learning unit may correct the correct answer data according to an external input.

[0013] Further, the comparison unit may determine whether the recognition result of the first recognition unit is different from the recognition result of the second recognition unit, and the collection unit may collect the input data as learning data when the recognition result of the first recognition unit is different from the recognition result of the second recognition unit.

[0014] Further, the comparison unit may determine whether the difference between the recognition result of the first recognition unit and the recognition result of the second recognition unit is equal to or greater than a predetermined threshold value, and the collection unit may collect the input data as learning data when the difference is equal to or greater than the threshold value.

[0016] Further, the timing determination unit may determine the timing at fixed intervals.

[0017] Further, the timing determination unit may determine the timing according to the learning proficiency in the first recognition unit.

[0018] Further, the timing determination unit may determine the timing according to an external input.

[0019] Further, the data collection device may be composed of an edge terminal including the first recognition unit and a server terminal including the second recognition unit.

[0020] Furthermore, it may further include one or more second edge terminals, and the second edge terminal may be provided with a recognition unit having the same configuration as the first recognition unit.

[0021] The first recognition unit and the second recognition unit may each perform image recognition, voice recognition, or natural language recognition.

[0022] Another aspect of the present disclosure is a data collection method in which a computer executes a process of collecting learning data for a data recognition model, the method including: a first recognition step of obtaining a recognition result by a first recognition unit for input data; a second recognition step of obtaining a recognition result by a second recognition unit different from the first recognition unit for the input data; a comparison step of comparing the recognition result of the first recognition unit and the recognition result of the second recognition unit; a collection step of collecting the input data as learning data according to the comparison result of the comparison step; and a determination step of determining the timing for operating the second recognition unit, wherein the first recognition unit takes less processing time to recognize the same object than the second recognition unit, the first recognition unit performs constant data recognition, and the determination step determines the timing for operating the second recognition unit to be less frequent than the constant data recognition by the first recognition unit. such that it becomes a fixed interval equal to or longer than the processing time of the second recognition unit The timing is determined such that the second recognition unit performs data recognition at the timing determined in the determination step.

[0023] Also, it may be that the calculation scale of the first recognition unit is smaller than that of the second recognition unit.

[0024] Also, in the collection step, the recognition result of the second recognition unit may be collected as learning data indicating correct data for the input data.

[0025] Furthermore, additional learning of the first recognition unit may be performed using the learning data collected in the collection step.

[0026] Also, the comparison step determines whether the recognition result of the first recognition unit and the recognition result of the second recognition unit are different, and in the collection step, the input data may be collected as the learning data when the recognition result of the first recognition unit and the recognition result of the second recognition unit are different.

[0027] Further, the comparison step determines whether the difference between the recognition result of the first recognition unit and the recognition result of the second recognition unit is equal to or greater than a predetermined threshold value, and the collection step may collect the input data as learning data when the difference is equal to or greater than the threshold value.

Effect of the Invention

[0029] When one of the recognition results by the first recognition unit and the second recognition unit is correct and the other is incorrect, for the recognition unit that made a mistake, the input data is classified as FP (False Positive) or FN (False Negative). Here, FP means identifying that the input data contains a detection target even though it does not, and FN means, conversely, identifying that the input data does not contain a detection target even though it does. Generally, the purpose of learning in data recognition is to reduce such FPs and FNs. And an effective method for reducing FPs and FNs is to assign correct labels to the data classified as FPs and FNs to generate learning data and perform additional learning so that similar data can be correctly recognized. According to the data collection device according to the present disclosure, data classified as such FPs and FNs can be easily collected as learning data. Also, for the correct labeling, by using the recognition result of the recognition unit that was correct in the recognition, it is not necessary for the user to manually perform the correct labeling. Therefore, it is possible to reduce the burden on the user related to the generation of learning data.

Brief Description of the Drawings

[0030]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0031] 1. Embodiment 1 Hereinafter, the image recognition system 1 according to Embodiment 1 will be described.

[0032] 1.1 Configuration FIG. 1 is a block diagram showing the configuration of the image recognition system 1. As shown in the figure, the image recognition system 1 includes an image recognition device 100 and a camera 190. The image recognition device 100 includes a control unit 110, a non-volatile storage unit 120, a first CNN 130 (first recognition unit), a second CNN 140 (second recognition unit), a recognition result comparison unit 150 (comparison unit), a data collection unit 160 (collection unit), a timing adjustment unit 170 (timing determination unit), and an additional learning unit 180 (learning unit).

[0033] Here, the first CNN 130, the second CNN 140, the recognition result comparison unit 150, the data collection unit 160, the timing adjustment unit 170, and the additional learning unit 180 constitute a data collection device.

[0034] The camera 190 includes an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor field-effect transistor) image sensor or a CCD (Charge-Coupled Device) image sensor, and outputs an image of a predetermined size by converting the light imaged on the imaging element into an electrical signal by photoelectric conversion.

[0035] The camera 190 outputs images at a predetermined rate. For example, it outputs images at 30 FPS.

[0036] The control unit 110 is composed of a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random access memory), etc. In the RAM, computer programs and data stored in the ROM and the storage unit 120 are loaded, and the CPU operates according to the computer programs and data on the RAM, thereby realizing each processing unit (the first CNN 130, the second CNN 140, the recognition result comparison unit 150, the data collection unit 160, the timing adjustment unit 170, the additional learning unit 180).

[0037] As an example, the storage unit 120 is composed of a hard disk. The storage unit 120 may be composed of a non-volatile semiconductor memory. The storage unit 120 stores the first learning parameter 121, the second learning parameter 122, and the additional learning data 123. The additional learning data 123 includes learning images 123a and correct answer data 123b.

[0038] 1.2 CNN As an example of a convolutional neural network, the neural network 200 shown in FIG. 2 will be described.

[0039] (1) Structure of the neural network 200 As shown in this figure, the neural network 200 is a hierarchical neural network having an input layer 200a, a feature extraction layer 200b, and an identification layer 200c.

[0040] Here, a neural network is an information processing system that mimics the human neural network. In the neural network 200, an engineering neuron model corresponding to a nerve cell is herein called a neuron U. The input layer 200a, the feature extraction layer 200b, and the identification layer 200c are each composed of a plurality of neurons U.

[0041] The input layer 200a usually consists of a single layer. Each neuron U in the input layer 200a receives, for example, the pixel values of each pixel constituting an image. The received image values are directly output from each neuron U in the input layer 200a to the feature extraction layer 200b. The feature extraction layer 200b extracts features from the data received from the input layer 200a and outputs them to the identification layer 200c. The feature extraction layer 200b is sometimes called a backbone network. The identification layer 300c performs identification using the features extracted by the feature extraction layer 300b.

[0042] As the neuron U, usually, as shown in Fig. 3(a), an element with multiple inputs and one output is used. The signal is transmitted only in one direction, and the input signal xi (i = 1, 2, ···, n) is multiplied by a certain neuron weight value (SUwi) and input to the neuron U. The neuron weight value can be changed by learning. From the neuron U, after the sum of each input value (SUwi × xi) multiplied by the neuron weight value SUwi undergoes transformation by the activation function f(X), it is output. That is, the output value y of the neuron U is represented by the following mathematical formula.

[0043] y = f(X) Here, X = Σ(SUwi × xi) That is. As the activation function, for example, ReLU or the sigmoid function can be used.

[0044] As a learning method of the neural network 200, for example, an error is calculated using a predetermined error function from the value indicating the correct answer (teacher data) and the output value of the neural network 200, and the error backpropagation method (backpropagation) is used to sequentially change the neuron weight values of the feature extraction layer 200b and the neuron weight values of the identification layer 200c using the steepest descent method or the like so that this error becomes minimum.

[0045] (2) Learning process The learning process in the neural network 200 will be described.

[0046] The learning process is a process of training the neural network 200. Fig. 4(a) schematically shows the data propagation model of the learning process.

[0047] For each learning image 123a, it is input into the input layer 200a of the neural network 200 and output from the input layer 200a to the feature extraction layer 200b. In each neuron U of the feature extraction layer 200b, an operation with neuron weight values is performed on the input data, and the data indicating the extracted features is output to the discrimination layer 200c. In each neuron U of the discrimination layer 200c, an operation with neuron weight values is performed on the input data (step S11). Based on the above features, object estimation is performed. The data indicating the result of the object estimation is output from the discrimination layer 200c.

[0048] The output value of the discrimination layer 200c is compared with the teacher data (correct data) 123b, and an error (loss) is calculated using a predetermined error function (step S12). The neuron weight values of the discrimination layer 200c and the neuron weight values of the feature extraction layer 200b are sequentially changed (backpropagation) so that this error is reduced (step S13). Thereby, the neural network 200 is trained.

[0049] (3) Learning Results The learning results are stored in the storage unit 120 as learning parameters. Fig. 3(b) shows the data structure of the learning parameters stored in the storage unit 120. As shown in Fig. 3(b), the learning parameter 210 is composed of a plurality of neuron information 211. Each neuron information 211 corresponds to each neuron U of the feature extraction layer 200b and the discrimination layer 200c.

[0050] Each neuron information 211 includes a neuron number 212 and a neuron weight value 213.

[0051] Neuron number 212 is a number that identifies each neuron U in the feature extraction layer 200b and the discrimination layer 200c.

[0052] Neuron weight value 213 is the neuron weight value of each neuron U in the feature extraction layer 200b and the discrimination layer 200c, respectively.

[0053] In this way, the learned model is called a data recognition model. The data recognition model is used to identify the objects contained in the data.

[0054] (4) Estimation step The estimation step in the neural network 200 will be described.

[0055] Figure 4(b) shows the data propagation model when object estimation is performed using the neural network 200 learned by the above learning step, with the image data obtained by the camera 190 as the input.

[0056] In the estimation step in the neural network 200, feature extraction and object estimation are performed using the learned feature extraction layer 200b and the learned discrimination layer 200c (step S14).

[0057] (5) First CNN 130, Second CNN 140 The image recognition system 1 includes two image recognizers (first CNN 130, second CNN 140). The first CNN 130 and the second CNN 140 are image recognizers that perform image recognition, for example, for person detection. If a person is detected in the image input from the camera 190, a recognition result indicating that a person is included is output, and if not detected, a recognition result indicating that no person is included is output.

[0058] The first CNN 130 and the second CNN 140 have the same configuration as the neural network 200. Depending on the scale of computation, even when learning with the same training data, the recognition speed (the time required to recognize one image) and the recognition accuracy (the accuracy of correctly recognizing the input image) of the CNNs will be different. The scale of computation varies depending on the CNN algorithm and the number of stages of the backbone network. Therefore, the recognition speed and recognition accuracy differ depending on the CNN algorithm. Also, even with the same algorithm, if the number of stages of the backbone network is different, the recognition speed and recognition accuracy will be different. Generally, the larger the scale of computation, the higher the recognition accuracy but the slower the recognition speed tends to be. Conversely, the smaller the scale of computation, the faster the recognition speed but the lower the recognition accuracy tends to be.

[0059] The larger the number of stages of the backbone network, the larger the scale of computation, and the smaller the number of stages of the backbone network, the smaller the scale of computation.

[0060] The first CNN 130 and the second CNN 140 have different scales of computation. The second CNN 140 has a larger scale of computation than the first CNN 130. That is, the second CNN 140 has a higher recognition accuracy than the first CNN 130, and the first CNN 130 has a higher recognition speed than the second CNN 140.

[0061] The first CNN 130 is an image recognizer that performs image recognition in real time and has a recognition speed such that it can complete image recognition within the interval of the images output by the camera 190. The second CNN 140 is an image recognizer that performs image recognition only when instructed by the timing adjustment unit 170.

[0062] The first CNN 130 and the second CNN 140 have both performed pre-training using the same training data, and the first learning parameter 121, which is the learning result of the first CNN 130, and the second learning parameter, which is the learning result of the second CNN 140, are stored in the storage unit 120.

[0063] (6) Additional learning unit 180 The additional learning unit 180 performs learning of the first CNN 130 using the additional learning data 123 stored in the storage unit 120, and updates the first learning parameter 121 using the learning result.

[0064] 1.3 Recognition result comparison unit 150 The recognition result comparison unit 150 acquires the recognition results of the first CNN 130 and the second CNN 140, compares the two, and outputs whether the recognition results match.

[0065] 1.4 Data collection unit 160 When the results of the comparison in the recognition result comparison unit 150 are different, the data collection unit 160 acquires the input image input to the first CNN 130 and the second CNN 140, and the recognition result of the second CNN 140, generates additional learning data 123 with the input image as the learning image 123a and the recognition result of the second CNN 140 as the correct data 123b for the learning image, and stores it in the storage unit 120.

[0066] 1.5 Timing adjustment unit 170 The timing adjustment unit 170 controls (determines) the timing for operating the second CNN 140 and the additional learning unit 180.

[0067] 1.6 Operations FIG. 5 is a flowchart showing the operations of the image recognition system 1.

[0068] At the start of processing, the control unit 110 substitutes 0 for a control variable n indicating the frame number of one frame of the image acquired from the camera as an initial setting (step S101).

[0069] The control unit 110 determines whether an interrupt for ending the processing has occurred (step S102), and if it has occurred (step S102: Yes), ends the processing.

[0070] If there is no interrupt indicating the end of processing (step S102: No), the control unit 110 acquires an image for one frame (camera image) from the camera 190 (step S103). The frame number of the camera image matches the control variable n. For example, when the control variable n is 1, the frame number of the camera image is 1.

[0071] The control unit 110 inputs the camera image with frame number n to the first CNN 130 to perform image recognition (step S104), and the first CNN 130 outputs a recognition result for the camera image with frame number n (step S105).

[0072] Next, the timing adjustment unit 170 determines whether the remainder obtained by dividing the control variable n by the threshold T1 is 0 (step S106). If the determination result is true (step S106: Yes), it is determined to operate the second CNN, and if the determination result is false (step S106: No), it is determined not to operate the second CNN. Here, the threshold T1 is a variable that specifies the interval for operating the second CNN 140. When the output speed of the camera 190 is 30 FPS and the threshold T1 is 1800, the second CNN 140 is operated once every 1800 frames, that is, once a minute.

[0073] When operating the second CNN 140, the control unit 110 inputs the camera image with frame number n to the second CNN 140 to perform image recognition (step S107), and the second CNN 140 outputs a recognition result for the camera image with frame number n.

[0074] The recognition result comparison unit 150 respectively acquires the recognition results for the camera images with frame number n of the first CNN 130 and the second CNN 140, compares the two, and outputs the comparison result (step S108).

[0075] The data collection unit 160 acquires the comparison result by the recognition result comparison unit. When the two are different (step S109: Yes), the camera image with frame number n is used as the learning image 123a, and the recognition result of the second CNN 140 for the camera image with frame number n is used as the correct data 123b for the learning image 123a. The additional learning data 123 formed by pairing the learning image 123a and the correct data 123b is generated and stored in the storage unit 120 (step S110).

[0076] Next, the timing adjustment unit 170 determines whether the remainder obtained by dividing the control variable n by the threshold value T2 is 0 (step S111). If the determination result is true (step S111: Yes), it is determined that additional learning of the first CNN 130 is to be performed. If the determination result is false (step S111: No), it is determined that additional learning of the first CNN 130 is not to be performed. Here, the threshold value T2 is a variable that specifies the interval for performing additional learning of the first CNN 130. For example, when the output rate of the camera 190 is 30 FPS and the threshold value T2 is 18144000, additional learning of the first CNN 130 is performed once every 18144000 (= 30 (frames) × 60 (seconds) × 60 (minutes) × 24 (hours) × 7 (days)) frames, that is, once a week.

[0077] When performing additional learning of the first CNN 130, the additional learning unit 180 performs additional learning of the first CNN 130 using the additional learning data 123 stored in the storage unit 120 (step S112).

[0078] The control unit 110 substitutes n + 1 into the control variable n and repeats the process from step S102.

[0079] Note that the following two processes can be executed in parallel.

[0080] (1) Processes of steps S102 to 105, 106, 111 to 113 (2) Processes of steps S107 to 110 Therefore, while the second CNN 140 is performing image recognition on the camera image with frame number n, the first CNN 130 performs image recognition on the camera images with frame numbers n+1, n+2, …….

[0081] 1.8 Effects The image recognition system 1 includes two image recognizers with different calculation scales, and executes image recognition on the same camera image. When one of the recognition results by the two image recognition units is correct and the other is incorrect, for the incorrect image recognition unit, the input image is classified as FP or FN. Here, FP means identifying that the input image contains a detection target even though it does not, and FN means, conversely, identifying that the input image does not contain a detection target even though it does. Generally, the purpose of learning in image recognition is to reduce such FPs and FNs. And an effective method for reducing FPs and FNs is to label the images classified as FPs and FNs correctly to generate learning data and perform additional learning so that similar images can be correctly recognized. According to the image recognition system 1 according to the present disclosure, images classified as such FPs and FNs can be easily collected as learning data. Also, for the labeling, by using the recognition result of the image recognition unit that was correct in the recognition, there is no need for the user to manually perform the labeling. Therefore, it is possible to reduce the burden on the user related to the generation of learning data.

[0082] 2. Supplementary As described above, the present invention has been described based on the embodiments, but it goes without saying that the present invention is not limited to the above-described embodiments, and it goes without saying that the following modification examples are included in the technical scope of the present invention.

[0083] (1) The image recognition system 1 in the above-described Embodiment 1 includes two image recognizers (the first CNN 130 and the second CNN 140) in the image recognition device 100 of the same housing. However, the configuration may be such that the two image recognizers are mounted on different terminal devices.

[0084] FIG. 6 is a block diagram showing the configuration of an image recognition system 2 in which two image recognizers are mounted on different terminal devices. As shown in the figure, the image recognition system 2 includes an edge terminal 300 and a server terminal 400.

[0085] The edge terminal 300 includes a control unit 310, a non-volatile storage unit 320, a sensor 330, a first CNN 340 (first recognition unit), an additional learning unit 350 (learning unit), a timing adjustment unit 360 (timing determination unit), and a communication unit 370.

[0086] The control unit 310 is composed of a CPU, a ROM, a RAM, etc. Computer programs and data stored in the ROM and the storage unit 320 are loaded into the RAM, and the CPU operates according to the computer programs and data on the RAM, thereby realizing each processing unit (the first CNN 340, the additional learning unit 350, and the timing adjustment unit 360), and controlling the sensor 330 and the communication unit 370.

[0087] The storage unit 320 is composed of a hard disk as an example. The storage unit 320 may be composed of a non-volatile semiconductor memory. The storage unit 320 stores the first learning parameter 321.

[0088] The sensor 330 is an imaging device such as a CMOS image sensor or a CCD image sensor, and outputs an image of a predetermined size by converting the light imaged on the imaging device into an electrical signal by photoelectric conversion. The sensor 330 outputs images at a predetermined rate. For example, it outputs images at 30 FPS.

[0089] The first CNN 340 has the same configuration as the first CNN 130 in Embodiment 1. The learning result of the first CNN 340 is stored in the storage unit 320 as the first learning parameter 321.

[0090] The additional learning unit 350 performs learning of the first CNN 340 using the additional learning data 422 received from the server terminal 400, and updates the first learning parameter 321 using the learning result.

[0091] The timing adjustment unit 360 controls the timing for operating the additional learning unit 350 and the second CNN 430 of the server terminal 400.

[0092] The communication unit 370 is a network interface for communicating with the server terminal 400. The edge terminal 300 transmits data such as captured images at the sensor 330 and recognition results of the first CNN 340 to the server terminal 400 via the communication unit 370. Also, the edge terminal 300 receives, via the communication unit 370, for example, additional learning data 422 from the server terminal 400.

[0093] The server terminal 400 includes a control unit 410, a non-volatile storage unit 420, a second CNN 430 (second recognition unit), a recognition result comparison unit 440 (comparison unit), a data collection unit 450 (collection unit), and a communication unit 460.

[0094] The control unit 410 is composed of a CPU, a ROM, a RAM, etc. Computer programs and data stored in the ROM and the storage unit 420 are loaded into the RAM, and the CPU operates according to the computer programs and data on the RAM to realize each processing unit (the second CNN 430, the recognition result comparison unit 440, and the data collection unit 450), and controls the communication unit 460.

[0095] The storage unit 420 is, for example, composed of a hard disk. The storage unit 420 may be composed of a non-volatile semiconductor memory. The storage unit 420 stores second learning parameters 421 and additional learning data 422. The additional learning data 422 includes learning images 422a and correct answer data 422b.

[0096] The second CNN 430 has the same configuration as the second CNN 140 of the first embodiment. The learning result of the second CNN 430 is stored in the storage unit 420 as the second learning parameters 421.

[0097] The recognition result comparison unit 440 has the same configuration as the recognition result comparison unit 150 in Embodiment 1. It acquires the recognition results of the first CNN 340 and the second CNN 430, compares the two, and outputs whether the recognition results match or not.

[0098] The data collection unit 450 has the same configuration as the data collection unit 160 in Embodiment 1. When the results of the comparison in the recognition result comparison unit 440 show that the two are different, it acquires the input images input to the first CNN 340 and the second CNN 430, and the recognition result of the second CNN 430, generates additional learning data 422 with the input images as learning images 422a and the recognition result of the second CNN 430 as correct data 422b for the learning images 422a, and stores it in the storage unit 420.

[0099] The communication unit 460 is a network interface for communicating with the edge terminal 300. The server terminal 400 receives data such as captured images by the sensor 330 and recognition results of the first CNN 340 from the edge terminal 300 via the communication unit 460. Also, the server terminal 400 transmits additional learning data 422, etc. to the edge terminal 300 via the communication unit 460.

[0100] FIG. 7 is a flowchart showing the operation of the edge terminal 300.

[0101] At the start of processing, the control unit 410 substitutes 0 for a control variable n indicating the frame number of one frame of image acquired from the camera as an initial setting (step S201).

[0102] The control unit 410 determines whether an interrupt for ending the process has occurred (step S202). If it has occurred (step S202: Yes), the process ends.

[0103] If an interrupt indicating the end of processing has not occurred (step S202: No), the control unit 410 acquires an image for one frame (sensor image) from the sensor 330 (step S203). The frame number of the sensor image matches the control variable n. For example, when the control variable n is 1, the frame number of the sensor image is 1.

[0104] The control unit 310 inputs the sensor image with frame number n to the first CNN 340 to execute image recognition (step S204), and the first CNN 340 outputs a recognition result for the camera image with frame number n (step S205).

[0105] Next, the timing adjustment unit 360 determines whether the remainder obtained by dividing the control variable n by the threshold T1 is 0 (step S206). If the determination result is true (step S206: Yes), it is determined to operate the second CNN 430, and if the determination result is false (step S206: No), it is determined not to operate the second CNN 430.

[0106] When operating the second CNN 430, the control unit 310 transmits the sensor image with frame number n and the recognition result of the first CNN 340 for the sensor image with frame number n to the server terminal 400 via the communication unit 370 (step S207).

[0107] Next, the timing adjustment unit 360 determines whether the remainder obtained by dividing the control variable n by the threshold T2 is 0 (step S208). If the determination result is true (step S208: Yes), it is determined to perform additional learning of the first CNN 340, and if the determination result is false (step S208: No), it is determined not to perform additional learning of the first CNN 340.

[0108] When performing additional learning of the first CNN 340, the control unit 310 transmits a request to acquire additional learning data 422 to the server terminal 400 via the communication unit 370, and receives the additional learning data 422 from the server terminal 400 as a response (step S209). The additional learning unit 350 performs additional learning of the first CNN 340 using the additional learning data 422 (step S210).

[0109] The control unit 310 substitutes n + 1 for the control variable n and repeats the process from step S202.

[0110] FIG. 8 is a flowchart showing the operation of the server terminal 400.

[0111] The control unit 410 waits until it receives data from the edge terminal 300 (step S301).

[0112] The control unit 410 determines whether it has received the sensor image of frame number n and the recognition result of the first CNN 340 for the sensor image of frame number n from the edge terminal 300 via the communication unit 460 (step S302).

[0113] When the control unit 410 has received the sensor image of frame number n and the recognition result of the first CNN 340 for the sensor image of frame number n (step S302: Yes), the control unit 410 inputs the sensor image of frame number n to the second CNN 430 and causes it to execute image recognition (step S303), and the second CNN 430 outputs the recognition result for the camera image of frame number n.

[0114] The recognition result comparison unit 440 respectively obtains the recognition results of the first CNN 430 and the second CNN 430 for the sensor image of frame number n, compares the two, and outputs the comparison result (step S304).

[0115] The data collection unit 450 acquires the comparison result by the recognition result comparison unit 440. When the two are different (step S305: Yes), the sensor image with frame number n is used as the learning image 422a, and the recognition result of the second CNN 430 for the sensor image with frame number n is used as the correct data 422b for the learning image 422a. Additional learning data 422 consisting of the learning image 422a and the correct data 422b is generated and stored in the storage unit 420 (step S306).

[0116] The control unit 410 determines whether it has received a request to acquire the additional learning data 422 from the edge terminal 300 via the communication unit 460 (step S307).

[0117] When a request to acquire the additional learning data 422 is received, the control unit transmits the additional learning data 422 to the edge terminal 300 via the communication unit 460 as a response.

[0118] Here, it is assumed that there is one edge terminal 300 for one server terminal 400, but a configuration with multiple edge terminals 300 may also be possible.

[0119] (2) In the above-described embodiment, it was described that the recognition result of the second CNN 140 is used as the correct data 123b. However, user input may be received and the correct data 123b may be corrected based on the user input.

[0120] (3) In the above-described embodiment, the first CNN 130 and the second CNN 140 are image recognizers for performing person detection. If a person is detected in the image input from the camera 190, a recognition result indicating that a person is included is output, and if not detected, a recognition result indicating that no person is included is output. However, the recognition result may be output as a numerical value such as a likelihood. In that case, when the difference between the two numerical values exceeds a predetermined threshold, the data collection unit 160 may determine that the recognition results are different and collect them as additional learning data.

[0121] (4) In the above-described embodiment, the timing adjustment unit 170 operates the second CNN 140 at a predetermined interval T1. However, the timing for operating the second CNN 140 is not limited to this. The timing for operating the second CNN 140 may be changed according to the learning proficiency of the first CNN 130. For example, in the initial stage of learning, the interval for operating the second CNN 140 may be shortened, and in the later stage of learning, the interval for operating the second CNN 140 may be lengthened. The learning proficiency may be, for example, based on the number of executions of image recognition in the first CNN 130. When the number of executions is less than a predetermined threshold, it may be regarded as the initial stage of learning, and when the number of executions is more than the predetermined threshold, it may be regarded as the later stage of learning. Also, based on the degree of coincidence between the results of the first CNN 130 and the second CNN 140, when the degree of coincidence is less than a predetermined threshold, it may be regarded as the initial stage of learning, and when the degree of coincidence is greater than the predetermined threshold, it may be regarded as the later stage of learning.

[0122] Further, user input may be received, and the interval T1 may be set based on the user input.

[0123] (5) In the above-described embodiment, in an image recognition system that performs image recognition using machine learning, two image recognizers with different calculation scales are provided, and image recognition is executed for the same camera image. However, this is not limited thereto.

[0124] (a) The object of learning and recognition may be voice data. In this case, examples of voice data may be music, human voices, natural sounds, etc. Examples of music may be classical music, ethnic music, pop music, Latin music, etc. Also, examples of human voices may be news voices, voices of lectures, voices of conversations, etc. Also, examples of natural sounds may be bird songs, wind sounds, river flow sounds, etc. In a voice recognition system that performs voice recognition using machine learning, two voice recognizers (first recognition unit and second recognition unit) with different calculation scales are provided, and voice recognition may be executed for the same voice data acquired from a voice data input device such as a microphone. For example, when the voice data is a human voice, a specific person's voice may be recognized.

[0125] (b) The object of learning and recognition may be character data in natural language recognition processing. In this case, examples of character data may be conversation sentences, literary works, newspaper articles, papers, etc. Also, examples of conversation sentences may be conversation sentences in Japanese, English, Italian, etc. Further, examples of literary works may be works such as poems, novels, stories, plays, reviews, essays, etc. Also, examples of newspaper articles may be political news, economic news, scientific news, etc. In a natural language processing system that performs natural language recognition using machine learning, two natural language recognizers (a first recognition unit and a second recognition unit) with different computational scales may be provided, and natural language recognition may be executed on the same character data. For example, when the character data is a newspaper article, the main part (subject) and the predicate part may be recognized and extracted from a sentence in the newspaper article.

[0126] (6) The above-described embodiments and modification examples may be combined respectively.

Industrial Applicability

[0127] The data collection device according to the present disclosure can reduce the burden on the user related to the generation of learning data and is useful as a data collection device for collecting learning data.

Explanation of Signs

[0128] 100 Image recognition device 110 Control unit 120 Storage unit 130 First CNN 140 Second CNN 150 Recognition result comparison unit 160 Data collection unit 170 Timing adjustment unit 180 Additional learning unit 190 Camera

Claims

1. A data collection device for collecting learning data for a data recognition model, comprising: a first recognition unit; a second recognition unit different from the first recognition unit; a comparison unit that compares the recognition result of the first recognition unit and the recognition result of the second recognition unit for input data; a collection unit that collects the input data as learning data according to the comparison result of the comparison unit; a timing determination unit that determines the timing for operating the second recognition unit; The first recognition unit takes less processing time to recognize the same object than the second recognition unit, The first recognition unit performs constant data recognition, The timing determination unit determines the timing for operating the second recognition unit to be less frequent than the constant data recognition by the first recognition unit and at a constant interval equal to or longer than the processing time of the second recognition unit, The second recognition unit performs data recognition at the timing determined by the timing determination unit. A data collection device.

2. The calculation scale of the first recognition unit is smaller than that of the second recognition unit. The data collection device according to Claim 1.

3. The collection unit collects the recognition result of the second recognition unit as learning data indicating correct data for the input data. The data collection device according to Claim 2.

4. Furthermore, it includes a learning unit that performs additional learning of the first recognition unit using the learning data collected by the collection unit. The data collection device according to Claim 3.

5. The learning unit corrects the correct data according to an external input. The data collection device according to Claim 4.

6. The comparison unit determines whether the recognition result of the first recognition unit and the recognition result of the second recognition unit are different, The collection unit collects the input data as learning data when the recognition result of the first recognition unit and the recognition result of the second recognition unit are different. The data collection device according to any one of Claims 1 - 5.

7. The comparison unit determines whether the difference between the recognition result of the first recognition unit and the recognition result of the second recognition unit is equal to or greater than a predetermined threshold value, The collection unit collects the input data as learning data when the difference is equal to or greater than the threshold value. The data collection device according to any one of Claims 1 - 5.

8. The timing determination unit determines the timing at fixed intervals. The data collection device according to any one of Claims 1 - 7.

9. The timing determination unit determines the timing according to the learning proficiency in the first recognition unit. ​ The data collection device according to any one of claims 1 to 7.

10. The timing determination unit determines the timing according to an external input The data collection device according to any one of claims 1 to 7.

11. An edge terminal including the first recognition unit, A server terminal including the second recognition unit, and is composed of The data collection device according to any one of claims 1 to 10.

12. Furthermore, it includes one or more second edge terminals, and the second edge terminal includes a recognition unit having the same configuration as the first recognition unit The data collection device according to claim 11.

13. The first recognition unit and the second recognition unit each perform image recognition, voice recognition, or natural language recognition The data collection device according to any one of claims 1 to 12.

14. A data collection method in which a computer executes a process of collecting learning data for a data recognition model, A first recognition step of obtaining a recognition result by a first recognition unit for input data, A second recognition step of obtaining a recognition result by a second recognition unit different from the first recognition unit for the input data, A comparison step of comparing the recognition result of the first recognition unit and the recognition result of the second recognition unit, A collection step of collecting the input data as learning data according to the comparison result of the comparison step, A determination step of determining the timing for operating the second recognition unit, Including, The first recognition unit takes less processing time to recognize the same object than the second recognition unit, The first recognition unit performs constant data recognition, The determination step determines the timing for operating the second recognition unit to be less frequent than the constant data recognition by the first recognition unit and at a constant interval equal to or longer than the processing time of the second recognition unit, The second recognition unit performs data recognition at the timing determined in the determination step Data collection method.

15. The calculation scale of the first recognition unit is smaller than that of the second recognition unit The data collection method according to claim 14.

16. The collection step collects the recognition result of the second recognition unit as learning data indicating correct data for the input data The data collection method according to claim 15.

17. Furthermore, additional learning of the first recognition unit is performed using the learning data collected in the collection step The data collection method according to claim 16.

18. The comparison step determines whether the recognition result of the first recognition unit is different from the recognition result of the second recognition unit. The collection step collects the input data as learning data when the recognition result of the first recognition unit is different from the recognition result of the second recognition unit. The data collection method according to any one of claims 14-17.

19. The comparison step determines whether the difference between the recognition result of the first recognition unit and the recognition result of the second recognition unit is equal to or greater than a predetermined threshold. The collection step collects the input data as learning data when the difference is equal to or greater than the threshold. The data collection method according to any one of claims 14-17.

Citation Information

Patent Citations

  • Neural network device

    JP1993061843A

  • Personal attribute estimation system, personal attribute estimation device and personal attribute estimation method

    JP2012252507A

  • Product recognition system, learned model, and product recognition method

    JP2018169752A

  • Information processing device, control method and program for information processing device

    JP2019046094A

  • Object recognition camera system, relearning system, and object recognition program

    JP2020052484A