Machine learning method, machine learning program, machine learning device, and information processing device
The method adjusts sequence data size to generate training data sets with varied intervals, improving robustness and accuracy in machine learning models without advanced preprocessing, addressing the limitations of existing methods.
Patent Information
- Application Number
- JP2023562139
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-11-18
- Filing Date
- 2022-08-19
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2042-08-19
AI Technical Summary
Existing machine learning methods require advanced preprocessing and are not robust against changes in sequence direction conditions, necessitating large amounts of high-quality training data for accurate object recognition.
A machine learning method that generates training data by adjusting sequence data size based on predetermined conditions, using supervised learning to create a learning model with improved robustness to sequence direction changes, without requiring advanced preprocessing.
Enables easy generation of multiple training data sets with different intervals, resulting in a learning model that is robust to sequence direction variations, enhancing accuracy in object recognition tasks.
Smart Images

Figure 0007806808000001 
Figure 0007806808000002 
Figure 0007806808000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a machine learning method, a machine learning program, a machine learning device, and an information processing device. [Background technology]
[0002] To achieve a certain level of object recognition accuracy through machine learning such as deep learning, it is generally necessary to use a large amount of high-quality training data for learning. To achieve this, there is a method for increasing data while maintaining quality, as described in Non-Patent Document 1, by setting a sampling rate according to the analysis results. In the method of Non-Patent Document 1, in order to suppress a decline in recognition accuracy due to differences in pronunciation (conversation, reading, speech), the entropy of the voice data is analyzed at 15 msec intervals, and the sampling rate is set according to the analysis results to generate training data for learning. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Amber Afshan, Jinxi Guo, Soo Jin Park, Vijay Ravi, Alan McCree, Abeer Alwan," Variable frame rate-based data augmentation to handle speaking-style variability for automatic speaker verification," Cornell University, USA, Sat, 8 Aug 2020, Internet (URL: https: / / arxiv.org / abs / 2008.03616) Summary of the Invention [Problem to be solved by the invention]
[0004] The technology of Non-Patent Document 1 is related to audio data, and requires advanced interpolation processing as preprocessing.
[0005] The present invention has been made to solve these problems. That is, an object of the present invention is to provide a machine learning device and a machine learning method that generate training data easily without requiring advanced preprocessing, and that generate a learning model with improved robustness against changes in conditions in the sequence direction by using the generated training data for training. [Means for solving the problem]
[0006] The above-mentioned problems of the present invention are solved by the following means.
[0007] (1) A machine learning method for generating a learning model for extracting features of a target, comprising: Step (a) of acquiring sequence data; Step (b) performing preprocessing of the sequence data to adjust the sequence direction size based on predetermined conditions, thereby generating a plurality of adjusted sequence data pieces with different sequence direction intervals from one sequence data piece; and (c) performing supervised learning using the plurality of adjusted series data generated in step (b) to generate a learning model.
[0008] (2) In the step (a), the label of the sequence data is acquired together with the sequence data; The machine learning method according to (1) above, wherein in step (c), the label of one of the sequence data is applied to the plurality of adjusted sequence data to perform supervised learning.
[0009] (3) The machine learning method according to (1) or (2) above, wherein in step (b), size adjustment conditions are automatically set based on the predetermined conditions.
[0010] (4) the series of data acquired in step (a) is time-series image data obtained by photographing a target object in a photographing area; The machine learning method according to any one of (1) to (3), wherein the learning model is a learning model for extracting features of a target object.
[0011] (5) In the machine learning method described in (4) above, in the step (b), the condition for adjusting the size is set according to the sampling rate or number of frames of the sequence data as the predetermined condition.
[0012] (6) further comprising a step (d) of acquiring external information relating to the imaging environment; The machine learning method according to (4) or (5) above, wherein in the step (b), the predetermined condition is set based on the external information to adjust the size.
[0013] (7) The machine learning method described in (6) above, wherein the external information is information regarding the moving speed of the object and the specifications of the camera that captures the capture area.
[0014] (8) The method further includes a step (e) of analyzing the sequence data based on a predetermined condition and detecting one or more key frames in which a point of interest of a target object exists from among a plurality of frames constituting the sequence data; A machine learning method described in any of (4) to (7) above, wherein in step (b), one reference frame is set from the key frames detected in step (e), and the size adjustment is performed based on the reference frame.
[0015] (9) The machine learning method according to (8) above, wherein in step (b), size adjustment conditions are set according to the number of key frames detected in step (e).
[0016] (10) The machine learning method according to (8) or (9), wherein in step (b), only the keyframes are targeted for size adjustment.
[0017] (11) A machine learning method according to any one of (8) to (10), wherein in step (b), the size adjustment method is made different before and after the reference frame in the arrangement direction of the sequence data.
[0018] (12) A machine learning device that generates a learning model for extracting features of a target, an acquisition unit that acquires sequence data; a pre-processing unit that performs pre-processing of the sequence data to adjust the sequence direction size based on predetermined conditions, thereby generating a plurality of adjusted sequence data pieces with different intervals in the sequence direction from one sequence data piece; a learning unit that performs supervised learning using the plurality of adjusted series data generated by the preprocessing unit to generate a learning model; A machine learning device comprising:
[0019] (13) The acquiring unit acquires the label of the sequence data together with the sequence data, The machine learning device according to (12), wherein the learning unit applies the label of one of the sequence data to the plurality of adjusted sequence data to perform supervised learning.
[0020] (14) The machine learning device according to (12) or (13), wherein the preprocessing unit automatically sets conditions for size adjustment based on the predetermined conditions.
[0021] (15) The series data acquired by the acquisition unit is time-series image data obtained by photographing a target object in a photographing area, The machine learning device according to any one of (12) to (14), wherein the learning model is a learning model for extracting features of a target object.
[0022] (16) The machine learning device according to (15) above, wherein the preprocessing unit sets the size adjustment conditions according to the sampling rate or number of frames of the sequence data as the predetermined conditions.
[0023] (17) The acquisition unit further acquires external information related to a shooting environment, The machine learning device according to (15) or (16), wherein the preprocessing unit sets a condition for the size adjustment based on the external information as the predetermined condition.
[0024] (18) The machine learning device according to (17) above, wherein the external information is information about the moving speed of the object and the specifications of the camera that captures the capture area.
[0025] (19) The apparatus further includes a detection unit that analyzes the sequence data based on a predetermined condition and detects one or more key frames in which a point of interest of a target object exists from among a plurality of frames constituting the sequence data, The machine learning device according to any one of (15) to (18), wherein the preprocessing unit sets one reference frame from among the key frames detected by the detection unit and performs the size adjustment based on the reference frame.
[0026] (20) The machine learning device according to (19), wherein the preprocessing unit sets conditions for size adjustment according to the number of keyframes detected by the detection unit.
[0027] (21) The machine learning device according to (19) or (20), wherein the preprocessing unit targets only the keyframes as targets for the size adjustment.
[0028] (22) The machine learning device according to any one of (19) to (21), wherein the preprocessing unit differs in the method of size adjustment before and after the reference frame in the arrangement direction of the sequence data.
[0029] (23) A machine learning program for causing a computer to execute the machine learning method described in any one of (1) to (11) above.
[0030] (24) an acquisition unit for acquiring sequence data; an extraction unit that extracts features of a target using a learning model trained by the machine learning method according to any one of (1) to (11) above; and an output unit that outputs the extraction result.
[0031] According to the machine learning method and machine learning device of the present invention, sequence data is acquired and preprocessed to adjust the sequence size based on predetermined conditions, thereby generating multiple adjusted sequence data sets with different intervals in the sequence direction from a single sequence data set, and then supervised learning is performed using the generated adjusted sequence data sets to generate a learning model. This allows multiple training data sets with different intervals in the sequence data to be easily generated without requiring advanced preprocessing, and training is performed using this data, thereby generating a learning model with improved robustness to changes in sequence direction conditions. [Brief explanation of the drawings]
[0032] [Figure 1] 1 is a diagram showing a schematic configuration of an information processing apparatus according to an embodiment of the present invention; [Figure 2] 2 is a side view showing an example of an object to be inspected by the information processing device shown in FIG. 1. [Figure 3] FIG. 1 is a block diagram showing a configuration of an information processing device. [Figure 4] FIG. 2 is a functional block diagram showing the flow of data in the machine learning device realized by the functioning of a control unit. [Figure 5] 10 is an example of sequence data. [Figure 6] 1 is a flowchart showing machine learning processing of the machine learning device. [Figure 7A] 10 is a subroutine flowchart showing the size adjustment condition setting process of step S53. [Figure 7B] 10 is a subroutine flowchart showing the size adjustment condition setting process of step S53 in another example. [Figure 8] 10 is an example of a plurality of adjusted series data generated by preprocessing. [Figure 9] 10 is an example of adjusted series data generated under different size adjustment conditions. [Figure 10] FIG. 10 is a schematic diagram illustrating a machine learning method using adjusted series data. [Figure 11] FIG. 1 is a functional block diagram showing the flow of data in inspection processing of an information processing device using a learning model generated by machine learning. [Figure 12] 10 is a flowchart showing an inspection process of the information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0033] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, the scope of the present invention is not limited to the disclosed embodiments. In the description of the drawings, the same elements are denoted by the same reference numerals, and duplicate explanations will be omitted. Furthermore, the dimensional proportions in the drawings are exaggerated for the convenience of explanation and may differ from the actual proportions.
[0034] FIG. 1 is a diagram showing a schematic configuration of an inspection system 1 including an information processing device according to this embodiment.
[0035] The inspection system 1 comprises a sequence data input device 30 and an information processing device 10, which are interconnected via a network 90 such as a LAN for mutual communication. The sequence data input device 30 generates and inputs sequence data. The sequence data input device 30 includes a camera 310. In addition to the camera 310, the sequence data input device 30 may also include a detection device that performs continuous observations and outputs detection data, such as a three-dimensional distance measurement sensor such as a LiDAR (Light Detection and Ranging), a temperature sensor or a pressure sensor installed in a factory, etc., and an HDD (hard disk drive) that records the sequence data obtained from these devices. The information processing device 10 functions as a machine learning device, performing machine learning using the sequence data from the sequence data input device 30 and generating a machine learning model.
[0036] (series data) The sequential data is a data group in which multiple data are arranged according to predetermined order information. For example, there are photographed data (time-series image data) obtained by photographing with the camera 310, three-dimensional data in which two-dimensional image data is arranged according to position information in a direction perpendicular to the two dimensions, audio data in which human voices are arranged in time series, and ranging point cloud data obtained from a three-dimensional ranging sensor. In the following, the sequential data will be described using photographed data (video) obtained by photographing with the camera 310 as an example.
[0037] FIG. 2 shows an example of a predetermined object to be inspected by the inspection system 1. In the example shown in FIG. 2, the object is a long sheet metal member, which is conveyed by a belt conveyor (not shown) from the right side to the left side in the conveying direction in FIG. 2. In this embodiment, the information processing device 10 of the inspection system 1 extracts defects in the surface paint of the sheet metal member (shown as points of interest in FIG. 2) as features of the object (object) and outputs the extraction results. Note that the object is not limited to this, and may be a plurality of vehicles or other products themselves, or some of the components for these products, continuously conveyed by a belt conveyor, and shape features of the object (such as product defects or missing parts) may be extracted and the extraction results may be output.
[0038] 3 is a block diagram showing the configuration of the information processing device 10. The information processing device 10 includes a control unit 11, a storage unit 12, an operation display unit 13, and a communication unit 14. These are interconnected via signal lines such as a bus for exchanging signals.
[0039] The control unit 11 functions as a machine learning device and includes multiple CPUs, multiple GPUs (Graphics Processing Units), RAM, ROM, etc., and controls each device and performs machine learning according to a program. The information processing device 10 may be an on-premise server or a cloud server using a commercial cloud service. Furthermore, some of the functions of the information processing device 10 (for example, only the machine learning device functions) may be implemented by the cloud server.
[0040] The storage unit 12 is composed of a semiconductor memory or a magnetic memory such as a hard disk, which stores various programs and data in advance. A machine learning model 200 (also referred to as a trained model) trained, generated, and updated through machine learning is stored in the storage unit 12. The storage unit 12 also stores the following three types of data d1 to d3: (d1) a large number of sequential data generated by the sequential data input device 30, (d2) external information, and (d3) extraction conditions for points of interest. Each piece of sequential data (d1) is associated with a label (correct answer label) and stored. The external information (d2) is information about the shooting environment, such as the sampling rate or frame rate (FPS) of the camera 310, or the moving speed of the object, i.e., the conveyor belt speed. Alternatively, if the sequential data is audio data, it is the sampling rate. The extraction conditions (d3) are preset rules, and a rule-based algorithm using these may be, for example, an image processing algorithm for detecting points of interest, such as pattern matching or edge detection. This extraction condition (d3) or an algorithm using this condition is used in the detection process of the detection unit 112, which will be described later.
[0041] The operation display unit 13 is, for example, a touch panel display that displays various information and accepts various inputs from the user. The user can set the above-mentioned shooting environment (external information) via the operation display unit 13. Labeling of each sequence data may be performed via the operation display unit 13, or may be performed in a pre-labeling process using a rule-based algorithm or a machine learning model. These settings or assigned information are stored in the storage unit 12.
[0042] The communication unit 14 is an interface for transmitting and receiving data via a network, and performs communication according to standards such as Ethernet, Bluetooth (registered trademark), and IEEE802.11 (Wi-Fi).
[0043] 4 is a functional block diagram showing the flow of data in the machine learning device realized by the functioning of the control unit 11. The control unit 11 functions as an acquisition unit 111 in cooperation with the communication unit 14. The control unit 11 also functions as a detection unit 112, a preprocessing unit 113, and a learning unit 114.
[0044] (Acquisition part 111) The acquisition unit 111 acquires external information and a plurality of training data from the sequence data input device 30 or from the storage unit 12. The training data is made up of a set of a plurality of sequence data and labels.
[0045] (Detection unit 112) The detection unit 112 receives sequence data from the acquisition unit 111. FIG. 5 shows an example of sequence data. The sequence data here is image data captured during a predetermined interval (time t-α to t+β). For example, if the sequence data is one second of video data captured by a camera 310 at 30, 60, or 120 FPS, one sequence data will consist of 30, 60, or 120 frames (still images). The predetermined interval and FPS can be set as appropriate. In the following description, one sequence data will consist of 60 frames. The sequence data used as training data is generated in advance by capturing images of an object to be inspected that has a target area of interest (for example, a paint unevenness defect) while moving it on a conveyor belt. In the example shown in FIG. 5, the target area of interest (paint unevenness) is shown in white for ease of understanding.
[0046] Furthermore, the detection unit 112 detects frames (hereinafter also referred to as key frames) containing points of interest from among the multiple frames constituting the sequence data based on a preset extraction condition (d3). The detection result is sent to the pre-processing unit 113. For example, if the sequence data is composed of 60 frames (numbered 1 to 60), the frame numbers of the key frames are sent.
[0047] (Preprocessing unit 113) The preprocessing unit 113 adjusts the size of the sequence data in the sequence direction based on predetermined conditions to generate multiple pieces of adjusted sequence data with different intervals in the sequence direction. The predetermined conditions include the following predetermined conditions A1 to A3 (hereinafter, these are also collectively referred to as predetermined condition A).
[0048] (A1) Sampling rate or number of frames, (A2) external information (for example, movement speed or camera specifications), and (A3) key frame information.
[0049] (A1) is information indicating the characteristics of the sequence data, which is stored in advance in the storage unit 12, and is set by the user, for example. (A2) External information is obtained from the sequence data input device 30. (A3) Key frame information is information on the number of key frames and / or the position of the reference frame (see below), and is determined based on the key frame information obtained from the detection unit 112.
[0050] Furthermore, the preprocessing unit 113 sets a reference frame from the sequence data. This reference frame is set from among the key frames detected by the detection unit 112. For example, in the example of FIG. 5, the key frame at time t is set as the reference frame. This reference frame is set according to a predetermined condition (hereinafter also referred to as predetermined condition B) that is set in advance. For example, predetermined condition B may be a method of setting the center position of a sequence of multiple key frames as the reference frame when multiple key frames are detected, or a method of setting the time (position) when the edge of the target location (the boundary between black and white in the figure) reaches near the center of the image as the reference frame.
[0051] The preprocessing unit 113 sets size adjustment conditions based on predetermined conditions A1 and A2. For example, if the speed range of the object's movement is predetermined in the inspection device, the preprocessing unit 113 increases the variety of images that can occur within that speed range (increasing the types of adjusted sequential data). Similarly, the preprocessing unit 113 increases the variety of images that can occur within the speed range to cover all frame rates depending on the camera specifications. As another example, the preprocessing unit 113 determines the number of frames in which the object of interest exists in the shooting area (hereinafter referred to as "existing frames" and "number of existing frames") based on the size of the object of interest (the size of the object of interest relative to the shooting area in the direction of movement) and the speed of movement based on the predetermined conditions A1 and A2, and performs size adjustment based on the number of frames to generate multiple pieces of adjusted sequential data. Note that the number of existing frames often coincides with the number of key frames. For example, the preprocessing unit 113 performs size adjustment by extracting several frames before and after the reference frame, or by thinning out one or two frames within the range of several frames before and after the reference frame.
[0052] Furthermore, size adjustment may be performed only on existing frames, or different size adjustment methods may be used for frames before and after the reference frame in the arrangement direction of the sequence data. Size adjustment may also involve interpolation or extrapolation, in addition to thinning. For example, if the number of existing frames is less than a predetermined number, an intermediate frame is generated by interpolating using frames before and after. Specific examples of size adjustment will be described later.
[0053] (Learning Section 114) The learning unit 114 performs machine learning through supervised learning using multiple pieces of adjusted sequence data with different intervals in the sequence direction after size adjustment and the labels assigned to the adjusted pieces of sequence data as training data, and generates or updates the machine learning model 200. Here, one label assigned to one piece of sequence data is commonly applied to multiple pieces of adjusted sequence data generated based on this sequence data.
[0054] (machine learning processing) The machine learning method according to this embodiment will be described below with reference to Fig. 6 to Fig. 11. In this embodiment, an example will be described in which, in photographic data consisting of 60 pieces of time-series image data as sequence data, the amount of data per piece is reduced by thinning processing in the time direction as size adjustment in the sequence direction.
[0055] Figure 6 is a flowchart showing the machine learning process executed by the control unit 11 functioning as a machine learning device. In the process of Figure 6, steps S51 to S55 are performed to generate multiple pieces of adjusted sequence data with different intervals from each piece of sequence data. This increases the number of samples (number of training data) and reduces the amount of each piece of data. Then, in step S56, machine learning is performed using the adjusted sequence data to generate and update a learning model.
[0056] (Step S51) Here, the acquisition unit 111 of the control unit 11 acquires external information. The external information is either acquired directly from the sequence data input device as described above, or set by the user via the operation / display unit 13 and stored in the storage unit 12.
[0057] (Step S52) Here, the acquiring unit 111 acquires training data directly from the sequence data input device 30 or stored in the storage unit 12. The training data is made up of multiple sequence data, and each sequence data is assigned a label.
[0058] (Step S53) Here, the preprocessing unit 113 automatically sets the size adjustment conditions either independently or in cooperation with the detection unit 112. Fig. 7A is a subroutine flowchart showing the size adjustment condition setting process of step S53 in one example, and Fig. 7B is a subroutine flowchart showing the size adjustment condition setting process of step S53 in another example.
[0059] (First example) (Step S611) As shown in FIG. 7A, the preprocessing unit 113 sets a plurality of size adjustment conditions based on a predetermined condition A. For example, the predetermined condition A is the number of frames constituting the sequential data (predetermined condition A3), and the larger the number of frames, the higher the thinning rate. For example, if there are 30 frames, the thinning rate is set to 1 or 2 frames, and if there are 60 frames, the thinning rate is set to 1 to 3 frames. For example, if there are 60 frames (0 to 59) and one frame is thinned out, odd-numbered frames are deleted and even-numbered frames (0, 2, 4, 6, ...) are used to generate adjusted sequential data with the data volume reduced to half. If two frames are thinned out, every third frame is used to generate adjusted sequential data with the data volume reduced to one-third (0, 3, 6, 9, ...). This completes the processing of FIG. 7A, and the process returns to the processing of FIG. 6 (RETURN).
[0060] (Another example) (Step S621) In another example shown in FIG. 7B, the detection unit 112 extracts key frames from the sequence data based on extraction condition (d3).
[0061] (Step S622) Here, the preprocessing unit 113 sets a reference frame. This reference frame is set from among the key frames detected by the detection unit 112 in step S621. For example, in Fig. 5, the frame at time t is set as the reference frame based on the above-mentioned predetermined condition B.
[0062] (Step S623) The preprocessing unit 113 sets a plurality of size adjustment conditions based on a combination of the number of existing frames determined by the predetermined condition A1 or A2 and the predetermined condition A3 (key frame information), or based on the predetermined condition A3 alone (see FIG. 8, which will be described later). This ends the processing in FIG. 7B, and the process returns to the processing in FIG. 6 (RETURN).
[0063] (Step S54) Referring again to Fig. 6, the preprocessing unit 113 performs size adjustment based on the size adjustment conditions set in step S53, and generates, from one piece of sequence data, a plurality of pieces of adjusted sequence data with different intervals.
[0064] Fig. 8 shows an example of multiple pieces of adjusted series data generated by preprocessing. The frames shown in Fig. 8 correspond to Fig. 5, and in Fig. 8, adjusted frames are enclosed in solid-line rectangular frames, and other frames (i.e., frames to be deleted) are shown in light gray. In the adjusted series data x1 shown in Fig. 8(a), as the adjustment condition set in step S623, a predetermined continuous section (three frames in the figure) of frames is extracted within the range of the number of existing frames, centered on the reference frame (time t) enclosed in a dashed rectangular frame (frames at times t-1, t, and t+1).
[0065] Furthermore, in the adjusted series data x2 in Fig. 8(b), as another adjustment condition set in step S623, three frames are extracted by thinning out one frame around the reference frame (times t-2, t, and t+2). Note that, although the example shown in Fig. 8 shows an example of adjusted series data consisting of three frames, this is not limiting and the data may consist of more than three frames. Furthermore, the adjusted series data may consist of only presence frames (or key frames) in which a point of interest exists in the shooting area, or may include frames other than the presence frames.
[0066] 9A and 9B show examples of adjusted sequential data generated under different size adjustment conditions. FIG. 9A shows adjusted sequential data generated by thinning out one frame centered on the reference frame (t), FIG. 9B shows adjusted sequential data generated by thinning out two frames centered on the reference frame (t), and FIG. 9C shows adjusted sequential data generated by a method (random thinning) in which the adjustment method is different before and after the reference frame (t). Specifically, in the example of FIG. 9C, the thinning rate is different before and after the reference frame. The adjustment conditions shown in FIG. 9 may be applied in combination with the adjustment conditions shown in FIG. 8 or in place of FIG. 8.
[0067] (Step S55) If the size adjustment for all the training data has not been completed, the control unit 11 returns the process to step S52 and repeats the subsequent processes. If the size adjustment for all the data sets of the training data has been completed, the control unit 11 proceeds to step S56.
[0068] (Step S56) The control unit 11, which is a machine learning device, reads the adjusted sequence data and labels after sample adjustment as training data and performs machine learning. FIG. 10 is a schematic diagram for explaining a machine learning method using adjusted sequence data. By the processing up to step S55, multiple pieces of adjusted sequence data x1 and x2 are generated from one piece of sequence data x to which a label X is associated. Furthermore, the label X associated with the original sequence data x is commonly applied to these adjusted sequence data x1 and x2. Note that while FIG. 10 shows an example in which two pieces of adjusted sequence data x1 and x2 are generated, three or more pieces of adjusted sequence data with different intervals may be generated and used for machine learning. For example, four pieces of adjusted sequence data x1 to x4, each spaced k apart in the sequence direction, may be generated as shown in FIGS. 8 and 9.
[0069] Similar size adjustments are also performed on numerous other sequence data, thereby adjusting the size of the sequence data and increasing the number of samples. These adjusted sequence data are then input to the neural network as training data for the machine learning device. The machine learning device (control unit 11) then compares the neural network's estimation results for the adjusted sequence data with the labels and adjusts parameters based on the comparison results. For example, by performing a process called back-propagation, the parameters are adjusted and updated so that the error in the comparison results is reduced. This is repeated for the target training data (adjusted sequence data) to advance machine learning. When machine learning using the target training data is completed, the learning model 200 is stored in the storage unit 12 and the process ends (END).
[0070] Although the machine learning method using a neural network configured by combining perceptrons has been described, the present invention is not limited to this, and various other supervised learning methods can be used, such as random forests, support vector machines (SVMs), boosting, Bayesian network linear discriminant analysis, and nonlinear discriminant analysis.
[0071] As described above, the machine learning method or machine learning device according to this embodiment acquires sequence data and labels, performs preprocessing on the sequence data to adjust the sequence size based on predetermined conditions, and generates multiple adjusted sequence data sets with different intervals in the sequence direction from a single sequence data set. Supervised learning is then performed using the labels and the multiple adjusted sequence data sets generated by the preprocessing unit to generate a learning model. This allows multiple training data sets with different intervals in the sequence data to be easily generated and used for training without requiring advanced preprocessing, thereby generating a learning model with improved robustness to changes in sequence direction conditions.
[0072] For example, when a learning model trained on a production line in one factory where products are moving on a conveyor belt is applied to another production line in another factory, it is expected that accuracy will decrease unless machine learning is performed for each conveyor belt with a different speed. Even in such a situation, by performing machine learning as in this embodiment, it is possible to use sequence data obtained from an object moving on a conveyor belt at a single speed and train using multiple adjusted sequence data with different intervals, thereby enabling a single learning model to handle diverse situations where the speeds are different. In particular, the machine learning device or machine learning method according to this embodiment can be preferably applied to generating a learning model for extracting features from objects for which moving speed or movement itself is not a primary parameter.
[0073] (Inspection processing using learning models) 11 and 12, an inspection process using the machine learning model 200 generated by the machine learning process of Fig. 6 will be described. Fig. 11 is a functional block diagram showing the flow of data in the inspection process of the information processing device 10, and Fig. 12 is a flowchart showing the inspection process of the information processing device 10.
[0074] As shown in FIG. 11, the control unit 11 of the information processing device 10 functions as an acquisition unit 116, an extraction unit 117, and an output unit 118. The acquisition unit 116 has the same function as the acquisition unit 111, and acquires sequence data obtained by photographing an object as shown in FIG. 2 using the camera 310 of the sequence data input device 30. The extraction unit 117 extracts features of the object (object) from the sequence data using a learning model 600. The output unit 118 outputs the extraction results.
[0075] (Step S71) The acquisition unit 116 acquires the sequence data. In the example of Fig. 2, captured images are sent from the camera 310 in real time, and are divided into sequence data for each predetermined period.
[0076] (Step S72) The extraction unit 117 develops the machine learning model 200 stored in the storage unit 12 and performs an appearance inspection using this model. The inspection result is output as a score.
[0077] (Step 73) The output unit 118 outputs a determination result according to the score. For example, the output unit 118 outputs a determination result of whether the object is defective or non-defective to the operation display unit 13 or the like according to the score of the object.
[0078] In this way, the information processing device 10 according to the present embodiment extracts features from sequence data containing objects using a learning model and outputs the extraction results, thereby enabling highly accurate determination of the features of the objects, i.e., whether the product is good or bad.
[0079] The configurations of the machine learning device and information processing device described above are the main configurations described in order to explain the features of the above-mentioned embodiments, but are not limited to the above configurations and can be modified in various ways within the scope of the claims. Furthermore, configurations that are included in general machine learning devices or information processing devices are not excluded.
[0080] In addition, some steps in the above-described flowcharts may be omitted, other steps may be added, the order of some of the steps may be changed, some steps may be executed simultaneously, or one step may be divided into multiple steps and executed.
[0081] The means and methods for performing the various processes in the information processing device 10 described above can be realized by either a dedicated hardware circuit or a programmed computer. The program may be provided by a computer-readable recording medium such as a USB memory or a DVD (Digital Versatile Disc)-ROM, or may be provided online via a network such as the Internet. In this case, the program recorded on the computer-readable recording medium is typically transferred to and stored in a storage unit such as a hard disk. The program may be provided as standalone application software, or may be incorporated into the software of the device as a function of the device.
[0082] This application is based on a Japanese patent application (Patent Application No. 2021-187584) filed on November 18, 2021, the disclosure of which is incorporated by reference in its entirety. [Explanation of symbols]
[0083] 1 inspection system, 10. Information processing equipment 11 Control unit (machine learning device) 111 Acquisition Department 112 Detector 113 Pretreatment section 114 Learning Department 116 Acquisition Department 117 Extraction part 118 Output section 12 Storage section 13 Operation display section 14 Communications Department 200 Learning Models 30-series data input device 310 Camera
Claims
1. A machine learning method for generating a learning model for extracting features of a target, comprising: (a) acquiring sequence data; (b) performing preprocessing of the sequence data to adjust the sequence-direction size based on predetermined conditions, thereby generating a plurality of adjusted sequence data pieces having different intervals in the sequence direction from one sequence data piece; and (c) performing supervised learning using the plurality of adjusted series data generated in the step (b) to generate a learning model; the series of data acquired in step (a) is time-series image data obtained by photographing a target object in a photographing area, and the learning model is a learning model for extracting features of the target object; The method further includes a step (e) of analyzing the sequence data based on a predetermined condition, and detecting one or more key frames in which a point of interest of a target object exists from among a plurality of frames constituting the sequence data, A machine learning method that executes processing in which, in step (b), one reference frame is set from among the key frames detected in step (e), and the size adjustment is performed based on the reference frame.
2. In the step (a), a label of the sequence data is acquired together with the sequence data; The machine learning method according to claim 1 , wherein in step (c), supervised learning is performed by applying the label of one of the sequence data to the plurality of adjusted sequence data.
3. The machine learning method according to claim 1 , wherein in the step (b), a condition for size adjustment is automatically set based on the predetermined condition.
4. The machine learning method according to claim 1 , wherein in the step (b), the condition for adjusting the size is set according to a sampling rate or a number of frames of the sequence data as the predetermined condition.
5. Further, the method includes a step (d) of acquiring external information related to a shooting environment, The machine learning method according to claim 1 , wherein in the step (b), a condition for the size adjustment is set as the predetermined condition based on the external information.
6. The machine learning method according to claim 5 , wherein the external information is information relating to a moving speed of the object and specifications of a camera that captures the image of the image capture area.
7. The machine learning method according to claim 1 , wherein in the step (b), a condition for size adjustment is set according to the number of the key frames detected in the step (e).
8. The machine learning method according to claim 1 , wherein in the step (b), only the key frames are subjected to the size adjustment.
9. The machine learning method according to claim 1 , wherein in the step (b), the size adjustment method is made different before and after the reference frame in the arrangement direction of the sequence data.
10. A machine learning device that generates a learning model for extracting features of a target, an acquisition unit that acquires sequence data; a pre-processing unit that performs pre-processing of the sequence data to adjust the size in the sequence direction based on predetermined conditions, thereby generating a plurality of adjusted sequence data pieces having different intervals in the sequence direction from one sequence data piece; a learning unit that performs supervised learning using the plurality of adjusted series data generated by the preprocessing unit to generate a learning model; the series data acquired by the acquisition unit is time-series image data obtained by photographing a target object in a photographing area, the learning model is a learning model for extracting features of a target object, a detection unit that analyzes the sequence data based on a predetermined condition and detects one or more key frames in which a point of interest of a target object exists from among a plurality of frames that constitute the sequence data; the preprocessing unit sets one reference frame from among the key frames detected by the detection unit, and performs the size adjustment based on the reference frame. Machine learning device.
11. the acquiring unit acquires the sequence data together with a label of the sequence data; The machine learning device according to claim 10 , wherein the learning unit performs supervised learning by applying the label of one of the sequence data to the plurality of adjusted sequence data.
12. The machine learning device according to claim 10 , wherein the preprocessing unit automatically sets a condition for size adjustment based on the predetermined condition.
13. The machine learning device according to claim 10 , wherein the preprocessing unit sets the condition for adjusting the size according to a sampling rate or a number of frames of the sequence data as the predetermined condition.
14. The acquisition unit further acquires external information related to a shooting environment, The machine learning device according to claim 10 , wherein the preprocessing unit sets a condition for the size adjustment as the predetermined condition based on the external information.
15. The machine learning device according to claim 14 , wherein the external information is information relating to a moving speed of the object and specifications of a camera that captures the image of the image capture area.
16. The machine learning device according to claim 10 , wherein the preprocessing unit sets a condition for size adjustment depending on the number of the key frames detected by the detection unit.
17. The machine learning device according to claim 10 , wherein the preprocessing unit performs the size adjustment only on the key frames.
18. The machine learning device according to claim 10 , wherein the preprocessing unit uses different size adjustment methods before and after the reference frame in the arrangement direction of the sequence data.
19. A machine learning program for causing a computer to execute the machine learning method according to any one of claims 1 to 9.
20. an acquisition unit that acquires sequence data; an extraction unit that extracts features of a target using a learning model trained by the machine learning method according to any one of claims 1 to 9; and an output unit that outputs the extraction result.
Citation Information
Patent Citations
Image processing device and learned model
WO2019069629A1
Teacher data extending device, teacher data extending method, and program
WO2020070876A1
Method for generating neural network model, and control device using neural network model
WO2020178936A1
Information processing device and information processing method
WO2021100267A1